Quote-ready evidence
Key Statistics
Use these figures with the named source and dated methodology. Published scale is not the same as usable training yield.
- 44,711 hDreamDojo-HV human video
Published egocentric pretraining scale across 6,015 tasks. Source [1]
- 1M+AgiBotWorld-Beta trajectories
The official card reports 2,976.4 hours from 100 robots. Source [2]
- 1M+Open X-Embodiment trajectories
Pooled from 60 datasets across 22 robot embodiments. Source [3]
- 76KDROID demonstrations
The official project reports 350 hours across 564 scenes. Source [4]
- 10K hEgocentric-10K factory video
Human first-person visual data, not robot-native control. Source [5]
- 3,670 hEgo4D daily-life video
Broad perception coverage from 923 participants. Source [6]
- 1,680 hEgoLive reported stereo video
A new workflow-focused release whose terms and annotations require review. Source [7]
- 12releases tracked
A selective register split by data class and preserved source unit.
Current Release Snapshot
| Release | Data class | Published scale | Access and rights signal |
|---|---|---|---|
| DreamDojo-HV | Human egocentric video | 44,711 hours | Project and paper report the corpus; code license does not establish video rights |
| Egocentric-10K | Human egocentric video | 10,000 hours | Gated Hugging Face access; card license and gated terms both need review |
| Ego4D | Human egocentric video | 3,670 hours | Approved credentials and dataset license agreement |
| EgoLive | Human stereo egocentric video | 1,680 hours | Marketplace route reported; complete dataset terms require verification |
| EgoVerse | Human demonstrations with mixed tooling | 1,362 hours | Living release; pin provenance and exact dataset terms |
| Ego-Exo4D V2 | Synchronized human ego/exo | 1,286.30 video hours | Agreement covers research and commercial use with restrictions |
| EgoDex | Human egocentric plus 3D pose | 829 hours | CC BY-NC-ND |
| AgiBotWorld-Beta | Robot-native demonstrations | 1M+ trajectories / 2,976.4 hours | Gated; CC BY-NC-SA 4.0 shown on card |
| Open X-Embodiment | Robot-native multi-embodiment | 1M+ trajectories | Component datasets retain their own terms |
| DROID | Robot-native teleoperation | 76K trajectories / 350 hours | Open dataset and quickstart; verify current component terms |
| HoloAssist | Interactive human assistance | 169 hours | Public download; CDLA v2 |
| EPIC-KITCHENS-100 | Human egocentric video | 100 hours | Public downloader; CC BY-NC 4.0 |
Robotics Data Is Scaling Along Two Different Axes
Source-backed context[1] DreamDojo official project and paper[2] AgiBotWorld-Beta official dataset card[3] Open X-Embodiment official project[4] DROID official dataset project
Human egocentric video now reaches tens of thousands of published hours, while robot-native collections reach one million or more trajectories. DreamDojo-HV reports 44,711 hours of human video. AgiBotWorld-Beta and Open X-Embodiment each report more than one million robot trajectories. DROID reports 76,000 trajectories across 350 hours.
Those top-line numbers cannot be ranked on one axis. Human video supplies visual, semantic, behavioral, and physical-world breadth. Robot demonstrations supply embodiment-specific states and actions. A trajectory is not a fixed amount of time, and an hour of human video is not a control sequence.
Human Video Provides Breadth Before Action Grounding
Source-backed context[1] DreamDojo official project and paper[5] Egocentric-10K official dataset card[6] Ego4D official project[7] EgoLive paper[8] EgoVerse official project[9] Ego-Exo4D V2 documentation[10] Apple EgoDex repository[11] HoloAssist official project[12] EPIC-KITCHENS official project
Human egocentric sources are strongest for world-model pretraining, task decomposition, language grounding, object-state understanding, hand-object interaction, visual representation learning, and motion priors. They normally stop short of robot-native joint state, gripper commands, rewards, or contact forces.
The practical bridge can be post-training on robot actions, a latent-action model, retargeted human motion, cross-embodiment co-training, or a separate policy-learning stage. The bridge must be stated; calling human video robot-action data is technically false.
Robot-Native Releases Add Control at a Different Cost
Source-backed context[2] AgiBotWorld-Beta official dataset card[3] Open X-Embodiment official project[4] DROID official dataset project
AgiBotWorld-Beta exposes action and proprioceptive structures at enormous scale, Open X-Embodiment standardizes data from 22 embodiments, and DROID holds hardware constant while increasing scene diversity. These sources are closer to policy learning because they carry a control or state signal.
They also inherit embodiment, teleoperation, schema, hardware, and component-license constraints. A million pooled trajectories do not guarantee compatibility with one target robot, and a single-platform dataset does not automatically generalize across embodiments.
Access and Rights Still Fragment the Market
Source-backed context[1] DreamDojo official project and paper[2] AgiBotWorld-Beta official dataset card[3] Open X-Embodiment official project[5] Egocentric-10K official dataset card[8] EgoVerse official project[11] HoloAssist official project[12] EPIC-KITCHENS official project
The tracked releases use signed agreements, gated dataset cards, non-commercial Creative Commons terms, component-specific licenses, public downloads, marketplaces, and project pages that do not establish complete corpus rights. Code and data remain separate legal objects.
A procurement record should pin the exact release, source agreement, commercial and model-training permissions, derivative rules, vendor access, redistribution, retention, and any obligation inherited from a pooled component dataset.
The Tracker Is Built to Change
Implementation guidanceEGXO guidance for translating the research into a project specification.
This is a dated release register, not a permanent ranking. EGXO reviews the CSV monthly and after material first-party announcements. New entries need an official source, preserved unit, access signal, license signal, and a limitation that prevents the headline figure from becoming empty hype.
- Use publisher scale as a discovery signal, not a usable-yield metric
- Compare hours with hours and trajectories with trajectories
- Separate human observation data from robot action data
- Verify the current license and access route before ingestion
- Pin the release and reproduce one representative loader test
Primary Sources and Further Reading
- [1] DreamDojo official project and paper ↗
- [2] AgiBotWorld-Beta official dataset card ↗
- [3] Open X-Embodiment official project ↗
- [4] DROID official dataset project ↗
- [5] Egocentric-10K official dataset card ↗
- [6] Ego4D official project ↗
- [7] EgoLive paper ↗
- [8] EgoVerse official project ↗
- [9] Ego-Exo4D V2 documentation ↗
- [10] Apple EgoDex repository ↗
- [11] HoloAssist official project ↗
- [12] EPIC-KITCHENS official project ↗
These sources inform the category-level guidance above. Project-specific requirements are defined with the buyer.