
Household and commercial environments
Egocentric Data forHomes & Workplaces.
License existing household or workplace video, or commission capture around your model, tasks, rights, and delivery.

Direct answer
What Is Egocentric Data?
Egocentric data is visual, sensor, and contextual information captured from the perspective of the person or machine performing a task. It is commonly called first-person, actor-view, or point-of-view data. A complete dataset can combine video with audio, IMU, depth, gaze, pose, task labels, timestamps, provenance, and rights metadata.
The defining feature is not the camera brand or mount. It is that the observation moves with the actor and exposes information available near the point of action. That makes egocentric data valuable for hand-object interaction and procedural activity, while also creating blind spots, motion, privacy, and embodiment constraints that must be designed around.
Updated

Fine Tool Use and Object Transformation
A dual-tool preparation sequence showing peeling, sectioning, and precision slicing with clear hand, object, and tool trajectories.
- Dual-tool sequence
- Fine motor control
- Visible state change
Peel the vegetable with controlled strokesSection and expose the interiorSlice prepared pieces into strips
Why Robotics Teams Use Egocentric Data
The value comes from mapping an observable signal to a specific learning or evaluation objective, not from collecting first-person footage by the hour.
| Observable signal | Potential use | What still may be missing |
|---|---|---|
| Hands, objects, tools, and contact transitions | Interaction representations, affordances, grasp or tool-use events | Force, torque, tactile feedback, precise 3D geometry |
| Task steps and visible state changes | Procedure understanding, temporal segmentation, world models | Hidden state, intent, success criteria not visible in frame |
| Language aligned with activity | VLA pretraining, narration grounding, retrieval | Robot-executable actions and embodiment constraints |
| Natural mistakes and recovery behavior | Robustness analysis, failure classification, evaluator data | Controlled counterfactuals and safety-certified behavior |
| Long-horizon activity in real environments | Memory, planning, state-transition learning | Complete scene coverage when objects leave the actor’s view |
Human demonstrations are useful, but they are not robot commands
A person’s video does not automatically contain the action space required by a robot. Human kinematics, camera motion, morphology, reach, contact, and control frequency differ from the target platform. Egocentric video may supply visual preconditions, task structure, language alignment, object-state supervision, or representation learning while robot-native rollouts supply executable actions.
Compare Human and Robot Data ↗See How Egocentric Data Fits Embodied AI ↗Write the learning objective before the recording protocol
Specify whether the model needs task retrieval, step boundaries, object interactions, action language, pose, imitation targets, failure cases, or evaluation footage. That decision controls mount position, frame rate, sensors, task variation, annotations, and which events must remain visible.
Read the Robot-Learning Guide ↗Egocentric, Exocentric, or Paired?
Choose the viewpoint from the information the model must observe. Two cameras are not automatically better; each added stream needs a model-facing reason and an acceptance test.
| Viewpoint | Best at showing | Common blind spots | Use it when |
|---|---|---|---|
| Egocentric | Local hand-object activity, actor-aligned context, near-field state changes | Whole-body pose, off-screen events, stable world coordinates, self-occlusion | The learning target depends on what is visible near the actor’s point of action |
| Exocentric | Body motion, workspace layout, approach paths, other agents, stable scene context | Fine contact, objects blocked by the actor, details outside camera coverage | Geometry, safety zones, multi-agent behavior, or full-body motion matters |
| Synchronized ego/exo | Cross-view correspondence, complementary visibility, human-to-observer mapping | Clock drift, calibration error, storage and review complexity | The model or evaluation explicitly needs aligned actor and observer views |
A useful pilot measures critical-action visibility, occlusion frequency, framing drift, synchronization error where applicable, reviewer cost, privacy exposure, and the percentage of takes that pass every required gate.
What an Egocentric Dataset Can Contain
Add only the signals that the consuming model and evaluation path can use. Every modality creates new synchronization, calibration, privacy, storage, and QA obligations.
RGB Video
Native media, resolution, frame rate, lens, exposure, motion, field of view, and decode checks.
Depth and Pose
Depth, hand pose, body pose, object pose, calibration, confidence, and coordinate conventions.
IMU
Timestamped accelerometer and gyroscope streams with alignment, units, sampling, and drift requirements.
Understand RGB + IMU Sync ↗Gaze
Calibration results, validity flags, synchronized gaze samples, coordinate mapping, and privacy-aware use.
Audio and Narration
Instructions, narration, transcription, language tags, timing, consent, and voice-handling policy.
Episodes and Tasks
Session, take, task, step, success, failure, interruption, and terminal-state definitions.
Annotations
Actions, objects, contacts, temporal segments, scene attributes, confidence, reviewer, and ontology version.
Rights and Provenance
Consent reference, privacy decision, permitted use, lineage, distribution tier, retention, and release status.
Schema and Dataset Card
Field definitions, checksums, splits, known limitations, transformations, version history, and loader evidence.
Separate facts from interpretation
Device, time, codec, file integrity, and checksums are capture facts. Task labels, interactions, pose, quality decisions, and privacy eligibility are derived records with their own methods and versions. Mixing everything into one mutable table destroys auditability.
Review Egocentric Metadata Types and Schema ↗Explore Data Annotation Services ↗Start with a written data contract
Define field names, units, allowed nulls, timebase, coordinate frames, episode boundaries, ontology versions, release compatibility, and buyer-loader expectations before collection scales.
Use the Dataset Specification Template ↗How Egocentric Data Collection Works
A defensible program defines acceptance before volume. Recorded, uploaded, technically valid, privacy-eligible, annotated, and release-approved are different states.
- 01
Translate the Model Gap
Define the target task, observable signal, environments, people, objects, variations, failure cases, and intended model or evaluation use.
- 02
Write the Data Specification
Fix required modalities, episode structure, metadata, annotations, rights, delivery, acceptance thresholds, and known unsupported conditions.
- 03
Run a Visibility and Privacy Pilot
Test mounts on real bodies and workspaces. Check critical actions, framing drift, bystanders, screens, documents, reflections, audio, and stop rules.
- 04
Collect With Versioned Instructions
Track protocol, hardware, task, environment, and instruction versions. Give contributors clear repeat, stop, safety, and privacy escalation rules.
- 05
Validate Media and Metadata
Check decode, duration, dimensions, frame rate, timestamps, required fields, duplicates, modality alignment, and file relationships.
- 06
Review Visibility, Task, and Privacy
Sample across tasks, contributors, environments, devices, and time. Quarantine ambiguous or ineligible material rather than silently repairing it.
- 07
Annotate and Release
Apply versioned guidelines, measure disagreement and rework, connect each asset to governance records, and publish a release manifest and dataset card.
- 08
Prove Buyer Ingest
Load a representative release in the target framework, decode every modality, traverse episodes, inspect batches, and record failures before scale approval.
Critical hands, objects, contacts, or state changes leave frame or remain occluded.
Corrupt media, unexpected frame rate, broken timestamps, missing files, or duplicate episodes.
Wrong task, incomplete terminal state, unsafe execution, uncontrolled interruption, or missing failure label.
Unapproved people, screens, documents, voices, locations, or environments enter the release.
Dataset sourcing
License Existing Data or Commission Custom Collection.
EGXO offers pre-existing egocentric datasets for controlled licensing. Current inventory covers household and commercial environments. This inventory was collected through licensed GIG Rewards collection programs operated with telco partners. Each delivery still requires buyer-specific confirmation of eligible media, permitted use, privacy status, and license terms.
Commission custom collection when the missing signal is specific: proprietary workflows, deployment environments, rare failure cases, required sensors, controlled variation, buyer-defined annotations, or a different rights package. A hybrid strategy can use existing data first, identify weak slices, and spend custom budget only where the mismatch matters.
Compare cost per accepted, ingestible, rights-eligible episode, not cost per recorded hour or downloaded terabyte.
Request Dataset Access ↗Plan a Custom Egocentric Dataset ↗Scope Custom Collection ↗Six Real Household Task Samples

Mechanical Disassembly and Reassembly
- Part disassembly
- Bimanual maintenance
- Reassembly sequence
Release the fan-guard fastenerClean the detached blade assemblyAlign and reseat the guard

Long-Horizon Meal Preparation
- Long-horizon task
- Ingredient state change
- Pan and tool control
Slice ingredients for the next stageTransfer aromatics into the heated panAdd and combine the vegetables

Contact-Rich Cleaning and Container Handling
- Contact-rich cleaning
- Object rotation
- Nesting and stacking
Rotate and scrub the container surfacesDry and inspect the cleaned containerAlign and nest containers for storage

Deformable-Object Manipulation
- Deformable objects
- Bimanual coordination
- Changing geometry
Lift and orient the garmentAlign fabric edges with both handsFold and compress the garment
What the Public Previews Prove
The six public previews are excerpts, not the complete dataset. Qualified organizations can purchase licensed access to existing EGXO datasets; evaluation, model-training, commercial-use, retention, redistribution, and other rights are defined separately in writing.
- Assets
- 6 video excerpts + posters
- Duration
- 10-15 seconds each
- Frame size
- 960 × 540
- Frame rate
- 24 FPS
- Tasks
- Mechanical maintenance, timed cooking, fine tool use, long-horizon cooking, cleaning and stacking, laundry folding
- Audio
- Removed
- Metadata
- Identifying source fields removed
- Approval
- Consented release privacy QA + site-owner publication approval
| Task family | Visible learning signal | Representative challenge |
|---|---|---|
| Fan maintenance | Disassembly, component cleaning, and ordered reassembly | Part state, fastener visibility, and alignment |
| Timed pancake cooking | Mixture state, controlled pouring, and tool-mediated turning | Heat timing and continuous visual state change |
| Vegetable preparation | Peeling, sectioning, precision slicing, and dual-tool control | Safety, contact visibility, and fast motion |
| Long-horizon stir-fry | Ingredient preparation, transfer to heat, and pan combination | Long-range phase dependencies and changing state |
| Container cleaning and stacking | Surface scrubbing, inspection, alignment, and nesting | Contact, reflections, and spatial fit |
| Laundry folding | Bimanual coordination and deformable-object state | Continuously changing geometry |
The machine-readable manifest contains sanitized descriptive metadata for the six previews. Training and redistribution rights require a separate written license.
Download Preview Manifest (.json)Dataset-card excerpt
Make the Release Understandable Without Reverse Engineering
A buyer should receive a machine-readable manifest and a human-readable dataset card describing the release, intended use, task coverage, modalities, transformations, limitations, rights, version, and validation evidence.
{
"release": "project-name-v1.0",
"viewpoint": "head-mounted-egocentric",
"modalities": ["rgb_video", "task_metadata"],
"episode_unit": "one task attempt",
"acceptance": {
"critical_actions_visible": true,
"privacy_review": "passed",
"buyer_ingest": "passed"
},
"known_limitations": [
"no force or torque measurements",
"human actions require embodiment mapping"
]
}Privacy, Rights, and Provider Due Diligence
First-person cameras can expose bystanders, screens, documents, voices, reflections, addresses, private spaces, and behavior unrelated to the requested task.
Collection authority
Which notice, consent, contractual, or other lawful framework covers the contributor, incidental people, environment, and intended use?
Rights scope
Does the release permit training, evaluation, annotation access, derived models, retention, vendor sharing, and the buyer’s intended deployment?
Privacy controls
How are capture minimization, exclusion zones, stop rules, review, redaction, access tiers, incidents, and deletion handled?
Traceability
Can each delivered asset be connected to protocol, capture, QA, annotation, privacy, rights, transformation, and release versions?
Acceptance evidence
What was measured, on which population, under which thresholds, and what proportion was rejected, quarantined, or left unverified?
Delivery proof
Has a representative sample passed the buyer’s loader, modality checks, episode traversal, batch inspection, and release documentation review?
From requirement to pilot
Bring the Model Gap. Leave With a Testable Data Brief.
Share the task, deployment environment, required viewpoint, modalities, annotations, rights, format, and success criteria. The next step is a focused pilot, not a fictional volume promise.
Common Egocentric Data Questions
These answers define the category. Project-specific commitments belong in the specification, pilot results, and written proposal.
What is egocentric data?
Egocentric data is recorded from the perspective of the person or machine performing a task. It usually includes first-person video and may add audio, depth, IMU, gaze, pose, language, task structure, and synchronized metadata.
Is egocentric data the same as first-person video?
First-person video is the most common egocentric modality, but a training dataset can include additional sensors, annotations, provenance, rights records, and delivery metadata. A video folder alone is not a complete egocentric dataset.
What is egocentric data used for in robotics?
Teams use it to study hand-object interaction, procedural task structure, state changes, affordances, long-horizon activity, language grounding, representation learning, world models, and evaluation. The useful target depends on what the capture actually makes observable.
Can human egocentric video train a robot directly?
Sometimes it supplies useful visual or task supervision, but ordinary human video does not contain robot joint commands, force, torque, or a compatible action space. Many systems still require embodiment mapping, teleoperation, robot rollouts, simulation, or robot-native demonstrations.
When is exocentric data better?
An external viewpoint is often better when whole-body pose, workspace geometry, other agents, stable world coordinates, approach trajectories, or events outside the actor’s field of view are central to the model objective.
When should ego and exo cameras be synchronized?
Use paired capture when the learning or evaluation task needs cross-view correspondence, recovery of occluded events, body and workspace context alongside actor-aligned detail, or transfer between human and robot perspectives. Define clock, drift, calibration, and dropped-frame tolerances before scaling.
Which camera position is best for egocentric data?
There is no universal best mount. Head and glasses views follow attention and viewing direction; chest views can be more stable; wrist views can reveal close manipulation but lose scene context. Run a visibility pilot on the real tasks, people, and workspaces.
What annotations can an egocentric dataset contain?
Common layers include task and step boundaries, actions, objects, hand-object interactions, narration, transcripts, gaze, pose, scene attributes, quality results, privacy decisions, protocol versions, and provenance. Annotation density should follow the training or evaluation objective.
Should a team use a public or custom egocentric dataset?
Use public data for shared benchmarks, research comparison, and fast baselines. Use custom collection when deployment tasks, environments, sensors, labels, licenses, privacy controls, or failure cases do not match. A hybrid strategy often provides the best balance.
How should egocentric data quality be measured?
Measure model-relevant visibility, media integrity, task completion, coverage, timestamp alignment, annotation fitness, privacy eligibility, rights status, schema validity, release lineage, and successful ingestion by the buyer’s actual pipeline.
Which delivery formats are available?
Native media with structured manifests, RLDS, LeRobot, and custom exports can be evaluated against the collection and buyer pipeline. The final format should be confirmed through a representative ingest test rather than assumed from a format name.
What should a buyer request before commissioning collection?
Request representative samples, a task and coverage specification, capture protocol, acceptance criteria, schema, dataset-card outline, privacy and rights summary, versioning policy, delivery plan, and a pilot that exercises the intended loader.
Can a company buy off-the-shelf egocentric video from EGXO?
Yes. Qualified organizations can request pricing and licensed access to EGXO's first-party household or commercial and workplace egocentric video. Current task coverage, available volume, permitted use, license terms, security requirements, and delivery are confirmed against the submitted use case.
Can EGXO run a custom egocentric data collection project?
Yes. EGXO can scope buyer-defined collection around required tasks, environments, viewpoints, contributor criteria, modalities, annotations, rights, acceptance gates, scale, and delivery format. Availability, timing, pricing, and final commitments are confirmed in writing after the project brief is reviewed.
Primary research context
Public Datasets and Formats That Define the Category
Choose a buying path
Start With Available Data or Build What Is Missing.
Send the short brief with your model objective, required tasks, rights, delivery needs, and timing.