Household and commercial environments

Egocentric Data forHomes & Workplaces.

License existing household or workplace video, or commission capture around your model, tasks, rights, and delivery.

First-person view of hands mixing batter, pouring it into a pan, and flipping a pancake.
Visible phaseMix batter to a consistent state960 × 540 · 24 FPS · 15s

Choose the Environment.

Household and commercial work create different tasks, privacy risks, scene structure, object distributions, and deployment questions.

Household data

Homes, Daily Tasks, and Manipulation

Review EGXO's growing first-party household catalogue, six public previews, and documented 10-hour evaluation release.

10 hours
Current evaluation release
111 videos
Privacy-reviewed release
49 families
Household task coverage
Explore Household Data

Commercial data

Real Workplaces and Multi-View Activity

Review current commercial source inventory or scope new collection around the workflows, sites, viewpoints, privacy rules, and rights your model needs.

  • POV, frontal, side, and whole-body source viewpoints
  • Existing inventory and buyer-defined collection routes
  • Paired claims released only after timing evidence passes
Explore Commercial Data

For procurement systems: choose household, commercial, or both, then select existing dataset access or custom collection. Final availability, pricing, rights, and timing are confirmed in writing.

Direct answer

What Is Egocentric Data?

Egocentric data is visual, sensor, and contextual information captured from the perspective of the person or machine performing a task. It is commonly called first-person, actor-view, or point-of-view data. A complete dataset can combine video with audio, IMU, depth, gaze, pose, task labels, timestamps, provenance, and rights metadata.

The defining feature is not the camera brand or mount. It is that the observation moves with the actor and exposes information available near the point of action. That makes egocentric data valuable for hand-object interaction and procedural activity, while also creating blind spots, motion, privacy, and embodiment constraints that must be designed around.

Updated

First-person view of hands peeling, sectioning, and slicing a vegetable.
Visible phasePeel the vegetable with controlled strokes960 × 540 · 24 FPS · 15s

Fine Tool Use and Object Transformation

A dual-tool preparation sequence showing peeling, sectioning, and precision slicing with clear hand, object, and tool trajectories.

  • Dual-tool sequence
  • Fine motor control
  • Visible state change

Peel the vegetable with controlled strokesSection and expose the interiorSlice prepared pieces into strips

Why Robotics Teams Use Egocentric Data

The value comes from mapping an observable signal to a specific learning or evaluation objective, not from collecting first-person footage by the hour.

Observable signalPotential useWhat still may be missing
Hands, objects, tools, and contact transitionsInteraction representations, affordances, grasp or tool-use eventsForce, torque, tactile feedback, precise 3D geometry
Task steps and visible state changesProcedure understanding, temporal segmentation, world modelsHidden state, intent, success criteria not visible in frame
Language aligned with activityVLA pretraining, narration grounding, retrievalRobot-executable actions and embodiment constraints
Natural mistakes and recovery behaviorRobustness analysis, failure classification, evaluator dataControlled counterfactuals and safety-certified behavior
Long-horizon activity in real environmentsMemory, planning, state-transition learningComplete scene coverage when objects leave the actor’s view

Human demonstrations are useful, but they are not robot commands

A person’s video does not automatically contain the action space required by a robot. Human kinematics, camera motion, morphology, reach, contact, and control frequency differ from the target platform. Egocentric video may supply visual preconditions, task structure, language alignment, object-state supervision, or representation learning while robot-native rollouts supply executable actions.

Compare Human and Robot Data See How Egocentric Data Fits Embodied AI

Write the learning objective before the recording protocol

Specify whether the model needs task retrieval, step boundaries, object interactions, action language, pose, imitation targets, failure cases, or evaluation footage. That decision controls mount position, frame rate, sensors, task variation, annotations, and which events must remain visible.

Read the Robot-Learning Guide

Egocentric, Exocentric, or Paired?

Choose the viewpoint from the information the model must observe. Two cameras are not automatically better; each added stream needs a model-facing reason and an acceptance test.

ViewpointBest at showingCommon blind spotsUse it when
EgocentricLocal hand-object activity, actor-aligned context, near-field state changesWhole-body pose, off-screen events, stable world coordinates, self-occlusionThe learning target depends on what is visible near the actor’s point of action
ExocentricBody motion, workspace layout, approach paths, other agents, stable scene contextFine contact, objects blocked by the actor, details outside camera coverageGeometry, safety zones, multi-agent behavior, or full-body motion matters
Synchronized ego/exoCross-view correspondence, complementary visibility, human-to-observer mappingClock drift, calibration error, storage and review complexityThe model or evaluation explicitly needs aligned actor and observer views

A useful pilot measures critical-action visibility, occlusion frequency, framing drift, synchronization error where applicable, reviewer cost, privacy exposure, and the percentage of takes that pass every required gate.

What an Egocentric Dataset Can Contain

Add only the signals that the consuming model and evaluation path can use. Every modality creates new synchronization, calibration, privacy, storage, and QA obligations.

Visual

RGB Video

Native media, resolution, frame rate, lens, exposure, motion, field of view, and decode checks.

Geometry

Depth and Pose

Depth, hand pose, body pose, object pose, calibration, confidence, and coordinate conventions.

Motion

IMU

Timestamped accelerometer and gyroscope streams with alignment, units, sampling, and drift requirements.

Understand RGB + IMU Sync
Attention

Gaze

Calibration results, validity flags, synchronized gaze samples, coordinate mapping, and privacy-aware use.

Language

Audio and Narration

Instructions, narration, transcription, language tags, timing, consent, and voice-handling policy.

Structure

Episodes and Tasks

Session, take, task, step, success, failure, interruption, and terminal-state definitions.

Labels

Annotations

Actions, objects, contacts, temporal segments, scene attributes, confidence, reviewer, and ontology version.

Governance

Rights and Provenance

Consent reference, privacy decision, permitted use, lineage, distribution tier, retention, and release status.

Delivery

Schema and Dataset Card

Field definitions, checksums, splits, known limitations, transformations, version history, and loader evidence.

Separate facts from interpretation

Device, time, codec, file integrity, and checksums are capture facts. Task labels, interactions, pose, quality decisions, and privacy eligibility are derived records with their own methods and versions. Mixing everything into one mutable table destroys auditability.

Review Egocentric Metadata Types and Schema Explore Data Annotation Services

Start with a written data contract

Define field names, units, allowed nulls, timebase, coordinate frames, episode boundaries, ontology versions, release compatibility, and buyer-loader expectations before collection scales.

Use the Dataset Specification Template

How Egocentric Data Collection Works

A defensible program defines acceptance before volume. Recorded, uploaded, technically valid, privacy-eligible, annotated, and release-approved are different states.

  1. 01

    Translate the Model Gap

    Define the target task, observable signal, environments, people, objects, variations, failure cases, and intended model or evaluation use.

  2. 02

    Write the Data Specification

    Fix required modalities, episode structure, metadata, annotations, rights, delivery, acceptance thresholds, and known unsupported conditions.

  3. 03

    Run a Visibility and Privacy Pilot

    Test mounts on real bodies and workspaces. Check critical actions, framing drift, bystanders, screens, documents, reflections, audio, and stop rules.

  4. 04

    Collect With Versioned Instructions

    Track protocol, hardware, task, environment, and instruction versions. Give contributors clear repeat, stop, safety, and privacy escalation rules.

  5. 05

    Validate Media and Metadata

    Check decode, duration, dimensions, frame rate, timestamps, required fields, duplicates, modality alignment, and file relationships.

  6. 06

    Review Visibility, Task, and Privacy

    Sample across tasks, contributors, environments, devices, and time. Quarantine ambiguous or ineligible material rather than silently repairing it.

  7. 07

    Annotate and Release

    Apply versioned guidelines, measure disagreement and rework, connect each asset to governance records, and publish a release manifest and dataset card.

  8. 08

    Prove Buyer Ingest

    Load a representative release in the target framework, decode every modality, traverse episodes, inspect batches, and record failures before scale approval.

Visibility

Critical hands, objects, contacts, or state changes leave frame or remain occluded.

Integrity

Corrupt media, unexpected frame rate, broken timestamps, missing files, or duplicate episodes.

Task

Wrong task, incomplete terminal state, unsafe execution, uncontrolled interruption, or missing failure label.

Privacy

Unapproved people, screens, documents, voices, locations, or environments enter the release.

Dataset sourcing

License Existing Data or Commission Custom Collection.

EGXO offers pre-existing egocentric datasets for controlled licensing. Current inventory covers household and commercial environments. This inventory was collected through licensed GIG Rewards collection programs operated with telco partners. Each delivery still requires buyer-specific confirmation of eligible media, permitted use, privacy status, and license terms.

Commission custom collection when the missing signal is specific: proprietary workflows, deployment environments, rare failure cases, required sensors, controlled variation, buyer-defined annotations, or a different rights package. A hybrid strategy can use existing data first, identify weak slices, and spend custom budget only where the mismatch matters.

Compare cost per accepted, ingestible, rights-eligible episode, not cost per recorded hour or downloaded terabyte.

Request Dataset Access Plan a Custom Egocentric Dataset Scope Custom Collection
First-person view of two hands disassembling, cleaning, and reassembling an electric fan.
Visible phaseRelease the fan-guard fastener960 × 540 · 24 FPS · 15s

Mechanical Disassembly and Reassembly

  • Part disassembly
  • Bimanual maintenance
  • Reassembly sequence

Release the fan-guard fastenerClean the detached blade assemblyAlign and reseat the guard

First-person view of hands preparing ingredients and cooking a stir-fry.
Visible phaseSlice ingredients for the next stage960 × 540 · 24 FPS · 15s

Long-Horizon Meal Preparation

  • Long-horizon task
  • Ingredient state change
  • Pan and tool control

Slice ingredients for the next stageTransfer aromatics into the heated panAdd and combine the vegetables

First-person view of hands cleaning, drying, and stacking household containers.
Visible phaseRotate and scrub the container surfaces960 × 540 · 24 FPS · 15s

Contact-Rich Cleaning and Container Handling

  • Contact-rich cleaning
  • Object rotation
  • Nesting and stacking

Rotate and scrub the container surfacesDry and inspect the cleaned containerAlign and nest containers for storage

First-person view of two hands folding a green shirt.
Visible phaseLift and orient the garment960 × 540 · 24 FPS · 10s

Deformable-Object Manipulation

  • Deformable objects
  • Bimanual coordination
  • Changing geometry

Lift and orient the garmentAlign fabric edges with both handsFold and compress the garment

What the Public Previews Prove

The six public previews are excerpts, not the complete dataset. Qualified organizations can purchase licensed access to existing EGXO datasets; evaluation, model-training, commercial-use, retention, redistribution, and other rights are defined separately in writing.

Assets
6 video excerpts + posters
Duration
10-15 seconds each
Frame size
960 × 540
Frame rate
24 FPS
Tasks
Mechanical maintenance, timed cooking, fine tool use, long-horizon cooking, cleaning and stacking, laundry folding
Audio
Removed
Metadata
Identifying source fields removed
Approval
Consented release privacy QA + site-owner publication approval
Task familyVisible learning signalRepresentative challenge
Fan maintenanceDisassembly, component cleaning, and ordered reassemblyPart state, fastener visibility, and alignment
Timed pancake cookingMixture state, controlled pouring, and tool-mediated turningHeat timing and continuous visual state change
Vegetable preparationPeeling, sectioning, precision slicing, and dual-tool controlSafety, contact visibility, and fast motion
Long-horizon stir-fryIngredient preparation, transfer to heat, and pan combinationLong-range phase dependencies and changing state
Container cleaning and stackingSurface scrubbing, inspection, alignment, and nestingContact, reflections, and spatial fit
Laundry foldingBimanual coordination and deformable-object stateContinuously changing geometry

The machine-readable manifest contains sanitized descriptive metadata for the six previews. Training and redistribution rights require a separate written license.

Download Preview Manifest (.json)

Dataset-card excerpt

Make the Release Understandable Without Reverse Engineering

A buyer should receive a machine-readable manifest and a human-readable dataset card describing the release, intended use, task coverage, modalities, transformations, limitations, rights, version, and validation evidence.

{
  "release": "project-name-v1.0",
  "viewpoint": "head-mounted-egocentric",
  "modalities": ["rgb_video", "task_metadata"],
  "episode_unit": "one task attempt",
  "acceptance": {
    "critical_actions_visible": true,
    "privacy_review": "passed",
    "buyer_ingest": "passed"
  },
  "known_limitations": [
    "no force or torque measurements",
    "human actions require embodiment mapping"
  ]
}

Privacy, Rights, and Provider Due Diligence

First-person cameras can expose bystanders, screens, documents, voices, reflections, addresses, private spaces, and behavior unrelated to the requested task.

01

Collection authority

Which notice, consent, contractual, or other lawful framework covers the contributor, incidental people, environment, and intended use?

02

Rights scope

Does the release permit training, evaluation, annotation access, derived models, retention, vendor sharing, and the buyer’s intended deployment?

03

Privacy controls

How are capture minimization, exclusion zones, stop rules, review, redaction, access tiers, incidents, and deletion handled?

04

Traceability

Can each delivered asset be connected to protocol, capture, QA, annotation, privacy, rights, transformation, and release versions?

05

Acceptance evidence

What was measured, on which population, under which thresholds, and what proportion was rejected, quarantined, or left unverified?

06

Delivery proof

Has a representative sample passed the buyer’s loader, modality checks, episode traversal, batch inspection, and release documentation review?

From requirement to pilot

Bring the Model Gap. Leave With a Testable Data Brief.

Share the task, deployment environment, required viewpoint, modalities, annotations, rights, format, and success criteria. The next step is a focused pilot, not a fictional volume promise.

Common Egocentric Data Questions

These answers define the category. Project-specific commitments belong in the specification, pilot results, and written proposal.

What is egocentric data?

Egocentric data is recorded from the perspective of the person or machine performing a task. It usually includes first-person video and may add audio, depth, IMU, gaze, pose, language, task structure, and synchronized metadata.

Is egocentric data the same as first-person video?

First-person video is the most common egocentric modality, but a training dataset can include additional sensors, annotations, provenance, rights records, and delivery metadata. A video folder alone is not a complete egocentric dataset.

What is egocentric data used for in robotics?

Teams use it to study hand-object interaction, procedural task structure, state changes, affordances, long-horizon activity, language grounding, representation learning, world models, and evaluation. The useful target depends on what the capture actually makes observable.

Can human egocentric video train a robot directly?

Sometimes it supplies useful visual or task supervision, but ordinary human video does not contain robot joint commands, force, torque, or a compatible action space. Many systems still require embodiment mapping, teleoperation, robot rollouts, simulation, or robot-native demonstrations.

When is exocentric data better?

An external viewpoint is often better when whole-body pose, workspace geometry, other agents, stable world coordinates, approach trajectories, or events outside the actor’s field of view are central to the model objective.

When should ego and exo cameras be synchronized?

Use paired capture when the learning or evaluation task needs cross-view correspondence, recovery of occluded events, body and workspace context alongside actor-aligned detail, or transfer between human and robot perspectives. Define clock, drift, calibration, and dropped-frame tolerances before scaling.

Which camera position is best for egocentric data?

There is no universal best mount. Head and glasses views follow attention and viewing direction; chest views can be more stable; wrist views can reveal close manipulation but lose scene context. Run a visibility pilot on the real tasks, people, and workspaces.

What annotations can an egocentric dataset contain?

Common layers include task and step boundaries, actions, objects, hand-object interactions, narration, transcripts, gaze, pose, scene attributes, quality results, privacy decisions, protocol versions, and provenance. Annotation density should follow the training or evaluation objective.

Should a team use a public or custom egocentric dataset?

Use public data for shared benchmarks, research comparison, and fast baselines. Use custom collection when deployment tasks, environments, sensors, labels, licenses, privacy controls, or failure cases do not match. A hybrid strategy often provides the best balance.

How should egocentric data quality be measured?

Measure model-relevant visibility, media integrity, task completion, coverage, timestamp alignment, annotation fitness, privacy eligibility, rights status, schema validity, release lineage, and successful ingestion by the buyer’s actual pipeline.

Which delivery formats are available?

Native media with structured manifests, RLDS, LeRobot, and custom exports can be evaluated against the collection and buyer pipeline. The final format should be confirmed through a representative ingest test rather than assumed from a format name.

What should a buyer request before commissioning collection?

Request representative samples, a task and coverage specification, capture protocol, acceptance criteria, schema, dataset-card outline, privacy and rights summary, versioning policy, delivery plan, and a pilot that exercises the intended loader.

Can a company buy off-the-shelf egocentric video from EGXO?

Yes. Qualified organizations can request pricing and licensed access to EGXO's first-party household or commercial and workplace egocentric video. Current task coverage, available volume, permitted use, license terms, security requirements, and delivery are confirmed against the submitted use case.

Can EGXO run a custom egocentric data collection project?

Yes. EGXO can scope buyer-defined collection around required tasks, environments, viewpoints, contributor criteria, modalities, annotations, rights, acceptance gates, scale, and delivery format. Availability, timing, pricing, and final commitments are confirmed in writing after the project brief is reviewed.

Primary research context

Public Datasets and Formats That Define the Category

Choose a buying path

Start With Available Data or Build What Is Missing.

Send the short brief with your model objective, required tasks, rights, delivery needs, and timing.