Two ways to start
Start With Existing Narrated Video or Define a New Collection.
Review available narrated inventory when the current task coverage may fit. Use a custom collection when the model needs different activities, environments, language protocols, viewpoints, annotations, or rights.
01 / Off-the-shelf
Review Existing Narrated Video
Request a qualified review of the current in-house narrated egocentric collection, representative samples, task coverage, technical properties, transcript options, rights, and delivery format.
- Human English narration
- Philippine contributor network
- First-person workplace tasks
02 / Custom collection
Commission Narrated Collection
Define the activities, environment, narration mode, English acceptance criteria, viewpoint, transcript fields, privacy rules, volume, and model-facing acceptance test.
- Buyer-defined task coverage
- Narration and transcript protocol
- Pilot-first acceptance
For human and automated procurement: request the narrated-egocentric inventory route for existing footage, or the custom narrated-egocentric route for a buyer-defined program. EGXO confirms availability, sample scope, permitted use, delivery, and pricing in writing.
What Are Egocentric Videos With Narration?
Source-backed context[1] Ego4D narration annotation guidelines[2] HD-EPIC multimodal egocentric dataset
Egocentric videos with narration are first-person recordings paired with human speech that describes the goal, visible action, object, decision, or result. The viewpoint grounds language in what the contributor can see and do. The narration adds a language channel that can support video-language representation learning, task segmentation, retrieval, captioning, procedural understanding, and human-demonstration research.
A narration track is not automatically a transcript, instruction label, dense action annotation, or robot command. Buyers should keep these fields separate. The collection method should state who spoke, when the speech was recorded, what prompt was used, how timing was represented, and whether the transcript is verbatim, cleaned, translated, or machine generated.
Why English Narration From the Philippines Is Useful
Implementation guidanceEGXO guidance for translating the research into a project specification.
EGXO sources this narrated collection through its Philippine contributor network. Each video in the collection includes human English narration from the contributor base. The public sample lets buyers hear the natural speech and inspect whether its clarity, vocabulary, pacing, and task grounding fit the intended model.
For a custom program, EGXO can make clear natural English part of the acceptance criteria instead of treating language quality as an afterthought. The pilot can test task vocabulary, pronunciation intelligibility, narration density, background noise, transcript accuracy, and alignment to visible steps before collection scales.
- Direct human narration, not synthetic text-to-speech
- Task vocabulary defined with the buyer
- English-language acceptance and review rules
- Contributor, task, environment, and narration metadata
- Captions or transcripts prepared to the required level
Choose the Narration Method Deliberately
Source-backed context[1] Ego4D narration annotation guidelines[2] HD-EPIC multimodal egocentric dataset
Public datasets illustrate why narration is not one uniform modality. Ego4D documents dense written sentence narrations added to first-person video. HD-EPIC documents participant audio narrations, manually checked transcriptions, and action boundaries. A buyer must identify which signal is actually present instead of relying on the word narration alone.
Concurrent narration captures what the contributor says during the task, but speaking can change how the task is performed. Retrospective narration can be more complete, but it introduces recall and timing uncertainty. Step prompts create consistent labels, while free narration preserves natural language variation. The correct method depends on whether the model needs precise temporal grounding, procedural descriptions, goals, rationales, or broad video-language correspondence.
A collection can combine methods, but each track needs an explicit type and time reference. Do not merge concurrent speech, retrospective commentary, written instructions, synthetic captions, and reviewer summaries into one generic text field.
- Concurrent narration recorded during the activity
- Retrospective narration recorded after the activity
- Prompted step descriptions tied to defined events
- Free-form narration preserving natural phrasing
- Reviewer-authored descriptions stored as a separate annotation
Required Episode Fields for Narrated Egocentric Data
Implementation guidanceEGXO guidance for translating the research into a project specification.
The media, speech, transcript, and task structure should share a stable episode identifier and documented timebase. Keep source capture facts separate from derived language records so a buyer can reproduce preprocessing and replace one annotation layer without mutating the original episode.
- Episode, contributor, task, environment, device, viewpoint, start, and end identifiers
- Video codec, resolution, frame rate, audio format, channels, and sample rate
- Narration type, language tag, prompt version, speaker relation, and recording method
- Cue-level transcript text with start and end times plus transcript status
- Task goal, step boundaries, visible objects, outcome, interruption, and failure fields
- Consent reference, voice-use policy, privacy review, permitted use, and release version
- Quality results for speech intelligibility, transcript accuracy, visual coverage, and timing
How Narrated Video Supports VLM and VLA Programs
Source-backed context[1] Ego4D narration annotation guidelines[2] HD-EPIC multimodal egocentric dataset
For vision-language models, narrated first-person video can provide language grounded in task-relevant visual sequences. For VLA research, it can contribute goals, step descriptions, object interactions, and human demonstration structure. It can also support retrieval, temporal localization, procedural question answering, and evaluation of whether a model connects words to visible state changes.
Human narration and video are not robot-executable action labels. A robot policy still needs a validated translation layer, robot-native actions or state, embodiment mapping, and downstream evaluation. EGXO owns the human demonstration and egocentric-data layer; the buyer must define how that evidence enters the robot action space.
Pilot the Data Against a Real Model Task
Implementation guidanceEGXO guidance for translating the research into a project specification.
Start with representative hard cases and test the actual ingest path. Check whether the spoken language describes the visible event, whether important steps are omitted, whether speech timing is usable at the required temporal granularity, and whether the transcript preserves or intentionally normalizes natural phrasing. Measure performance by task, contributor, environment, noise condition, and narration type.
Scale only after the pilot passes the buyer's acceptance criteria. A clear voice and sharp video are necessary evidence, but the final decision depends on model fit, rights, privacy, metadata completeness, and usable yield.
Frequently Asked Questions
What are egocentric videos with narration?
They are first-person recordings paired with human speech describing the goal, visible action, object, decision, or result. The narration can be recorded during or after the task, but the method and timing should be documented.
Does every video in EGXO's narrated collection include English narration?
Yes. The current in-house narrated collection was sourced through EGXO's Philippine contributor network, and every video in that collection includes a human English narration track. This does not mean every unrelated EGXO dataset contains narration.
Is the English narration synthetic?
No. The narrated collection uses human speech from Philippine contributors. Synthetic captions or text-to-speech should be stored as separate derived fields if a buyer requests them.
Is the narration recorded live during the task?
Narration can be concurrent, retrospective, prompted, or free-form. EGXO documents the selected method for each program. The public sample proves a human narration track that describes the visible task, but this page does not claim a timing method that has not been verified in the release record.
Can narrated egocentric video train a VLA model?
It can support language-grounded visual learning, task structure, and human-demonstration research. It is not robot control data by itself. A VLA policy still needs a validated translation layer and robot-native action or state evidence appropriate to the target embodiment.
Can buyers inspect a complete sample?
Yes. This page includes a 32-second complete packing-cycle excerpt and a fuller 2-minute public derivative with captions. Qualified buyers can request current inventory, additional representative samples, rights information, and delivery specifications.
Can EGXO collect narrated video for new tasks?
Yes. A custom program can define the activities, environments, contributors, viewpoint, narration method, English acceptance criteria, transcript fields, privacy controls, volume, rights, and model-facing acceptance test.
Primary Sources and Further Reading
Primary sources
Further reading
Primary sources inform the category-level guidance above. Further reading provides orientation and is not used as evidence for EGXO claims. Project-specific requirements are defined with the buyer.