Evaluate Evidence, Not Provider Claims

Source-backed context[1] Hugging Face LeRobotDataset v3 documentation[2] MLCommons Croissant 1.1 metadata specification[3] Project Aria multimodal data formats and calibration[4] W3C PROV-O provenance ontology

An egocentric data provider should be evaluated across six gates: deployment fit, critical-action observability, time and geometry integrity, provenance and permitted use, measured quality, and delivery into the target loader. A provider passes only when representative episodes and production records support each claim.

Start with a small paid pilot built around the real model objective. Inspect accepted, rejected, repaired, and ambiguous examples; run the delivery through the intended preprocessing and training stack; then scale only after the agreed acceptance tests pass. Hours collected are an input. Usable, rights-eligible, ingestible episodes are the product.

1. Define the Training or Evaluation Contract

Source-backed context[1] Hugging Face LeRobotDataset v3 documentation[2] MLCommons Croissant 1.1 metadata specification

Write one representative record before comparing suppliers. Specify what the model observes, what it predicts, the episode boundary, the target tasks and environments, required modalities, output format, and the result the pilot must support. This prevents a provider from substituting easy-to-collect video for the signals the model actually needs.

  • Target model, policy, perception task, or evaluation claim
  • Required tasks, objects, tools, environments, actors, and failure cases
  • Observation fields, prediction or action target, and outcome labels
  • Required train, validation, test, and held-out generalization slices
  • Target delivery format, loader, schema version, and acceptance command

2. Test Viewpoint and Critical-Action Visibility

Source-backed context[3] Project Aria multimodal data formats and calibration[5] Ego-Exo4D capture and dataset documentation

Egocentric does not automatically mean useful. Head, chest, wrist, glasses, and robot-mounted cameras produce different occlusion, motion, field-of-view, and scale. Ask the provider to show that the hands, manipulated object, contact transition, terminal state, and relevant surrounding context remain visible for the target task.

Measure visibility rather than reviewing only attractive clips. A pilot should report critical-action visibility, hand-object occlusion, framing drift, motion blur, lighting failures, off-screen events, and the percentage of takes that pass every required gate.

3. Verify Time, Calibration, and Metadata Integrity

Source-backed context[2] MLCommons Croissant 1.1 metadata specification[3] Project Aria multimodal data formats and calibration

For multimodal capture, ask which clock is authoritative and how camera, IMU, depth, audio, gaze, pose, and external streams are aligned. Review native timestamps, clock-domain conversion, drift, dropped samples, resampling, calibration version, coordinate frames, units, uncertainty, and missing-data masks.

A metadata dictionary should identify whether every field is measured, derived, annotated, or synthetic. The provider should be able to reconstruct a complete episode without relying on undocumented filename conventions or manual repair.

  • Timestamp unit, clock domain, offset, drift, and synchronization method
  • Camera intrinsics, distortion, extrinsics, axis conventions, and calibration date
  • Sensor sample rates, gaps, interpolation policy, and valid masks
  • Episode, task, actor, environment, device, and release identifiers
  • Field-level origin, version, units, uncertainty, and transformation history

4. Separate Provenance, Consent, and Permitted Use

Source-backed context[4] W3C PROV-O provenance ontology[6] NIST AI Risk Management Framework Playbook[7] Ego4D license agreement and access documentation

A technically strong clip may still be unusable. Require a record of who or what produced the data, the collection context, consent status, bystander handling, privacy review, transformations, reviewer actions, release decision, and the terms that authorize the intended model use.

Do not accept a repository’s code license as evidence of data rights. Record model-training permission, commercial use, redistribution, subcontractor access, retention, deletion, derivative outputs, geographic constraints, and revocation procedures separately. Pin the reviewed terms to the dataset version.

5. Demand Measured QA and a Rejection Trail

Source-backed context[1] Hugging Face LeRobotDataset v3 documentation[2] MLCommons Croissant 1.1 metadata specification[7] Ego4D license agreement and access documentation

A provider should define acceptance tests before collection and report results at episode and release level. Quality cannot be reduced to resolution or annotation accuracy. It includes task correctness, observability, synchronization, calibration, metadata completeness, privacy, rights eligibility, and successful ingest.

Ask for rejected and repaired examples. A supplier that shows only accepted media hides the actual operating system. The rejection taxonomy, reviewer agreement, repair history, and recurrence rate reveal whether quality improves as the program scales.

  • Usable-episode yield and rejection reasons by task, device, and environment
  • Critical-action visibility and occlusion frequency
  • Timestamp, calibration, missing-stream, and schema-validation failures
  • Annotation agreement, ambiguity, correction, and reviewer escalation
  • Privacy, consent, rights, and release-approval failures
  • Loader validation, checksums, manifest completeness, and replay success

6. Run a Real Loader Acceptance Test

Source-backed context[1] Hugging Face LeRobotDataset v3 documentation[2] MLCommons Croissant 1.1 metadata specification

The strongest procurement test is boring: receive a representative versioned release and run it through the actual loader. Confirm checksums, manifests, schema validation, media decode, episode reconstruction, modality alignment, splits, masks, and batch sampling. Time the manual work required to make the release usable.

Delivery labels such as LeRobot, RLDS, HDF5, or custom JSON are not sufficient on their own. Verify the exact schema and version, because two releases can share a format name while disagreeing on actions, observation keys, units, timestamps, and episode boundaries.

Score the Pilot Before Scaling

Implementation guidanceEGXO guidance for translating the research into a project specification.

Use a weighted scorecard and a small number of hard gates. A high average must not compensate for missing commercial rights, failed ingest, invisible critical actions, or undocumented synchronization. Record evidence links, owner, result, limitation, and remediation for every criterion.

  • 20 points: deployment, task, environment, and diversity fit
  • 15 points: viewpoint and critical-action observability
  • 15 points: synchronization, calibration, and metadata integrity
  • 20 points: provenance, consent, privacy, and permitted use
  • 15 points: QA evidence, rejection trail, and reviewer controls
  • 15 points: versioned delivery, loader validation, and support
  • Hard fail: unresolved rights, material privacy gap, failed representative ingest, or hidden critical action

Provider Red Flags

Implementation guidanceEGXO guidance for translating the research into a project specification.

Walk away when the sales claim cannot be traced to a record or reproducible test. The most dangerous providers are not necessarily low quality; they are providers whose evidence is too vague to let a buyer discover where the limits are.

  • Hours claimed without usable yield, task distribution, or acceptance criteria
  • A modality checklist without timestamps, calibration, coordinate frames, or missing-data policy
  • Only polished sample clips; no rejected, repaired, or ambiguous records
  • Commercial-use assurances without pinned terms and provenance
  • Annotation accuracy without taxonomy, agreement method, or error slices
  • Format promises without a representative release tested in the buyer’s loader
  • No versioning, change log, checksums, or dataset-level release approval

Primary Sources and Further Reading

  1. [1] Hugging Face LeRobotDataset v3 documentation ↗
  2. [2] MLCommons Croissant 1.1 metadata specification ↗
  3. [3] Project Aria multimodal data formats and calibration ↗
  4. [4] W3C PROV-O provenance ontology ↗
  5. [5] Ego-Exo4D capture and dataset documentation ↗
  6. [6] NIST AI Risk Management Framework Playbook ↗
  7. [7] Ego4D license agreement and access documentation ↗

These sources inform the category-level guidance above. Project-specific requirements are defined with the buyer.