Delivery format
LeRobot format for robot training data
LeRobot format is useful for developer-friendly robot learning datasets and policy training pipelines. Define episode metadata, observation tensors, action tensors, timestamps, and repo-compatible manifest before reviewing samples so you can verify that delivery matches the training pipeline.
Quick facts
- Origin
- Hugging Face — LeRobot framework released 2024, Apache-2.0 license, github.com/huggingface/lerobot.
- Datasets
- 181 datasets in the public LeRobot collection on Hugging Face Hub.
- Supported policies
- 10 architectures: ACT, Diffusion, VQ-BeT, HIL-SERL, TDMPC, π0, π0.5, GR00T N1.5, SmolVLA, XVLA.
- Format
- Synchronized MP4 videos + Parquet tables for state/action streams; LeRobotDataset v2.0 / v2.1.
- Simulators
- LIBERO and MetaWorld benchmarks supported.
Comparison
| Format choice | Strength | Risk |
|---|---|---|
| LeRobot format | developer-friendly robot learning datasets and policy training pipelines | Needs exact schema agreement before capture |
| Raw files | Fast supplier export | High buyer cleanup burden |
| Custom schema | Matches internal pipeline | Harder supplier onboarding |
What is LeRobot format?
LeRobot format should be requested when the buyer's training or evaluation pipeline already expects developer-friendly robot learning datasets and policy training pipelines. Anchor the bounty to the canonical specification before suppliers submit samples [1], then use implementation documentation to make the expected file layout reviewable [2]. Robotics teams should also name the dataset or paper lineage they expect suppliers to support [3].
[1]"LeRobot aims to provide models, datasets, and tools for real-world robotics in PyTorch."
For truelabel buyers, that quote matters because it turns LeRobot format from a generic delivery preference into a source-backed requirement the supplier can test against a sample file.
Using LeRobot format with robot data
A useful LeRobot format sample should prove episode metadata, observation tensors, action tensors, timestamps, and repo-compatible manifest, plus file naming, manifest completeness, timestamp behavior, and rejected-example traceability. Include at least one workflow or converter reference so the supplier can show how the files load in practice [4], one interoperability reference for adjacent formats [5], and one comparison source for why this format is preferable to a raw folder dump [6].
Minimum LeRobot delivery schema
LeRobot-compatible delivery should include episode IDs, task or language instruction, observation streams, video/image paths, action tensors, robot state where available, timestamps, episode boundaries, success/failure labels, calibration or embodiment metadata, split metadata, stats or manifest, and license/provenance records. In procurement terms, LeRobot is a packaging target; it is not proof that the data has commercial rights or enough signals for policy learning.
| Field | Buyer question | Reject if |
|---|---|---|
| Observations | Do videos/images load and align? | Missing paths or mismatched frame counts |
| Actions/state | Are shape, units, and cadence documented? | Video-only data presented as robot episodes |
| Metadata | Are embodiment, task, split, and stats present? | No manifest or dataset card |
| Rights | Are license and provenance files attached? | Unknown upstream terms |
LeRobot file layout and manifest
A buyer-ready LeRobot handoff should name the repository or folder root, episode table, video shards or image paths, metadata/stats files, split file, dataset card, license/provenance bundle, and any conversion script. The manifest should include episode_id, task, observation keys, action shape, state shape, FPS, timestamp source, success label, split, and rejection reason for excluded samples. When public Hub data does not fit the embodiment or rights posture, compare the custom option against the LeRobot dataset alternative page before sourcing.
| Field | Why it matters | Acceptance check |
|---|---|---|
| episode_id | joins videos, state, actions, labels | unique and stable |
| observations | names cameras/sensors consumed by the model | paths load and frame counts match |
| actions/state | turns footage into policy data | shape, units, cadence documented |
| split | separates train/eval/leakage review | held-out episodes are not reused |
| license/provenance | keeps commercial review attached | files present per source |
Validation workflow before accepting delivery
Do not accept a LeRobot-labeled folder until it loads in the buyer's environment. Sample validation should open the dataset, count episodes, verify video/action alignment, inspect timestamp monotonicity, confirm metadata and dataset card fields, check success/failure labels, and verify that license/provenance files match the manifest. If conversion from RLDS, HDF5, MCAP, or ROS bag was used, require the conversion log and any dropped-field report. For custom capture, send the same manifest to the robotics data marketplace; for policy data, align the format fields with VLA training data, teleoperation, or robot demonstration requirements before approval.
LeRobot vs RLDS vs HDF5 vs MCAP
Choose LeRobot for Hugging Face/PyTorch robot-learning workflows, RLDS for episodic TFDS-style pipelines, HDF5 for compact hierarchical arrays, MCAP or ROS bag for timestamped robot-native logs, and raw folders only when higher cleanup burden is acceptable. Conversion cannot recreate missing actions, timestamps, calibration, or consent artifacts.
| Format | Best for | Watch out for |
|---|---|---|
| LeRobot | PyTorch/Hugging Face policy training | public datasets may not match rights or embodiment |
| RLDS | episode/step sequential-decision pipelines | conversion details must preserve actions and metadata |
| HDF5 | compact lab trajectory arrays | needs explicit schema documentation |
| MCAP/ROS bag | robot-native timestamped multimodal logs | training stack may need conversion |
Related pages
Use these to move from category-level context into specific task, dataset, format, and comparison detail.
External references and source context
- LeRobot repository
LeRobot provides robotics models, datasets, and tools for PyTorch workflows.
GitHub ↩ - LeRobot documentation
Hugging Face publishes LeRobot documentation for robotics dataset workflows.
Hugging Face ↩ - LeRobot dataset documentation
LeRobot dataset documentation defines dataset packaging expectations.
Hugging Face ↩ - LeRobot: State-of-the-art Machine Learning for Real-World Robotics in Pytorch
The LeRobot paper frames the library as real-world robotics tooling.
arXiv ↩ - RLDS: Reinforcement Learning Datasets
RLDS is a related robotics episode format for conversion planning.
GitHub ↩ - HDF5 1.14 documentation
HDF5 is relevant to LeRobot-compatible robot episode storage.
The HDF Group ↩ - truelabel egocentric data licensing hub
Internal contextual link to egocentric data licensing and provenance guidance.
truelabel.ai - truelabel egocentric data glossary
Internal contextual link to the egocentric data definition.
truelabel.ai - truelabel sourcing brief intake
Internal contextual link to Truelabel's sourcing intake workflow.
truelabel.ai - truelabel egocentric warehouse video sourcing spec
Internal contextual link to warehouse egocentric video sourcing.
truelabel.ai - truelabel egocentric kitchen video sourcing spec
Internal contextual link to kitchen egocentric video sourcing.
truelabel.ai - truelabel industrial egocentric video sourcing spec
Internal contextual link to industrial egocentric video sourcing.
truelabel.ai - truelabel warehouse robotics data sourcing
Internal contextual link to warehouse robotics data sourcing.
truelabel.ai - truelabel kitchen manipulation data sourcing
Internal contextual link to kitchen manipulation data sourcing.
truelabel.ai - truelabel LeRobot format guide
Internal contextual link to the LeRobot format guide.
truelabel.ai - truelabel eval data for robotics hub
Internal contextual link to robotics eval data sourcing.
truelabel.ai - truelabel hand-object interaction data page
Internal contextual link to hand-object interaction training data requirements.
truelabel.ai - truelabel egocentric video datasets hub
Internal contextual link to the egocentric video datasets hub.
truelabel.ai
FAQ
What is LeRobot format used for?
LeRobot format is used for developer-friendly robot learning datasets and policy training pipelines.
What fields should LeRobot format delivery require?
At minimum, require episode metadata, observation tensors, action tensors, timestamps, and repo-compatible manifest, plus a delivery manifest and validation notes.
Can suppliers convert into this format?
Some suppliers can deliver directly in the requested format; others may need conversion. Buyers should require a small sample before full delivery.
Should the format be decided before capture?
Yes. Deciding the format before capture prevents missing fields, timestamp drift, and expensive post-delivery cleanup.
Is LeRobot a dataset or a format?
LeRobot is both an open-source robotics ecosystem and a dataset format. Public Hub datasets may use LeRobot-compatible layouts, and a custom supplier delivery can also be packaged as LeRobotDataset.
What should a LeRobot delivery include?
Require loadable episodes, videos or image paths, action/state arrays, timestamps, metadata or stats, dataset card, split fields, license/provenance files, and a validation note showing the sample opens in your stack.
Can video-only data be converted to LeRobot?
It can be packaged with metadata, but it is not policy-ready LeRobot robot data unless action/state fields, timestamps, task labels, and rights artifacts are present.
Working with LeRobot format
Truelabel normalizes LeRobot format across capture partners so you can ingest one consistent schema instead of writing per-vendor adapters.
Request LeRobot format data