Dataset alternative
Open X-Embodiment alternative
Open X-Embodiment is a strong public cross-embodiment research baseline: it pools data from 21 institutions across 22 embodiments, but the cited paper and project do not document one commercial license or a unified consent grant for the pool — so commercial suitability is not verified from those sources, and each upstream dataset must be reviewed before commercial use. Use it for benchmarking and representation learning; commission custom, rights-cleared collection when you need a specific embodiment, a private environment, a consent package, a controlled task distribution, or commercial deployment terms the public mixture cannot guarantee.
Verdict by buyer scenario
How we selected and evaluated the options
How to choose an Open X-Embodiment alternative. This is a bottleneck decision, not a ranking — the right alternative is defined by which gap OXE leaves for your deployment.
| Criterion | Why it matters | Evidence we require |
|---|---|---|
| Cross-embodiment coverage | OXE's strength is breadth; a single-arm program trades breadth for fit. | The official corpus/embodiment description. |
| Task / environment match | A generalist mixture rarely matches your workcell. | The task distribution documented for the baseline. |
| Data format | RLDS/LeRobot/MCAP fit avoids a hidden ETL project. | The stated record schema. |
| Licensing / consent status | The cited OXE paper and project do not document one commercial license or a unified consent grant, so commercial suitability is not verified from them. | Each upstream dataset's own license and consent — recorded as found, not found, or source silent — before commercial use. |
| Consent visibility | The cited OXE sources do not expose per-contributor consent. | A stated consent/provenance posture; otherwise 'not verified'. |
| Commercial-deployment risk | Research posture ≠ product-training clearance. | Explicit commercial-use terms, dated. |
| Custom-collection ability | Custom capture can match a private embodiment + rights. | A supplier's ability to deliver a rights-cleared sample first. |
Choose by bottleneck; no single winner and no weighting.
- Inclusion rules
- Included: Open X-Embodiment, closely comparable public baselines (DROID, BridgeData V2), and custom rights-cleared collection as the alternative when public data cannot clear fit or rights.
- Exclusion rules
- Excluded: generic annotation vendors with no capture evidence, public datasets presented as uniformly commercial, and tools mislabeled as data sources. Broad cross-dataset commercial-use shortlists route to their own owner.
- Source basis
- Primary papers for OXE, DROID, and BridgeData V2, each with a checked date, plus TrueLabel's own marketplace documentation for the custom-collection row. No pricing, turnaround, or capability numbers are invented.
- Update cadence
- Robotics datasets move fast; re-checked monthly, with row-level checked dates on every license/commercial claim.
- Disclosure
- TrueLabel publishes this comparison and offers custom rights-cleared collection — a commercial interest you should read this knowing. Public baselines are assessed from primary sources, not paid placement, and TrueLabel is compared only as custom rights-cleared collection, never as a public dataset or a reseller of one. This is not legal advice; commercial suitability requires review of each upstream dataset's exact license, consent, and intended use.
- Scoring caveat
- Choose by bottleneck; this is a shortlist, not a ranking. The cited OXE paper and project do not document one commercial license or a unified consent grant, so commercial suitability is not verified from them — review each upstream dataset.
Evidence matrix
| Option | Supported claim | Official source | Checked | Confidence | Limitation |
|---|---|---|---|---|---|
| Public baselines | |||||
| Open X-Embodiment | Pools more than 1,000,000 trajectories across 22 embodiments and 21 institutions, spanning 527 skills (160,266 tasks), to study cross-embodiment policy transfer. | Open X-Embodiment: Robotic Learning Datasets and RT-X Models · Project site | 2026-07-21 · 2026-07-21 | High (paper + project page) | Aggregation of 60 datasets pooled into one release; the cited OXE paper and project do not document a commercial license or a unified consent grant, so commercial suitability is not verified from them — review each upstream dataset. |
| DROID | A single-Franka-Panda real-world baseline: 76,000 demonstrations across 564 scenes and 86 tasks, captured by 50 operators at 13 institutions. | Project site | 2026-07-21 | High (project page) | One embodiment and research scenes; no per-buyer object, workcell, or commercial-consent coverage. The arXiv abstract states 84 tasks; the official DROID project page states 86 — this row cites the project page for the 86 figure. |
| BridgeData V2 | A large real-robot manipulation dataset collected on a WidowX 250 arm. Its code repository (github.com/rail-berkeley/bridge_data_v2) carries an MIT LICENSE. | BridgeData V2: A Dataset for Robot Learning at Scale · BridgeData V2 dataset repository | 2026-07-19 · 2026-07-21 | High (paper + repo LICENSE) | The MIT license grants rights in the repository's Software only, not the captured media or contributor rights; the dataset/media is downloaded separately and its terms must be verified separately. Research tasks/hardware, not universal deployment coverage. |
| Custom collection alternative | |||||
| TrueLabel custom collection | A physical AI data marketplace that sources rights-cleared, embodiment-specific demonstrations with contributor consent and delivery in the buyer-required RLDS, LeRobot, MCAP, or custom schema — not a public dataset and not a reseller of OXE. | truelabel physical AI data marketplace bounty intake | 2026-07-21 | first-party | We publish this page; ask for a rights-cleared sample from your embodiment and environment before scale. |
The cited OXE paper and project do not document one commercial license or a unified consent grant for the pool, so commercial suitability is not verified from those sources — review each upstream dataset before commercial use. Public availability of the mixture settles none of it.
Buyer decision checklist
- Choose when
- Pretraining a generalist policy, scenes close enough, no commercial-rights blocker → OXE / DROID / BridgeData V2. · You need a specific embodiment, private environment, or controlled task distribution → custom collection. · Your legal team requires one harmonized commercial license across the corpus → custom collection under a single buyer-owned license.
- Avoid when
- You need commercial-training rights that the OXE paper and project do not document (commercial suitability is not verified from them, so each upstream dataset must be reviewed first), or your exact embodiment is sparsely represented in the mixture.
- Proof to request
- For every OXE component you train on, that upstream dataset's own license and whatever consent/provenance the upstream documents — or a written finding that it is silent; plus a documented task/embodiment/environment match; the delivery format; and, for custom capture, one rights-cleared accepted sample from your target embodiment before any scale commitment.
Limitations and caveats
Quick facts
- OXE scale
- 1M+ real robot trajectories pooled across 22 distinct embodiments and 21 institutions, 527 skills / 160,266 tasks (October 2023)
- Format
- RLDS — common record schema unifying contributing datasets so downstream models can train across robots.
- Where it fits
- Cross-embodiment pretraining (RT-X, RT-2-X) and policy generalization research.
- Commercial gap
- The cited OXE paper and project do not document one commercial license or a unified consent grant across the 60 contributing datasets; commercial suitability is not verified from those sources, and embodiment, environment, and task coverage may not match the buyer's robot or workcell.
- What to source instead
- Embodiment-specific demonstrations on the buyer's robot, in the buyer's environment, with rights-cleared delivery, contributor-consent artifacts, and RLDS, LeRobot, MCAP, or custom schema.
Comparison
| Criteria | Open X-Embodiment | truelabel sourcing |
|---|---|---|
| Best use | cross-embodiment robot learning research baseline | net-new robot demonstrations for the buyer's embodiment or task |
| Rights | Check public license and restrictions | Buyer-defined commercial terms |
| Fresh capture | Fixed public corpus | Supplier samples against a new spec |
| Metadata | Dataset-defined | Buyer-required manifest and QA fields |
When Open X-Embodiment is the right baseline
Open X-Embodiment is a cross-embodiment robotics corpus: it unifies more than one million trajectories across 22 embodiments and 21 institutions, spanning 527 skills, to support generalist policy training [1]. It can serve as a research-grade benchmark of broad embodiment coverage, a pretraining substrate for cross-robot generalization, and a baseline against which deployment-specific data is measured. The cited paper and project document its scale and research role but do not document one commercial license or a unified consent grant, so commercial suitability is not verified from those sources [1]. If you are pretraining a generalist policy and your scenes are close enough to what the mixture already covers, use OXE as the baseline and spend your data budget where the mixture is thin.
What the OXE sources do—and do not—document
The procurement gap is not OXE's research quality; it is that OXE pools dozens of contributing datasets behind one release, and the OXE paper and project describe that scale without stating one commercial license or a unified consent grant for it [1]. Commercial suitability therefore cannot be read off the aggregation: a team training a paid product against the unified release has to review each upstream dataset's own terms before relying on it, and public availability of the mixture settles none of that. Undocumented composite corpora are exactly where downstream risk hides [2].
[2]"The machine learning community currently has no standardized process for documenting datasets, which can lead to severe consequences in high-stakes domains."
When to replace or complement OXE with custom collection
A custom alternative is warranted when the buyer needs commercial-training rights under one harmonized license, contributor-consent artifacts, deployment-environment fidelity, or fresh demonstrations on the buyer's exact embodiment. Commercial vendors market manipulation-data collection and annotation with delivery terms suited to product deployment [3] commercial collection programs; whether a given vendor also supplies contributor-consent and licensing artifacts is a separate question to verify, not something a collection offer proves. The decision is rarely "OXE or custom" — one possible deployment pattern is to pretrain on a public baseline and fine-tune on rights-cleared, embodiment-specific data collected for the deployment. Treat the public mixture as the floor a paid program has to clear, then source only what the mixture does not cover or document.
How to scope an Open X-Embodiment alternative
Scope the replacement around the exact gaps OXE does not cover for your deployment: commercial license terms, target embodiment, capture rig, accepted tasks, contributor-consent coverage, and sample-level annotation requirements. A strong request specifies dataset motivation, composition, collection process, and recommended uses before any supplier begins capture [4]. Attach a structured Data Card summary to each delivered batch so buyers can audit dataset origin, development, and intent. Buyers can still point suppliers to the Open X-Embodiment paper so everyone understands the cross-embodiment baseline being complemented — but the accepted sample, not the aggregation, is what proves commercial terms and buyer-specific metadata before scale-up.
Public baselines that complement OXE: DROID and BridgeData V2
When breadth matters less than a tight single-embodiment baseline, two public corpora can serve as complements. DROID contributes 76,000 real-world demonstrations across 564 scenes and 86 tasks, captured by 50 operators at 13 institutions on a Franka Panda arm — a useful reference for the deployment-fidelity bar a commercial replacement typically has to meet [5]. BridgeData V2 is a large real-robot manipulation dataset collected on a WidowX 250 arm [6]; its code repository carries an MIT LICENSE [7], which grants rights in the repository's Software — not the captured media or the rights of the people who produced it; the dataset/media is downloaded separately and its terms must be verified separately. Both are genuine baselines, and neither carries your exact objects, workcell, or contributor consent. Use them to benchmark; source custom data when the embodiment, environment, or rights diverge.
Acceptance gates before you scale a custom corpus
Before scaling any custom collection into a deployment corpus, run a structured acceptance protocol on every batch rather than sizing the program off a public dataset. The gates that matter, in order: embodiment match (the correct robot, gripper, and calibration); action-schema match (RLDS-compliant records, time-aligned observations, actions, state, and terminal flags); license harmonization (every episode under one buyer-owned commercial-training license, or mapped to a subset license that has cleared review); per-contributor consent (a signed commercial-training agreement and per-session consent for every operator); sensor fidelity (RGB, depth, and end-effector pose synchronized within a stated tolerance); task-success labeling (human-verified success with a documented reviewer-agreement process); and coverage (enough distinct objects, lighting, background, and operator variation for the task). Reject any batch that misses a gate, and run a small pilot before funding scale — skipping the pilot can make late gate failures expensive, because they surface after collection when re-collection is far costlier than a first-batch eval.
Related pages
Use these to move from category-level context into specific task, dataset, format, and comparison detail.
External references and source context
- Open X-Embodiment: Robotic Learning Datasets and RT-X Models
Open X-Embodiment unifies more than 1 million trajectories across 22 embodiments and 21 institutions, spanning 527 skills; the cited paper and project do not document one commercial license or a unified consent grant.
arXiv ↩ - Datasheets for Datasets
Verbatim Datasheets for Datasets framing for why undocumented composite corpora create downstream procurement risk in commercial deployment.
arXiv ↩ - encord
Encord markets manipulation-data collection and annotation; whether a program includes contributor-consent and licensing artifacts is a separate question to verify.
encord.com ↩ - Data Cards: Purposeful and Transparent Dataset Documentation for Responsible AI
Data Cards capture dataset origins, development, intent, and ethical considerations buyers need before commercial training.
arXiv ↩ - Project site
DROID: 76,000 demonstrations across 564 scenes and 86 tasks, 50 operators, 13 institutions (Franka Panda).
droid-dataset.github.io ↩ - BridgeData V2: A Dataset for Robot Learning at Scale
BridgeData V2 is a large real-robot manipulation dataset collected on a WidowX 250 arm.
Proceedings of Machine Learning Research ↩ - BridgeData V2 dataset repository
The BridgeData V2 GitHub repository carries an MIT LICENSE that grants rights in the repository's Software only; the dataset/media is downloaded separately and its terms must be verified separately.
RAIL, UC Berkeley ↩
FAQ
Can Open X-Embodiment be used commercially?
The cited OXE paper and project describe the pooled corpus and its scale, but they do not document one commercial license or a unified consent grant for it, so commercial suitability is not verified from those sources — review each upstream dataset you train on. Public availability of the mixture does not grant commercial rights to any part of it.
When is Open X-Embodiment enough?
When you are pretraining a generalist cross-embodiment policy, doing RT-X-style reproducibility or ablations, and your scenes are close enough that you do not need commercial-exclusive or embodiment-specific coverage. OXE can serve as a research baseline. The cited paper and project do not document one commercial license or a unified consent grant, so its commercial suitability is not verified from those sources — that is a licensing and fit question, not a scale question.
When do I need custom demonstrations instead?
When your embodiment is sparsely represented in the mixture, your environment is private, you need contributor-consent artifacts, you require one harmonized commercial license, or your task distribution and failure modes are not covered. One possible pattern is to pretrain on a public baseline and fine-tune on rights-cleared, embodiment-specific data collected for the deployment.
How should I document the subsets I trained on?
Keep a per-subset manifest: for every OXE component in your training mixture, record what its own license and consent state — found, not found, or source silent — with the date you checked each. Attach a structured Data Card to any custom batch you add. That manifest is what a legal or procurement review will ask for, and it is far cheaper to build as you go than to reconstruct later.
How does TrueLabel differ from Open X-Embodiment?
Open X-Embodiment is a public research aggregation; TrueLabel is a physical AI data marketplace for custom, rights-cleared collection. TrueLabel does not resell or relicense OXE or any public dataset — it sources embodiment-specific demonstrations with contributor consent and delivery in the buyer-required RLDS, LeRobot, MCAP, or custom schema, so a buyer can clear rights and fit that the public mixture does not establish. Use OXE as the baseline; use custom collection where the cited sources do not document commercial rights or the baseline does not cover your fit.
Still choosing between alternatives?
Send the dimensions that matter most — license, modality, scale, contributor consent — and truelabel routes you to the dataset or partner that actually fits.
Request an Open X-Embodiment alternative