Robot action data
Teleoperation data marketplace
A teleoperation data marketplace lets robotics teams source synchronized camera streams, joint states, end-effector poses, action traces, task labels, and success/failure metadata captured while a human operator controls the robot — and Truelabel benchmarks teleop sourcing against public references like RoboSet's 9,500 teleoperated trajectories. Truelabel matches teleop sourcing requests to candidate capture suppliers reviewed against the buyer spec and routes samples through buyer review before scale. The decision is a scenario, not a winner: custom teleop capture on your embodiment, a public baseline for research, a managed enterprise program, or tooling-plus-capture. Public datasets (DROID, BridgeData V2, RoboSet, AgiBot World) are baselines — they rarely satisfy your exact deployment distribution, consent, or embodiment. TrueLabel is the custom-capture path; it is not the fit when a public baseline already matches your robot or when you want a fully managed enterprise program.
Verdict by buyer scenario
How we selected and evaluated the options
How we evaluate teleoperation sources. Tuned to what makes teleop data usable, not generic labeling. Weights are ours.
| Criterion | Weight | What we check |
|---|---|---|
| Embodiment match | 20% | Same arm/gripper/DoF as your deployment, or a documented transfer gap |
| Control frequency / telemetry | 15% | Action + state logged at a usable, documented rate, with end-effector pose telemetry |
| Camera views | 10% | Wrist, egocentric, and/or external streams, time-synced |
| Force/torque / depth | 15% | Contact-rich signal (F/T, tactile, depth) where the task needs it |
| Operator QA | 15% | Human-verified success/failure labels, inter-reviewer agreement |
| Consent / provenance | 10% | Per-session consent artifacts, chain of custody |
| Pilot turnaround | 10% | Time to a reviewable sample before scale |
| Total cost | 5% | Quote-based against the spec, not a flat rate |
Weights sum to 100%.
- Inclusion rules
- Included if the source serves teleoperation/robot-action data and has a primary or official source we cite. Public datasets are labeled baselines; only paid/service operations are scored as providers.
- Exclusion rules
- Excluded sim-only frameworks positioned as capture supply (noted separately), and unsourced claims.
- Source basis
- Dataset project sites/papers and official vendor pages, each dated below.
- Disclosure
- TrueLabel runs this marketplace and this page — weigh that conflict. We separate public baselines from paid providers so a dataset is never scored as if it were a vendor, and we state where TrueLabel is not the fit. No pay-to-play ordering; public info + buyer-fit criteria. Absence of public evidence is not proof a source lacks a capability.
- Scoring caveat
- Dataset counts and vendor scope drift; scores are directional and dated. Verify in a pilot.
Evidence matrix
| Option | Supported claim | Official source | Checked | Confidence | Limitation |
|---|---|---|---|---|---|
| Paid / service providers | |||||
| Scale AI | Managed physical-AI data programs (data-engine work for robotics customers) | scale.com physical ai | 2026-07-19 | Medium (vendor) | Generalist; enterprise minimums — not for small niche pilots on a tight timeline |
| TrueLabel (marketplace) | Routes teleop sourcing requests to candidate suppliers reviewed against the buyer spec; sample review before scale | truelabel teleoperation data marketplace | 2026-07-19 | Medium (first-party) | We publish this page. Not for a public-baseline fit or a fully managed enterprise program; ask for relevant sample evidence and capacity before scale |
| Appen | Lists paid Physical AI services spanning egocentric data, robot evaluation, trajectory annotation, LiDAR, and sensor fusion | Appen Physical AI Training Data | 2026-06-10 | Medium (vendor) | Broad physical-AI service line; confirm teleoperation rig, action/state schema, delivery format, and capacity in a pilot |
| ClarU (commercial teleoperation dataset provider) | Offers a commercial warehouse robot-arm teleoperation dataset with RGB, depth, force/torque, action, force, and success fields | Teleoperation Warehouse Dataset for Robotics AI | Claru | 2026-05-04 | Medium (vendor self-description) | ClarU is a separate commercial vendor with no implied relationship to TrueLabel. Verify current availability, license, embodiment fit, and field-level samples directly |
| Silicon Valley Robotics Center | Offers custom robot teleoperation data collection scoped by robot, objects, scenes, modalities, success criteria, and delivery format | Custom Robot Teleoperation Data Collection Service | Silicon Valley Robotics Center | 2026-05-04 | Medium (vendor) | Custom collection is spec-dependent; request a loadable sample, written rights terms, and evidence of capacity for the target embodiment |
| Tooling / ecosystem | |||||
| Hugging Face LeRobot | Open robotics framework + dataset hub with teleop benchmarks (PushT, ALOHA, xArm); Parquet + video conventions | LeRobot documentation | 2026-07-19 | High (platform) | Ecosystem/tooling, not a managed capture SLA — not for buyer-owned rights out of the box |
| Public baselines — references, not vendors | |||||
| DROID | 76k teleoperated Franka demonstrations, 564 scenes, 13 institutions, synchronized observations + actions | DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset | 2026-07-19 | High (paper) | Single Franka embodiment; research scenes — see /compare/droid-dataset-alternative |
| BridgeData V2 | 60,096 teleoperated trajectories, 24 environments, WidowX 250 (dataset facts); repository published under the MIT license | BridgeData V2: A Dataset for Robot Learning at Scale · BridgeData V2 dataset repository | 2026-07-19 · 2026-07-14 | High (paper + official repo license) | WidowX tabletop tasks; MIT applies to the repository — still confirm it covers your intended use |
| RoboSet | 30,050 trajectories total, of which 9,500 are teleoperated | Dataset page | 2026-05-05 | High (project) | Kitchen-scale manipulation; research corpus |
| Open X-Embodiment | 1M+ trajectories across 22 embodiments, 21 institutions, 527 skills | Open X-Embodiment: Robotic Learning Datasets and RT-X Models | 2026-07-19 | High (paper) | 60+ per-dataset licenses; heterogeneous embodiment coverage |
| AgiBot World (Beta) | 1,000,000+ trajectories from 100 robots across 2,976.4 hours | AgiBotWorld-Beta | 2026-07-14 | Medium (dataset card) | Verify license + embodiment fit before commercial use |
| Mobile ALOHA | Open-hardware bimanual mobile-manipulation platform + public demonstration data | Mobile ALOHA: Learning Bimanual Mobile Manipulation with Low-Cost Whole-Body Teleoperation | 2026-07-19 | High (project) | Requires replicating the hardware platform |
Buyer decision checklist
- Choose when
- Public baseline: Research, imitation-learning starting point, embodiment close enough, license permits → DROID / BridgeData V2 / RoboSet / AgiBot World. · Marketplace: Your exact embodiment, control frequency, force/torque, and commercial rights matter → post a spec, request a 10–25 episode pilot.
- Avoid when
- Marketplace: A public baseline already matches your robot and you don't need exclusivity; or you want a fully managed enterprise program with no supplier selection.
- Proof to request
- Embodiment + gripper spec match; action/state logged at your control frequency; camera-view sync sample; force/torque or depth where the task needs it; human-verified success labels with reviewer agreement; per-session consent artifacts; delivery in MCAP/HDF5/RLDS/LeRobot you can load.
Limitations and caveats
Quick facts
- Robot state
- Joint positions, velocities, end-effector pose
- Video
- Synced wrist, egocentric, or external camera streams
- Task
- Pick, place, open, close, sort, assemble, recover
- Format
- MCAP, HDF5, RLDS, LeRobot, or buyer-defined schema
- QA
- Sync tolerance, completed task segments, metadata completeness
Comparison
| Data type | Contains | Best for |
|---|---|---|
| Egocentric video | Human POV footage | World-model and perception pretraining |
| Robot demonstrations | Human task examples | Imitation and behavior cloning |
| Teleoperation data | Robot actions and synchronized observations | Policy learning and VLA fine-tuning |
What makes teleop data useful
Useful teleoperation data is more than a video export. It keeps synchronized observations and actions [1], explicit episode and step boundaries [2], and timestamped multimodal logs [3] so buyers can audit whether each accepted sample can train or evaluate policies.
[4]"Overall we have 30,050 trajectories in the dataset, out of which 9,500 are collected through teleoperation."
That public dataset pattern is the minimum bar for a marketplace spec: ask for trajectory counts, camera viewpoints, task and scene coverage, and failure labels before funding scale-up [5].
How truelabel routes teleop sourcing requests
The sourcing request captures robot embodiment, teleoperation interface, sensor package, delivery format, and acceptance criteria. truelabel routes suppliers according to whether their rigs can export policy-ready action data [6], whether the proposed collection fits real-world deployment environments [7], and whether the capture partner can support physical-AI data operations rather than generic annotation [8]. Candidate suppliers should be reviewed against the buyer's capability vector before any scale-up is funded.
Why public teleop datasets are baselines, not procurement
DROID, BridgeData V2, RoboSet, and AgiBot World are genuinely useful — they set the shape of what good teleop data looks like: synchronized observations and actions, explicit episode boundaries, timestamped multimodal logs. But a public teleop dataset is a reference distribution, not a supply contract. It rarely matches your exact arm, gripper, control frequency, or workcell; it carries a research or per-dataset license rather than buyer-owned commercial rights; and it ships no per-contributor consent artifacts scoped to your product. Treat them the way you'd treat a benchmark: measure your gap against them, then decide whether custom capture is needed to close it. That's the boundary between this page and /compare/droid-dataset-alternative, which handles the DROID-specific public-vs-custom decision in depth.
Related pages
Use these to move from category-level context into specific task, dataset, format, and comparison detail.
External references and source context
- Project site
Teleop data should pair synchronized observations and robot actions for policy learning.
droid-dataset.github.io ↩ - RLDS: Reinforcement Learning Datasets
Teleop specs should define episodes, steps, observations, actions, and metadata.
GitHub ↩ - MCAP file format
MCAP stores timestamped multimodal robotics logs for delivery and replay.
mcap.dev ↩ - Dataset page
RoboSet reports 9.5 thousand teleoperated trajectories.
robopen.github.io ↩ - Teleoperation datasets are becoming the highest-intent physical AI content category
Teleop sourcing requests should specify embodiment, interface, cameras, rate, and success bar.
tonyzhaozh.github.io ↩ - Project site
Robot policies benefit from action data paired with observations across tasks and embodiments.
robotics-transformer-x.github.io ↩ - Figure + Brookfield humanoid pretraining dataset partnership
Commercial humanoid teams pursue real-world training data from deployment environments.
figure.ai ↩ - scale.com physical ai
Physical AI vendors route custom robotics data collection and data-engine workflows.
scale.com ↩
FAQ
What is teleoperation data?
Teleoperation data is data recorded while a human remotely controls a robot. It usually includes robot state, actions, camera observations, timestamps, and task metadata that can train or evaluate robot policies.
What formats can teleoperation data use?
Common formats include MCAP, HDF5, RLDS, LeRobot datasets, ROS bag exports, JSON, CSV, and buyer-specific schemas. The sourcing request should define the required format before suppliers submit samples.
How much teleoperation data should I request?
The right volume depends on the task, robot embodiment, success criteria, and model architecture. A small eval request can validate sample quality before the buyer funds a larger capture program.
Can teleop data be exclusive?
Yes. Net-new teleop sourcing requests can specify exclusive rights. Off-the-shelf datasets are typically non-exclusive unless the buyer pays for exclusivity.
Should I use a public teleoperation dataset or commission custom capture?
Start public. If DROID (Franka), BridgeData V2 (WidowX), RoboSet, or AgiBot World match your embodiment and task closely enough and you don't need commercial exclusivity, use them — they're free and well-documented. Commission custom capture when your arm/gripper, control frequency, force-torque needs, environment, or commercial rights diverge from anything public. Measure the gap against the baseline first, then decide.
When is TrueLabel not the right teleop source?
When a public baseline already fits your robot and you don't need exclusive, buyer-owned rights — buy nothing. And when you want a single fully managed enterprise teleop program with no supplier selection, a managed data vendor will feel smoother than a marketplace. TrueLabel fits when embodiment match, control-frequency telemetry, operator QA, and buyer-owned commercial rights are the binding constraints, and you'd rather gate on a 10–25 episode pilot than buy blind.
Looking for teleoperation data marketplace?
Specify modality, task, environment, requested rights posture, and delivery format. Truelabel routes the request to candidate capture partners and helps scope consent/provenance artifacts and commercial licensing requirements for buyer review before delivery.
Request teleoperation data