Glossary
RLDS: Reinforcement Learning Dataset Standard
RLDS (Reinforcement Learning Datasets) is an episode-based data format and toolchain from Google DeepMind for sequential decision-making data in robotics and reinforcement learning. Built on TensorFlow Datasets, it stores each trajectory as an episode: an ordered sequence of steps, where every step carries an observation, action, reward, discount, and is_first/is_last boundary flags. Because observations and actions are nested dictionaries, one schema holds multi-camera, multi-robot, language-conditioned data, which is why Open X-Embodiment, RT-X, and OpenVLA all standardized on it.
Quick facts
- Topic
- Rlds
- Audience
- Procurement leads, ML ops, robotics engineers
- Deliverable
- Buyer-facing reference + procurement guidance
What RLDS Is and Why It Exists
Google DeepMind released RLDS in 2021 as a data spec, storage format, and toolchain for sequential decision-making datasets (RLDS: an Ecosystem to Generate, Share and Use Datasets in Reinforcement Learning[1]). Before it, every robot-learning group shipped its own layout: CSVs, pickled Python objects, custom HDF5 schemas, or ROS bag files. Reusing another lab's data cost more engineering than the dataset was worth.
RLDS fixes that with two ideas. First, a dataset is a stream of episodes, each an ordered sequence of steps, so episode boundaries and resets are first-class instead of inferred from timestamps. Second, the per-step fields are nested dictionaries, so one schema absorbs three RGB cameras, a 7-DOF joint vector, a language instruction, and a task embedding without a format change. The payoff is concrete: the Open X-Embodiment collaboration pooled data from 22 robot embodiments across 21 institutions into a single RLDS corpus (527 skills, 160,266 tasks), and that corpus made RT-X possible, a single policy that beat single-dataset baselines by 50% on novel tasks[2]. Hugging Face LeRobot, TensorFlow Datasets, and the major annotation platforms read RLDS natively.
The Step and Episode Schema
An RLDS dataset is a two-level tree: episodes at the top, an ordered sequence of steps inside each. The step schema is fixed, which is what makes cross-dataset code possible; the flexibility lives inside the nested dicts.
Every step carries six fields. `observation` maps sensor names to tensors, for example `{'wrist_rgb': [224,224,3], 'base_rgb': [480,640,3], 'joint_pos': [7]}`. `action` is nested the same way and holds either continuous vectors or discrete indices. `reward` is a scalar float, `discount` is a float in [0,1] for temporal credit assignment, and `is_first` / `is_last` mark episode boundaries[3]. The non-obvious part for newcomers: because boundaries are explicit booleans rather than implicit gaps, a loader can shuffle and window steps without ever crossing an episode reset.
Episodes wrap their steps in a `tf.data.Dataset` and attach metadata like `episode_id`, `success`, `language_instruction`, and `scene_info`. The DROID dataset adds `collector_id`, `robot_serial`, and `timestamp_utc` for data provenance tracking. On disk, RLDS uses TFRecord serialization with sharded `.tfrecord` files plus `features.json` and `dataset_info.json`, so the TensorFlow Datasets integration can stream, shard, and checksum straight from a cloud bucket. Real datasets span 50 GB (BridgeData V2, 60,000 episodes) to 2 TB (Open X-Embodiment aggregate).
Converting Your Dataset to RLDS
Conversion means mapping your source format onto the canonical step schema. The RLDS repository ships reference converters for the common starting points: ROS bags, HDF5 archives, and pickled trajectories. Two source formats cover most teams. ROS bags are timestamped topic streams, so the converter synchronizes topics by timestamp, groups messages into steps, and nests sensors under `observation`; the rosbag2_storage_mcap plugin preserves nanosecond timestamps on the way in. HDF5 sources, common for teleoperation, use groups for episodes and datasets for timesteps, so the converter walks the hierarchy and flattens arrays into dicts. The Furniture Bench dataset ships a complete HDF5 reference.
The step most teams skip is validation, and it is the one Open X-Embodiment enforced hardest: its curation pass rejected 18% of submissions, almost always for missing `is_first`/`is_last` flags or action dimensions that drift between episodes[2].
- 01
Map episodes and steps
Treat each HDF5 group (or each bag recording) as one episode and each timestep as one step. Synchronize multi-rate topics to a single step clock before grouping.
- 02
Define the features dict
Declare tensor shapes and dtypes for every observation and action key. Keys must be identical across all episodes; drifting action dimensions are the top rejection cause.
- 03
Nest observations and actions
Write sensors under observation (wrist_rgb, joint_pos) and controls under action. Keep robot-native units and defer normalization to training time.
- 04
Set boundary flags
Emit is_first on the first step of each episode and is_last on the final step. Missing flags let a policy train straight through a reset.
- 05
Write and validate
Serialize with the RLDS builder API, then run rlds.validate_dataset to check schema, tensor shapes, NaNs, and boundary flags before you ship.
Loading RLDS into Training Frameworks
On TensorFlow, RLDS drops straight into `tf.data`: load sharded TFRecords, apply frame stacking and normalization, batch, and prefetch. The RT-1 training code runs exactly this pattern. PyTorch needs a bridge. The LeRobot library exposes `RLDSDataset`, a `Dataset` subclass that wraps the TFRecords behind `__getitem__`, so a standard `DataLoader` with multi-worker prefetch works unchanged; the Diffusion Policy example runs RLDS to LeRobot to PyTorch Lightning.
Two loader-side transforms matter more than the storage format. Frame stacking gives the policy velocity and acceleration: RLDS stores single frames, and the loader stacks the last k with `tf.data.Dataset.window` or LeRobot's `delta_timestamps`. RT-2 stacks 6 frames at 3 Hz for a 2-second window. Action chunking trades one-step control for an n-step trajectory: the ACT notebook emits 100-step chunks, executes 10, then re-queries, which cuts compounding error and lifts long-horizon success by 23%[4]. Neither transform is stored in the dataset, so the same RLDS files feed reactive and chunked policies without a rewrite.
Multi-Robot and Cross-Embodiment Training
The nested-dict schema is what lets one policy train across incompatible robots. Open X-Embodiment spans 22 embodiments (7-DOF arms, mobile manipulators, quadrupeds, humanoids), each with its own sensors and action dimensions, keyed apart as `franka_wrist_rgb`, `stretch_head_depth`, `spot_proprioception`. To tell the policy which body it is driving, the RT-X procedure injects a learned `embodiment_id` embedding into the observation dict; the transformer attends to that token, learning body-specific priors while sharing vision and language representations. RT-X reached 50% average success on 6 unseen robots after training on 15 source robots[2].
Two conventions make mixing work. Store actions in robot-native units and normalize at training time: LeRobot's normalization module computes per-dataset mean, std, min, and max, then applies z-score or min-max scaling, so no robot's raw scale dominates the loss. And weight the mixture by dataset, not by size. Open X-Embodiment found uniform per-dataset sampling beat size-proportional sampling by 12% on average, because proportional sampling lets a few huge datasets swamp rare skills that never accumulate gradient[2].
RLDS vs ROS Bags, MCAP, HDF5, and Parquet
No format wins on every axis. ROS bags are the default sensor log but carry no episode boundaries or task labels, so converting to RLDS adds structure at a storage cost. MCAP modernizes the bag with indexed seeks and language-agnostic schemas, which makes it far better for interactive exploration but leaves you writing custom parsers per dataset. HDF5 is the teleoperation workhorse, used by ALOHA, and supports in-place writes, but its schemas are per-dataset and it has no built-in versioning or cloud-streaming story. Apache Parquet is the newest contender, strong on columnar filtering and schema evolution and now experimentally supported in LeRobot, though its row-group model fits tabular data better than nested episodes.
The practical rule: pick RLDS when you train TensorFlow or LeRobot pipelines across many datasets, and reach for MCAP or Parquet when random access or schema churn dominates.
| Format | Episode + step structure | Random access | Schema evolution | Best fit |
|---|---|---|---|---|
| RLDS | Built in (fixed step schema) | Sequential, slow seeks | Regenerate dataset | Multi-dataset TF/LeRobot training |
| ROS bag | None (flat topic stream) | Replay-oriented | Ad hoc | Raw on-robot sensor logging |
| MCAP | None (indexed messages) | Chunk-indexed seeks | Language-agnostic schemas | Interactive exploration, active learning |
| HDF5 | Per-dataset groups | Partial reads | In-place writes | Teleoperation capture |
| Parquet | Row groups (not native) | Columnar projection | Backward-compatible columns | Filtering, columnar analytics |
Limitations and When to Reach for Something Else
RLDS's strengths carry matching costs. TFRecord framing (length prefixes, CRC checksums, padding) adds roughly 15-25% over raw HDF5 or MCAP, so a 100 GB HDF5 dataset lands near 120 GB. Random access is the sharper limit: TFRecords are append-only, so seeking to episode n means scanning everything before it, which is why the MCAP format with chunk indexes wins for active-learning loops that sample specific episodes. Schema evolution is brittle too. Adding one observation key means regenerating the whole dataset, whereas the Parquet format lets old readers ignore new columns; the LeRobot library is testing Parquet for exactly this reason. RLDS softens the pain with semantic versioning, where a minor bump adds backward-compatible fields (the BridgeData V2 release added 60,000 episodes to V1 as a minor version), but a renamed key is still a breaking major bump.
High-rate multimodal streams strain the tensor-centric schema hardest. The HOI4D dataset logs 6-axis force-torque at 1 kHz as variable-length tensors, and RLDS's fixed schema forces every step to pad or truncate to a maximum length. Proposed fixes point the same way: an in-progress RLDS-Video profile stores observations as compressed H.264/VP9 streams instead of per-frame tensors to shrink video-heavy datasets, and world models like NVIDIA Cosmos are pushing the format toward native video. If your data is variable-rate audio, tactile, or force-torque, evaluate MCAP with protobuf schemas before committing.
RLDS in Foundation Models and Commercial Pipelines
RLDS is now the default training format for robot foundation models. RT-1 trained on 130,000 RLDS episodes from one robot and hit 97% success on 17 tasks; RT-2 added web-scale vision-language pretraining on top of the same robot data. The open-source OpenVLA model, trained only on the public, all-RLDS Open X-Embodiment corpus, outperformed the closed 55B-parameter RT-2-X by 16.5% absolute task success across 29 evaluation tasks, and its release ships the RLDS loaders and normalization statistics so anyone can reproduce it. Simulation follows the same convention: the NVIDIA GR00T N1 report mixes simulated and real trajectories under one standardized loader, so sim and real share the same pipeline[5].
That standardization is why data vendors converge on RLDS. Scale AI's Physical AI platform ingests arbitrary formats and exports RLDS, and annotation tools like Encord and Segments.ai now emit it alongside COCO and Pascal VOC. Truelabel's physical AI marketplace delivers robot-manipulation datasets in RLDS, alongside LeRobot, MCAP, and custom schemas, to S3, GCS, or Azure, with contributor consent artifacts and per-trajectory provenance metadata on each delivery. For regulated buyers in medical robotics, food handling, or aerospace, that provenance record of who collected and labeled each trajectory is what makes FDA 510(k) and EU AI Act audits tractable[6].
Related pages
Use these to move from category-level context into specific task, dataset, format, and comparison detail.
External references and source context
- RLDS: an Ecosystem to Generate, Share and Use Datasets in Reinforcement Learning
Formalizes RLDS specification and ecosystem in 2021 Google DeepMind paper
arXiv ↩ - Open X-Embodiment: Robotic Learning Datasets and RT-X Models
Open X-Embodiment collaboration consolidating 60 datasets, 160,266 tasks, 22 embodiments in RLDS format
arXiv ↩ - RLDS GitHub repository
RLDS GitHub repository with conversion scripts and validation tools
GitHub ↩ - LeRobot: State-of-the-art Machine Learning for Real-World Robotics in Pytorch
LeRobot paper reporting 23% higher success with action chunking
arXiv ↩ - NVIDIA GR00T N1 technical report
NVIDIA GR00T N1 technical report mixing simulated and real trajectories under one standardized loader
arXiv ↩ - Regulation (EU) 2024/1689 laying down harmonised rules on artificial intelligence
EU AI Act Regulation 2024/1689 requiring data provenance for high-risk AI systems
EUR-Lex ↩
More glossary terms
FAQ
What is the difference between RLDS and ROS bags?
ROS bags are timestamped message streams with no semantic structure; they store raw sensor data and control commands as they arrive. RLDS adds episode boundaries, task labels, success flags, and a standardized step schema (observation, action, reward, discount). Converting a ROS bag to RLDS requires synchronizing message topics by timestamp, grouping them into steps, and nesting sensor data under the observation dictionary. RLDS datasets are 15-25% larger than equivalent ROS bags due to TFRecord framing overhead, but they integrate directly with TensorFlow and PyTorch training pipelines without custom parsing logic.
Can I use RLDS datasets with PyTorch?
Yes, via bridge libraries. The LeRobot library provides RLDSDataset, a PyTorch Dataset subclass that wraps RLDS TFRecords and exposes them through standard __getitem__ indexing. This enables PyTorch DataLoader usage with multi-worker prefetching and GPU pinning. LeRobot handles frame stacking, action chunking, and normalization automatically. Performance is comparable to native PyTorch formats when training Diffusion Policy on the Open X-Embodiment dataset.
How do I convert my HDF5 robot dataset to RLDS?
The RLDS GitHub repository provides reference conversion scripts. The typical workflow: read HDF5 groups as episodes, read HDF5 datasets as timesteps, map your observation and action arrays to the RLDS step schema (nested dictionaries), add is_first and is_last flags at episode boundaries, and write TFRecords using the RLDS builder API. You must define a features dictionary specifying tensor shapes and dtypes for all observation and action keys. After conversion, run rlds.validate_dataset to check schema conformance and detect NaN values. The Furniture Bench dataset provides a complete HDF5-to-RLDS reference implementation.
What metadata should I include in an RLDS dataset?
Minimum required metadata: dataset name, version, description, citation, splits (train/val/test), and features (tensor shapes/dtypes). Best practice adds robot-specific fields: embodiment (robot model), control frequency (Hz), camera names, action space type (continuous/discrete), task categories, collector demographics, known failure modes, and licensing terms. The Open X-Embodiment project and DROID dataset provide exemplary metadata. For commercial datasets, include data provenance fields: collector IDs, robot serial numbers, calibration timestamps, and annotation lineage to support regulatory audits.
Why is RLDS larger than my original dataset?
RLDS uses TFRecord serialization, which adds per-record framing (length prefix, CRC checksum, padding). This incurs 15-25% overhead compared to raw formats like HDF5 or MCAP. A 100 GB HDF5 dataset typically becomes 120 GB in RLDS. The overhead buys benefits: TFRecords support efficient sharding for distributed training, automatic checksumming for corruption detection, and native integration with TensorFlow's tf.data pipeline. For bandwidth-constrained scenarios, consider MCAP or Parquet as alternatives; both offer comparable training performance with lower storage overhead.
How does RLDS handle multi-robot datasets with different action spaces?
RLDS uses nested dictionaries for observations and actions, allowing per-robot keys. The Open X-Embodiment dataset contains 21 robot morphologies, each with unique observation keys (franka_wrist_rgb, stretch_head_depth) and action dimensions (7-DOF joint velocities vs 2-DOF mobile base commands). Training code injects embodiment tokens, learned embedding vectors added to the observation dictionary, so policies can condition on robot identity. Action normalization is applied per-dataset during training, not during RLDS creation. This preserves robot-native units in the dataset and enables flexible mixing strategies.
Find datasets covering RLDS
Truelabel surfaces vetted datasets and capture partners working with RLDS. Send the modality, scale, and rights you need and we route you to the closest match.
List Your Robot Dataset on Truelabel