truelabelRequest dataEarnRequest

Physical AI Data Engineering

How to Convert Data to RLDS Format

RLDS (Reinforcement Learning Datasets) is a TensorFlow Datasets extension that standardizes robot demonstration data into episode-structured TFRecords. Converting to RLDS requires auditing source data (HDF5, ROS bags, MCAP), defining a schema with observation/action/reward fields, implementing a TFDS DatasetBuilder that extracts episodes, and validating output against policy training requirements. The format powers Open X-Embodiment, which spans 22 embodiments and 527 skills, and models like RT-1, RT-2, and OpenVLA.

Updated 2026-07-1412 min read
By Truelabel Team
Reviewed by Truelabel Team ·
convert data to RLDS format

Quick facts

Topic
HOW TO Convert Data TO Rlds Format
Audience
Procurement leads, ML ops, robotics engineers
Deliverable
Operational playbook with sample workflow + accept-rule criteria

Why RLDS won as the robot-data interchange format

RLDS (Reinforcement Learning Datasets) came out of Google Research in 2021 to end the format sprawl that made robot demonstrations impossible to pool[1]. The Open X-Embodiment collaboration then standardized 22 embodiments and 527 skills on it, which is why RT-1, RT-2, and OpenVLA can train across embodiments that used to sit in incompatible silos[2].

Three properties make it an interchange layer rather than just another container. Episode boundaries are explicit: is_first, is_last, and is_terminal live on every step, so nobody reverse-engineers where a trajectory starts. Schema metadata (shapes, dtypes, coordinate frames) is embedded in the DatasetInfo, which kills the README archaeology that custom HDF5 layouts force on every new reader. And the on-disk layout loads straight into LeRobot and TF-Agents with no adapter code.

truelabel delivers robot-demonstration datasets in RLDS, LeRobot, MCAP, or a custom schema, each with per-trajectory provenance and contributor consent artifacts attached. The table below is the mental model for why a buyer asks for RLDS in the first place.

PropertyRLDSLeRobotHDF5 / ROS bag
ContainerSharded TFRecordsParquetHDF5 file / rosbag2
Episode boundariesExplicit is_first / is_last / is_terminalepisode_index columnImplicit; caller infers
Schema metadataEmbedded in DatasetInfoEmbedded in dataset cardREADME or out-of-band
Native loadersTFDS, TF-AgentsHugging Face Datasets, PyTorchh5py / rosbags, custom
Random accessPer-shardPer-row-groupSlow on large files
RLDS vs LeRobot vs raw HDF5/ROS for robot demonstration data

Set up the toolchain, then check your conventions

You need Python 3.9+ with TensorFlow 2.15+ and tensorflow-datasets ≥4.9.0 (`pip install tensorflow tensorflow-datasets`). Add h5py for HDF5, rosbags for ROS 1/2 bags, or mcap for MCAP logs. TFRecord and protobuf familiarity helps, but TFDS hides most of the serialization.

The part that actually burns days is not the tooling. It is the conventions your source encodes implicitly: which frame actions live in, how quaternions are ordered, whether images are raw or already JPEG. Get one wrong and the dataset builds cleanly, loads cleanly, and trains a policy that fails on the robot with no error anywhere in the stack. Budget 2-5 days for a first pass: about a day to audit, a day on the builder, a day chasing episode boundaries, one to two validating. The RLDS repository ships reference builders for the common formats.

The conversion loop, end to end

Every conversion runs the same five-step loop. The steps below are the skeleton; the two sections after this one cover the parts that are easy to get subtly wrong (episode extraction and validation).

The schema declaration is the one piece worth seeing in full. For a 7-DoF arm with a wrist camera, `_info` returns a FeaturesDict shaped like `{'steps': Dataset({'observation': {'image': Image(shape=(256,256,3), encoding_format='jpeg'), 'state': Tensor(shape=(7,), dtype=tf.float32)}, 'action': Tensor(shape=(8,), dtype=tf.float32), 'reward': tf.float32, 'is_first': tf.bool, 'is_last': tf.bool, 'is_terminal': tf.bool})}`, with `supervised_keys=None`. Populate homepage, citation, and description too, since those surface in the TFDS catalog. If your source ships predefined splits, emit one SplitGenerator each; otherwise split 90/10 by episode count.

  1. 01

    Audit the source

    Enumerate every field, shape, and dtype: h5py for HDF5, rosbag info or the rosbags library for ROS bags, the mcap CLI for MCAP. Record image resolution, proprioception dimension (a 7-DoF arm is 7), action dimension, and any language strings.

  2. 02

    Define the target schema

    Map each source field to an RLDS feature. A step is observation, action, reward (optional for imitation), and the three boolean flags. Pin the action shape exactly: a 7-DoF arm plus gripper is (8,) concatenated, or {'arm':(7,),'gripper':(1,)} split.

  3. 03

    Scaffold the DatasetBuilder

    Run tfds new my_dataset, subclass tfds.core.GeneratorBasedBuilder, and fill _info, _split_generators, and _generate_examples.

  4. 04

    Extract episodes

    Yield (episode_id, {'steps': steps_list}) per episode. Set is_first on step 0, is_last on the final step, is_terminal only on a real success or failure, never a timeout.

  5. 05

    Build and validate

    Run tfds build, reload, and check lengths, shapes, and action ranges; then train a policy for a few epochs. If the loss never drops, suspect normalization, not the model.

Episode extraction: encoding, normalization, and the three flags

Extraction is where the real decisions happen. For HDF5 with explicit markers (`data['episode_0']`), iterate the keys. For ROS bags with no markers, infer boundaries from topic gaps: a gap over two seconds in `/joint_states` starts a new episode, the heuristic the RoboNet conversion uses across 7 platforms.

Image encoding has one rule people miss: never JPEG a depth map or a segmentation mask. JPEG is lossy 8-bit, so float32 depth collapses to 256 levels and fine geometry is gone. Pass raw uint8 RGB straight to `tfds.features.Image` and let TFDS compress it; store depth as 16-bit PNG or a raw Tensor. The Segments.ai tooling handles depth and LiDAR without lossy compression if you need a reference.

Normalization is the other trap. If the source logs joint velocities in rad/s but the policy wants delta positions, integrate with dt from the timestamps before you write anything. Store the min/max constants you used in the dataset metadata so the policy can invert them at inference; without that, a policy trained on one dataset misfires on another with a different action scale.

The three boolean flags have exact semantics. is_first is step 0, is_last is the final step, is_terminal is true only when the episode ended in a genuine success or failure, not a timeout. Many datasets, DROID included, have no terminal states, so is_terminal stays false everywhere. Reward is optional for imitation; write 0.0 when you have none. A few Open X-Embodiment conventions are worth adopting from the start[2]:

  • Standardize RGB to 256x256 or 224x224 so ImageNet-pretrained encoders load without a resize step.
  • Normalize actions to [-1, 1] and keep the per-dimension min/max constants in dataset metadata.
  • Add a language_instruction field even for non-language policies; generic captions still help zero-shot transfer.
  • Filter obvious junk early: episodes under 10 steps, grasps where the gripper never closes, clips dominated by motion blur.

Validate against a policy, not just shapes

Shape and dtype checks catch the crashes; they do not catch the dataset that trains a broken policy. After `tfds build`, reload with `tfds.load('my_dataset', split='train')` and confirm four things: episode lengths match the source (off-by-one is the usual bug), image shapes are right, action ranges are plausible (gripper in [0,1], joints in [-π,π]), and the is_first/is_last flags land on real boundaries. The Open X-Embodiment validation suite runs exactly these shape, dtype, and range checks across its datasets.

Then actually train. Load the set through the LeRobot Diffusion Policy example and run 10 epochs on 100 episodes. If the loss will not move, the fault is almost always action normalization or observation preprocessing, not the architecture. The OpenVLA codebase ships RLDS dataloaders with built-in sanity checks (action-magnitude histograms, image mean and std) that surface these faster than reading the raw data by hand.

Write the decisions down. A short README covering coordinate frames, action representation, the normalization constants, and the filtering rule ("dropped 5% of episodes under 10 steps") is what a buyer needs for procurement diligence; the Datasheets for Datasets template is a reasonable skeleton.

Pitfalls that produce silent failures

Quaternion order. ROS emits [x,y,z,w]; many libraries expect [w,x,y,z], and the RT-2 codebase wants [x,y,z,w]. An orientation 180 degrees off is this bug nine times out of ten; `scipy.spatial.transform.Rotation` converts.

Off-by-one boundaries. is_last on step N when the source holds N+1 steps means the policy never learns the terminal behavior. Print lengths before and after and assert equality.

RGBA slipping in. TFDS expects (256,256,3) and your images are (256,256,4). Slice `image[:,:,:3]` before writing.

Empty episodes. `_generate_examples` yields a zero-length steps list when a filter is too aggressive or parsing failed quietly; log episode IDs and step counts inside the loop.

Normalization fit on train only. Constants computed on the training split clip validation actions past [-1, 1]. Fit them on train plus validation.

A build that hangs usually means source data on a slow network filesystem or O(n²) extraction; copy to local NVMe and profile `_generate_examples`. RLDS dataloaders also saturate SATA during training, so keep the working set on NVMe. When something is genuinely stuck, the RLDS and LeRobot issue trackers hold most of the known fixes.

Distribute, license, and list the dataset

To share publicly, open a pull request against tensorflow/datasets with your builder and a card following the Hugging Face template; the Data Cards paper adds the procurement fields (collection method, consent, known biases). Host the TFRecords on GCS or S3, and for anything over 100 GB add a torrent or rsync mirror, since academic labs run on thin cloud budgets.

Licensing is per-dataset, not per-format. RLDS itself is Apache 2.0, but content carries its own terms: RoboNet is BSD-3-Clause and commercial-friendly, while EPIC-KITCHENS annotations are CC BY-NC 4.0 and non-commercial. Check every source before you redistribute. If people are identifiable and the data was collected in the EU, GDPR Article 7 consent applies[3]; Ego4D blurs every face and license plate as the reference approach. Industrial physical-AI systems fall under the EU AI Act's high-risk tier, which requires documented collection method, quality control, and known bias[4]. Machine-readable usage terms via the ODRL model let compliance checks run without a lawyer in the loop.

For commercial distribution, list on truelabel's physical AI marketplace: buyers post a spec, matched suppliers return sample packets with QA evidence before any scale commitment, and every trajectory ships with data provenance and consent metadata attached.

Worked example: DROID from HDF5 to RLDS

DROID (350K trajectories, 76 hours) shipped as HDF5. The audit: delta end-effector actions (7-DoF pose plus gripper) in the robot base frame, a 128x128 wrist camera, 10 Hz. The schema follows directly: `observation={'image': Image(128,128,3), 'state': Tensor(7,)}, action=Tensor(8,)`. Extraction iterates the HDF5 episodes and sets is_first/is_last on the boundaries.

The one real snag: DROID stores images as JPEG bytes inside HDF5, and TFDS wants raw arrays or JPEG bytes with specific headers. Decoding to a numpy array and letting TFDS re-encode costs about 10% more build time but removes the ambiguity. As an illustrative sanity check, a Diffusion Policy trained for a few dozen epochs on the converted build should reach real-robot success within a few points of the rate reported in the DROID paper; a wider gap points back at the conversion, not the model. Treat any single success figure as illustrative, since real-robot numbers move with hardware, seed, and evaluation protocol. That build is now the canonical RLDS version behind OpenVLA and LeRobot benchmarks.

The transferable lessons: double your first time estimate, validate with a training run rather than shape checks, and put the normalization constants in the dataset card so the next person does not have to reverse-engineer them.

Use these to move from category-level context into specific task, dataset, format, and comparison detail.

External references and source context

  1. RLDS: an Ecosystem to Generate, Share and Use Datasets in Reinforcement Learning

    RLDS paper introducing the format and ecosystem for standardizing robot learning datasets

    arXiv ↩
  2. Open X-Embodiment: Robotic Learning Datasets and RT-X Models

    Open X-Embodiment collaboration aggregating 22 embodiments and 527 skills in RLDS format

    arXiv ↩
  3. GDPR Article 7 — Conditions for consent

    GDPR Article 7 specifying consent requirements for personal data collection

    GDPR-Info.eu ↩
  4. Regulation (EU) 2024/1689 laying down harmonised rules on artificial intelligence

    EU AI Act Regulation 2024/1689 requiring dataset documentation for high-risk AI systems

    EUR-Lex ↩
  5. scale.com physical ai

    Scale AI Physical AI platform consuming RLDS datasets natively for robot training

    scale.com
  6. NVIDIA Cosmos World Foundation Models

    NVIDIA Cosmos world foundation models using RLDS format for physical AI training

    NVIDIA Developer
  7. rosbag2_storage_mcap

    ROS 2 MCAP storage plugin bridging ROS bags and MCAP format

    GitHub
  8. BridgeData V2: A Dataset for Robot Learning at Scale

    BridgeData V2 paper documenting large-scale robot learning dataset with RLDS conversion

    arXiv
  9. Reading from a ROS bag file

    ROS documentation for reading bag files programmatically

    docs.ros.org
  10. MCAP guides

    MCAP CLI tools and guides for inspecting multi-sensor logs

    MCAP
  11. EPIC-KITCHENS-100 dataset page

    EPIC-KITCHENS-100 dataset page documenting egocentric video dataset structure

    epic-kitchens.github.io
  12. RLDS with TensorFlow Datasets

    TensorFlow RLDS documentation specifying episode structure and validation requirements

    TensorFlow
  13. CALVIN GitHub repository

    CALVIN GitHub repository with language-conditioned manipulation dataset

    GitHub
  14. LeRobot dataset documentation

    LeRobot dataset documentation showing Parquet-based episode structure

    Hugging Face
  15. RT-1: Robotics Transformer for Real-World Control at Scale

    RT-1 paper reporting coordinate frame mismatch impact on task success rates

    arXiv
  16. Project site

    RoboNet dataset wiki with GCS and torrent distribution options

    robonet.wiki
  17. Rescaling Egocentric Vision: Collection, Pipeline and Challenges for EPIC-KITCHENS-100

    EPIC-KITCHENS-100 paper documenting 3 major dataset versions with schema changes

    arXiv
  18. Training ACT with LeRobot Notebook

    ACT training notebook demonstrating multi-view camera setup in LeRobot

    GitHub
  19. Domain Randomization for Transferring Deep Neural Networks from Simulation to the Real World

    Domain randomization paper showing 30% sim-to-real improvement with mixed datasets

    arXiv
  20. RLBench: The Robot Learning Benchmark & Learning Environment

    RLBench paper providing 100 simulated manipulation tasks in RLDS format

    arXiv
  21. DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset

    DROID paper documenting episode filtering improving policy success by 18%

    arXiv
  22. PointNet: Deep Learning on Point Sets for 3D Classification and Segmentation

    PointNet paper processing point clouds directly for 3D classification

    arXiv
  23. Project site

    Dex-YCB dataset with tactile sensor data showing 25% manipulation improvement

    dex-ycb.github.io
  24. Project site

    RH20T dataset with 4-camera setup for multi-view manipulation

    rh20t.github.io
  25. RLDS with TensorFlow Datasets

    TFDS Image feature documentation for automatic JPEG compression

    TensorFlow
  26. RoboNet dataset license

    RoboNet dataset BSD-3-Clause license permitting commercial use

    GitHub raw content
  27. EPIC-KITCHENS-100 annotations license

    EPIC-KITCHENS annotations non-commercial CC BY-NC 4.0 license

    GitHub
  28. Subpart 27.4 - Rights in Data and Copyrights

    FAR Subpart 27.4 specifying data rights in government contracts

    acquisition.gov
  29. World Models

    World Models paper showing 50% real data reduction with learned world models

    worldmodels.github.io
  30. NVIDIA GR00T N1 technical report

    NVIDIA GR00T N1 technical report documenting 10M hours simulated pretraining

    arXiv

FAQ

What is the difference between RLDS and LeRobot dataset formats?

RLDS uses TensorFlow TFRecords with a nested episode structure (Dataset of steps), while LeRobot uses Parquet files with a flat table structure (one row per step, episode_id column for grouping). RLDS integrates natively with TensorFlow Datasets and TF-Agents; LeRobot integrates with Hugging Face Datasets and PyTorch. Both support the same observation/action/reward schema. LeRobot can load RLDS datasets via a backend adapter. For new datasets, choose based on your training framework: TensorFlow → RLDS, PyTorch → LeRobot. The Open X-Embodiment project uses RLDS; Hugging Face robotics benchmarks use LeRobot.

How do I handle datasets with variable-length episodes in RLDS?

RLDS supports variable-length episodes natively: each episode is a separate example in the Dataset, and episodes can have different step counts. In your DatasetBuilder, yield episodes as lists of steps: `yield (episode_id, {'steps': steps_list})` where `len(steps_list)` varies per episode. TFDS serializes each episode independently. During training, use `tf.data.Dataset.padded_batch` to pad episodes to the same length within a batch, or use `bucket_by_sequence_length` to group similar-length episodes. Padding to a common length within a batch is standard practice and does not degrade training when the episodes batched together are of similar length.

Can I convert ROS 2 bags directly to RLDS without intermediate formats?

Yes, use the rosbags library (Python) to read ROS 2 bags and extract messages. Install via `pip install rosbags`. Iterate over topics (/camera/image_raw, /joint_states, /cmd_vel), deserialize messages, and map them to RLDS observation/action fields. For multi-topic synchronization, use message timestamps to align camera frames with joint states (typically within 10 ms). The rosbag2_storage_mcap plugin allows reading ROS 2 bags as MCAP files, which have better random-access performance for large datasets. The DROID and BridgeData V2 conversion scripts demonstrate ROS bag → RLDS pipelines.

What is the recommended image resolution for RLDS datasets?

Use 224×224 or 256×256 for RGB images. These resolutions match ImageNet pretraining (224×224) and are large enough to preserve manipulation-relevant details (object edges, gripper position) while keeping dataset size manageable. The Open X-Embodiment project standardized on 256×256 across its 22 embodiments. Higher resolutions (512×512, 640×480) increase storage by 4-9× and training time by 2-3× with minimal accuracy gain for manipulation tasks. For navigation or fine-grained assembly, 512×512 may be justified. Store original resolution in a separate 'image_highres' field if needed for future use.

How do I add language annotations to an existing RLDS dataset?

Load the RLDS dataset with `tfds.load`, iterate over episodes, generate captions (manually or with GPT-4V/CLIP), and rebuild the dataset with an updated schema that includes `language_instruction: tfds.features.Text`. Store captions in a separate JSON file (episode_id → caption mapping) during generation, then merge during rebuild. The RT-2 paper shows that even generic captions ("pick up object", "place in container") improve zero-shot transfer. For datasets with 10K+ episodes, use a GPT-4V batch API or BLIP-2 (free, lower quality). Document caption source in dataset card.

What are the storage requirements for a typical RLDS manipulation dataset?

A 10K-episode dataset with 256×256 RGB images (JPEG quality 95), 10 Hz sampling, 20-second average episode length, 7-DoF proprioception, and 8-DoF actions requires approximately 15-25 GB. Breakdown: images dominate at ~10 KB per JPEG-compressed frame; a 20-second episode at 10 Hz is about 200 frames, so ~2 MB per episode × 10K episodes ≈ 20 GB. Proprioception and actions are ~1 KB per step, negligible. Add roughly 20% overhead for TFRecord metadata and sharding. Corpus-scale collections such as Open X-Embodiment run into the terabytes. Use cloud storage with lifecycle policies (move to Coldline after 90 days) to cut standing costs.

Looking for convert data to RLDS format?

Specify modality, task, environment, requested rights posture, and delivery format. Truelabel routes the request to candidate capture partners and helps scope consent/provenance artifacts and commercial licensing requirements for buyer review before delivery.

List Your RLDS Dataset on Truelabel