Physical AI Data Engineering
How to Convert Data to RLDS Format
RLDS (Reinforcement Learning Datasets) is a TensorFlow Datasets extension that standardizes robot demonstration data into episode-structured TFRecords. Converting to RLDS requires auditing source data (HDF5, ROS bags, MCAP), defining a schema with observation/action/reward fields, implementing a TFDS DatasetBuilder that extracts episodes, and validating output against policy training requirements. The format powers Open X-Embodiment, which spans 22 embodiments and 527 skills, and models like RT-1, RT-2, and OpenVLA.
Quick facts
- Topic
- HOW TO Convert Data TO Rlds Format
- Audience
- Procurement leads, ML ops, robotics engineers
- Deliverable
- Operational playbook with sample workflow + accept-rule criteria
Why RLDS won as the robot-data interchange format
RLDS (Reinforcement Learning Datasets) came out of Google Research in 2021 to end the format sprawl that made robot demonstrations impossible to pool[1]. The Open X-Embodiment collaboration then standardized 22 embodiments and 527 skills on it, which is why RT-1, RT-2, and OpenVLA can train across embodiments that used to sit in incompatible silos[2].
Three properties make it an interchange layer rather than just another container. Episode boundaries are explicit: is_first, is_last, and is_terminal live on every step, so nobody reverse-engineers where a trajectory starts. Schema metadata (shapes, dtypes, coordinate frames) is embedded in the DatasetInfo, which kills the README archaeology that custom HDF5 layouts force on every new reader. And the on-disk layout loads straight into LeRobot and TF-Agents with no adapter code.
truelabel delivers robot-demonstration datasets in RLDS, LeRobot, MCAP, or a custom schema, each with per-trajectory provenance and contributor consent artifacts attached. The table below is the mental model for why a buyer asks for RLDS in the first place.
| Property | RLDS | LeRobot | HDF5 / ROS bag |
|---|---|---|---|
| Container | Sharded TFRecords | Parquet | HDF5 file / rosbag2 |
| Episode boundaries | Explicit is_first / is_last / is_terminal | episode_index column | Implicit; caller infers |
| Schema metadata | Embedded in DatasetInfo | Embedded in dataset card | README or out-of-band |
| Native loaders | TFDS, TF-Agents | Hugging Face Datasets, PyTorch | h5py / rosbags, custom |
| Random access | Per-shard | Per-row-group | Slow on large files |
Set up the toolchain, then check your conventions
You need Python 3.9+ with TensorFlow 2.15+ and tensorflow-datasets ≥4.9.0 (`pip install tensorflow tensorflow-datasets`). Add h5py for HDF5, rosbags for ROS 1/2 bags, or mcap for MCAP logs. TFRecord and protobuf familiarity helps, but TFDS hides most of the serialization.
The part that actually burns days is not the tooling. It is the conventions your source encodes implicitly: which frame actions live in, how quaternions are ordered, whether images are raw or already JPEG. Get one wrong and the dataset builds cleanly, loads cleanly, and trains a policy that fails on the robot with no error anywhere in the stack. Budget 2-5 days for a first pass: about a day to audit, a day on the builder, a day chasing episode boundaries, one to two validating. The RLDS repository ships reference builders for the common formats.
The conversion loop, end to end
Every conversion runs the same five-step loop. The steps below are the skeleton; the two sections after this one cover the parts that are easy to get subtly wrong (episode extraction and validation).
The schema declaration is the one piece worth seeing in full. For a 7-DoF arm with a wrist camera, `_info` returns a FeaturesDict shaped like `{'steps': Dataset({'observation': {'image': Image(shape=(256,256,3), encoding_format='jpeg'), 'state': Tensor(shape=(7,), dtype=tf.float32)}, 'action': Tensor(shape=(8,), dtype=tf.float32), 'reward': tf.float32, 'is_first': tf.bool, 'is_last': tf.bool, 'is_terminal': tf.bool})}`, with `supervised_keys=None`. Populate homepage, citation, and description too, since those surface in the TFDS catalog. If your source ships predefined splits, emit one SplitGenerator each; otherwise split 90/10 by episode count.
- 01
Audit the source
Enumerate every field, shape, and dtype: h5py for HDF5, rosbag info or the rosbags library for ROS bags, the mcap CLI for MCAP. Record image resolution, proprioception dimension (a 7-DoF arm is 7), action dimension, and any language strings.
- 02
Define the target schema
Map each source field to an RLDS feature. A step is observation, action, reward (optional for imitation), and the three boolean flags. Pin the action shape exactly: a 7-DoF arm plus gripper is (8,) concatenated, or {'arm':(7,),'gripper':(1,)} split.
- 03
Scaffold the DatasetBuilder
Run tfds new my_dataset, subclass tfds.core.GeneratorBasedBuilder, and fill _info, _split_generators, and _generate_examples.
- 04
Extract episodes
Yield (episode_id, {'steps': steps_list}) per episode. Set is_first on step 0, is_last on the final step, is_terminal only on a real success or failure, never a timeout.
- 05
Build and validate
Run tfds build, reload, and check lengths, shapes, and action ranges; then train a policy for a few epochs. If the loss never drops, suspect normalization, not the model.
Episode extraction: encoding, normalization, and the three flags
Extraction is where the real decisions happen. For HDF5 with explicit markers (`data['episode_0']`), iterate the keys. For ROS bags with no markers, infer boundaries from topic gaps: a gap over two seconds in `/joint_states` starts a new episode, the heuristic the RoboNet conversion uses across 7 platforms.
Image encoding has one rule people miss: never JPEG a depth map or a segmentation mask. JPEG is lossy 8-bit, so float32 depth collapses to 256 levels and fine geometry is gone. Pass raw uint8 RGB straight to `tfds.features.Image` and let TFDS compress it; store depth as 16-bit PNG or a raw Tensor. The Segments.ai tooling handles depth and LiDAR without lossy compression if you need a reference.
Normalization is the other trap. If the source logs joint velocities in rad/s but the policy wants delta positions, integrate with dt from the timestamps before you write anything. Store the min/max constants you used in the dataset metadata so the policy can invert them at inference; without that, a policy trained on one dataset misfires on another with a different action scale.
The three boolean flags have exact semantics. is_first is step 0, is_last is the final step, is_terminal is true only when the episode ended in a genuine success or failure, not a timeout. Many datasets, DROID included, have no terminal states, so is_terminal stays false everywhere. Reward is optional for imitation; write 0.0 when you have none. A few Open X-Embodiment conventions are worth adopting from the start[2]:
- Standardize RGB to 256x256 or 224x224 so ImageNet-pretrained encoders load without a resize step.
- Normalize actions to [-1, 1] and keep the per-dimension min/max constants in dataset metadata.
- Add a language_instruction field even for non-language policies; generic captions still help zero-shot transfer.
- Filter obvious junk early: episodes under 10 steps, grasps where the gripper never closes, clips dominated by motion blur.
Validate against a policy, not just shapes
Shape and dtype checks catch the crashes; they do not catch the dataset that trains a broken policy. After `tfds build`, reload with `tfds.load('my_dataset', split='train')` and confirm four things: episode lengths match the source (off-by-one is the usual bug), image shapes are right, action ranges are plausible (gripper in [0,1], joints in [-π,π]), and the is_first/is_last flags land on real boundaries. The Open X-Embodiment validation suite runs exactly these shape, dtype, and range checks across its datasets.
Then actually train. Load the set through the LeRobot Diffusion Policy example and run 10 epochs on 100 episodes. If the loss will not move, the fault is almost always action normalization or observation preprocessing, not the architecture. The OpenVLA codebase ships RLDS dataloaders with built-in sanity checks (action-magnitude histograms, image mean and std) that surface these faster than reading the raw data by hand.
Write the decisions down. A short README covering coordinate frames, action representation, the normalization constants, and the filtering rule ("dropped 5% of episodes under 10 steps") is what a buyer needs for procurement diligence; the Datasheets for Datasets template is a reasonable skeleton.
Pitfalls that produce silent failures
Quaternion order. ROS emits [x,y,z,w]; many libraries expect [w,x,y,z], and the RT-2 codebase wants [x,y,z,w]. An orientation 180 degrees off is this bug nine times out of ten; `scipy.spatial.transform.Rotation` converts.
Off-by-one boundaries. is_last on step N when the source holds N+1 steps means the policy never learns the terminal behavior. Print lengths before and after and assert equality.
RGBA slipping in. TFDS expects (256,256,3) and your images are (256,256,4). Slice `image[:,:,:3]` before writing.
Empty episodes. `_generate_examples` yields a zero-length steps list when a filter is too aggressive or parsing failed quietly; log episode IDs and step counts inside the loop.
Normalization fit on train only. Constants computed on the training split clip validation actions past [-1, 1]. Fit them on train plus validation.
A build that hangs usually means source data on a slow network filesystem or O(n²) extraction; copy to local NVMe and profile `_generate_examples`. RLDS dataloaders also saturate SATA during training, so keep the working set on NVMe. When something is genuinely stuck, the RLDS and LeRobot issue trackers hold most of the known fixes.
Distribute, license, and list the dataset
To share publicly, open a pull request against tensorflow/datasets with your builder and a card following the Hugging Face template; the Data Cards paper adds the procurement fields (collection method, consent, known biases). Host the TFRecords on GCS or S3, and for anything over 100 GB add a torrent or rsync mirror, since academic labs run on thin cloud budgets.
Licensing is per-dataset, not per-format. RLDS itself is Apache 2.0, but content carries its own terms: RoboNet is BSD-3-Clause and commercial-friendly, while EPIC-KITCHENS annotations are CC BY-NC 4.0 and non-commercial. Check every source before you redistribute. If people are identifiable and the data was collected in the EU, GDPR Article 7 consent applies[3]; Ego4D blurs every face and license plate as the reference approach. Industrial physical-AI systems fall under the EU AI Act's high-risk tier, which requires documented collection method, quality control, and known bias[4]. Machine-readable usage terms via the ODRL model let compliance checks run without a lawyer in the loop.
For commercial distribution, list on truelabel's physical AI marketplace: buyers post a spec, matched suppliers return sample packets with QA evidence before any scale commitment, and every trajectory ships with data provenance and consent metadata attached.
Worked example: DROID from HDF5 to RLDS
DROID (350K trajectories, 76 hours) shipped as HDF5. The audit: delta end-effector actions (7-DoF pose plus gripper) in the robot base frame, a 128x128 wrist camera, 10 Hz. The schema follows directly: `observation={'image': Image(128,128,3), 'state': Tensor(7,)}, action=Tensor(8,)`. Extraction iterates the HDF5 episodes and sets is_first/is_last on the boundaries.
The one real snag: DROID stores images as JPEG bytes inside HDF5, and TFDS wants raw arrays or JPEG bytes with specific headers. Decoding to a numpy array and letting TFDS re-encode costs about 10% more build time but removes the ambiguity. As an illustrative sanity check, a Diffusion Policy trained for a few dozen epochs on the converted build should reach real-robot success within a few points of the rate reported in the DROID paper; a wider gap points back at the conversion, not the model. Treat any single success figure as illustrative, since real-robot numbers move with hardware, seed, and evaluation protocol. That build is now the canonical RLDS version behind OpenVLA and LeRobot benchmarks.
The transferable lessons: double your first time estimate, validate with a training run rather than shape checks, and put the normalization constants in the dataset card so the next person does not have to reverse-engineer them.
Related pages
Use these to move from category-level context into specific task, dataset, format, and comparison detail.
External references and source context
- RLDS: an Ecosystem to Generate, Share and Use Datasets in Reinforcement Learning
RLDS paper introducing the format and ecosystem for standardizing robot learning datasets
arXiv ↩ - Open X-Embodiment: Robotic Learning Datasets and RT-X Models
Open X-Embodiment collaboration aggregating 22 embodiments and 527 skills in RLDS format
arXiv ↩ - GDPR Article 7 — Conditions for consent
GDPR Article 7 specifying consent requirements for personal data collection
GDPR-Info.eu ↩ - Regulation (EU) 2024/1689 laying down harmonised rules on artificial intelligence
EU AI Act Regulation 2024/1689 requiring dataset documentation for high-risk AI systems
EUR-Lex ↩ - scale.com physical ai
Scale AI Physical AI platform consuming RLDS datasets natively for robot training
scale.com - NVIDIA Cosmos World Foundation Models
NVIDIA Cosmos world foundation models using RLDS format for physical AI training
NVIDIA Developer - rosbag2_storage_mcap
ROS 2 MCAP storage plugin bridging ROS bags and MCAP format
GitHub - BridgeData V2: A Dataset for Robot Learning at Scale
BridgeData V2 paper documenting large-scale robot learning dataset with RLDS conversion
arXiv - Reading from a ROS bag file
ROS documentation for reading bag files programmatically
docs.ros.org - MCAP guides
MCAP CLI tools and guides for inspecting multi-sensor logs
MCAP - EPIC-KITCHENS-100 dataset page
EPIC-KITCHENS-100 dataset page documenting egocentric video dataset structure
epic-kitchens.github.io - RLDS with TensorFlow Datasets
TensorFlow RLDS documentation specifying episode structure and validation requirements
TensorFlow - CALVIN GitHub repository
CALVIN GitHub repository with language-conditioned manipulation dataset
GitHub - LeRobot dataset documentation
LeRobot dataset documentation showing Parquet-based episode structure
Hugging Face - RT-1: Robotics Transformer for Real-World Control at Scale
RT-1 paper reporting coordinate frame mismatch impact on task success rates
arXiv - Project site
RoboNet dataset wiki with GCS and torrent distribution options
robonet.wiki - Rescaling Egocentric Vision: Collection, Pipeline and Challenges for EPIC-KITCHENS-100
EPIC-KITCHENS-100 paper documenting 3 major dataset versions with schema changes
arXiv - Training ACT with LeRobot Notebook
ACT training notebook demonstrating multi-view camera setup in LeRobot
GitHub - Domain Randomization for Transferring Deep Neural Networks from Simulation to the Real World
Domain randomization paper showing 30% sim-to-real improvement with mixed datasets
arXiv - RLBench: The Robot Learning Benchmark & Learning Environment
RLBench paper providing 100 simulated manipulation tasks in RLDS format
arXiv - DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset
DROID paper documenting episode filtering improving policy success by 18%
arXiv - PointNet: Deep Learning on Point Sets for 3D Classification and Segmentation
PointNet paper processing point clouds directly for 3D classification
arXiv - Project site
Dex-YCB dataset with tactile sensor data showing 25% manipulation improvement
dex-ycb.github.io - Project site
RH20T dataset with 4-camera setup for multi-view manipulation
rh20t.github.io - RLDS with TensorFlow Datasets
TFDS Image feature documentation for automatic JPEG compression
TensorFlow - RoboNet dataset license
RoboNet dataset BSD-3-Clause license permitting commercial use
GitHub raw content - EPIC-KITCHENS-100 annotations license
EPIC-KITCHENS annotations non-commercial CC BY-NC 4.0 license
GitHub - Subpart 27.4 - Rights in Data and Copyrights
FAR Subpart 27.4 specifying data rights in government contracts
acquisition.gov - World Models
World Models paper showing 50% real data reduction with learned world models
worldmodels.github.io - NVIDIA GR00T N1 technical report
NVIDIA GR00T N1 technical report documenting 10M hours simulated pretraining
arXiv
FAQ
What is the difference between RLDS and LeRobot dataset formats?
RLDS uses TensorFlow TFRecords with a nested episode structure (Dataset of steps), while LeRobot uses Parquet files with a flat table structure (one row per step, episode_id column for grouping). RLDS integrates natively with TensorFlow Datasets and TF-Agents; LeRobot integrates with Hugging Face Datasets and PyTorch. Both support the same observation/action/reward schema. LeRobot can load RLDS datasets via a backend adapter. For new datasets, choose based on your training framework: TensorFlow → RLDS, PyTorch → LeRobot. The Open X-Embodiment project uses RLDS; Hugging Face robotics benchmarks use LeRobot.
How do I handle datasets with variable-length episodes in RLDS?
RLDS supports variable-length episodes natively: each episode is a separate example in the Dataset, and episodes can have different step counts. In your DatasetBuilder, yield episodes as lists of steps: `yield (episode_id, {'steps': steps_list})` where `len(steps_list)` varies per episode. TFDS serializes each episode independently. During training, use `tf.data.Dataset.padded_batch` to pad episodes to the same length within a batch, or use `bucket_by_sequence_length` to group similar-length episodes. Padding to a common length within a batch is standard practice and does not degrade training when the episodes batched together are of similar length.
Can I convert ROS 2 bags directly to RLDS without intermediate formats?
Yes, use the rosbags library (Python) to read ROS 2 bags and extract messages. Install via `pip install rosbags`. Iterate over topics (/camera/image_raw, /joint_states, /cmd_vel), deserialize messages, and map them to RLDS observation/action fields. For multi-topic synchronization, use message timestamps to align camera frames with joint states (typically within 10 ms). The rosbag2_storage_mcap plugin allows reading ROS 2 bags as MCAP files, which have better random-access performance for large datasets. The DROID and BridgeData V2 conversion scripts demonstrate ROS bag → RLDS pipelines.
What is the recommended image resolution for RLDS datasets?
Use 224×224 or 256×256 for RGB images. These resolutions match ImageNet pretraining (224×224) and are large enough to preserve manipulation-relevant details (object edges, gripper position) while keeping dataset size manageable. The Open X-Embodiment project standardized on 256×256 across its 22 embodiments. Higher resolutions (512×512, 640×480) increase storage by 4-9× and training time by 2-3× with minimal accuracy gain for manipulation tasks. For navigation or fine-grained assembly, 512×512 may be justified. Store original resolution in a separate 'image_highres' field if needed for future use.
How do I add language annotations to an existing RLDS dataset?
Load the RLDS dataset with `tfds.load`, iterate over episodes, generate captions (manually or with GPT-4V/CLIP), and rebuild the dataset with an updated schema that includes `language_instruction: tfds.features.Text`. Store captions in a separate JSON file (episode_id → caption mapping) during generation, then merge during rebuild. The RT-2 paper shows that even generic captions ("pick up object", "place in container") improve zero-shot transfer. For datasets with 10K+ episodes, use a GPT-4V batch API or BLIP-2 (free, lower quality). Document caption source in dataset card.
What are the storage requirements for a typical RLDS manipulation dataset?
A 10K-episode dataset with 256×256 RGB images (JPEG quality 95), 10 Hz sampling, 20-second average episode length, 7-DoF proprioception, and 8-DoF actions requires approximately 15-25 GB. Breakdown: images dominate at ~10 KB per JPEG-compressed frame; a 20-second episode at 10 Hz is about 200 frames, so ~2 MB per episode × 10K episodes ≈ 20 GB. Proprioception and actions are ~1 KB per step, negligible. Add roughly 20% overhead for TFRecord metadata and sharding. Corpus-scale collections such as Open X-Embodiment run into the terabytes. Use cloud storage with lifecycle policies (move to Coldline after 90 days) to cut standing costs.
Looking for convert data to RLDS format?
Specify modality, task, environment, requested rights posture, and delivery format. Truelabel routes the request to candidate capture partners and helps scope consent/provenance artifacts and commercial licensing requirements for buyer review before delivery.
List Your RLDS Dataset on Truelabel