Dataset profile
IPEC-COMMUNITY/language_table_lerobot Dataset Profile
IPEC-COMMUNITY/language_table_lerobot is an Apache-2.0 licensed robotics dataset containing 442,226 episodes captured with xArm manipulators, totaling over 7 million frames across 127,605 distinct language-conditioned tasks recorded at 10 FPS. Built using the LeRobot v2.0 framework and derived from the Language Table benchmark via OpenX RLDS conversion, this dataset enables physical-AI teams to train vision-language-action models for tabletop manipulation tasks where natural language instructions guide robotic pick-and-place, sorting, and rearrangement behaviors. The permissive Apache-2.0 license permits commercial deployment, making it suitable for production VLA systems, teleoperation bootstrapping, and world-model pretraining in warehouse automation, collaborative assembly, and household robotics applications.
Quick facts
- Scale
- 442,226 episodes / 7M+ frames
- License
- Apache-2.0 (commercial use permitted)
- Robot Platform
- xArm manipulator
- Task Count
- 127,605 language-conditioned tasks
- Format
- LeRobot v2.0 / Parquet chunks
- Frame Rate
- 10 FPS
Dataset Composition and Structure
The dataset organizes 442,226 episodes into 443 chunks of approximately 1,000 episodes each, stored as Parquet files following the LeRobot v2.0 specification. Each episode represents a complete manipulation trajectory on an xArm robot responding to a natural language instruction, with observations and actions recorded at 10 frames per second. The 127,605 unique tasks span language-conditioned tabletop scenarios including object placement, spatial sorting, color-based rearrangement, and instruction-following pick-and-place operations characteristic of the Language Table benchmark paradigm.
All 442,226 episodes include synchronized video recordings stored alongside trajectory data, enabling multimodal learning pipelines that consume both visual observations and proprioceptive state. The chunked Parquet architecture supports efficient random access during training while maintaining compatibility with the broader LeRobot ecosystem, including data loaders, normalization utilities, and policy training scripts. This structural design allows teams to subset tasks by language complexity, sample episodes for few-shot evaluation, or stream the full corpus for large-scale VLA pretraining without custom preprocessing.
Licensing and Commercial Use
Released under the Apache License 2.0, this dataset grants explicit permission for commercial use, modification, and redistribution with minimal restrictions. Robotics teams can train proprietary models on this data, deploy resulting policies in commercial products, and distribute derivative datasets without royalty obligations, provided they include the original license notice and state any modifications made. The permissive terms eliminate procurement friction common with research-only datasets, accelerating time-to-deployment for startups and enterprise AI labs building production manipulation systems.
The Apache-2.0 grant covers both the trajectory data and accompanying video recordings, meaning visual encoders, inverse dynamics models, and end-to-end VLA architectures trained on this corpus inherit no restrictive downstream licensing. Teams should verify that any additional datasets mixed during multi-source pretraining carry compatible licenses, but the IPEC-COMMUNITY release itself imposes no constraints on model weights, inference services, or robotic hardware integrations that leverage learned representations from these 442,226 episodes.
Procurement and Integration Considerations
The dataset resides on HuggingFace Hub with 919,489 recorded downloads, indicating broad community adoption and stable hosting infrastructure suitable for production data pipelines. Teams can ingest the chunked Parquet files directly via the LeRobot library or implement custom loaders using standard Arrow/Pandas tooling, with the 1,000-episode chunk size balancing storage efficiency against parallel loading performance on multi-GPU training clusters. The 10 FPS capture rate matches real-time control frequencies for many commercial manipulators, reducing temporal downsampling needs during policy rollout.
Integration with existing VLA pipelines requires validating observation and action space alignment between the xArm platform and target deployment hardware. The Language Table task distribution emphasizes tabletop manipulation with discrete objects rather than contact-rich assembly or deformable material handling, so teams targeting those domains should plan for sim-to-real transfer, domain randomization, or supplemental data collection. The dataset's origin as an OpenX RLDS conversion means metadata schemas follow established robotics data standards, easing interoperability with RT-1, RT-2, and Octo-style transformer policies.
Known Limitations and Scope Boundaries
While the 127,605 tasks provide substantial language diversity, the underlying physical scenarios remain constrained to tabletop environments with rigid objects and overhead camera viewpoints typical of the Language Table benchmark. Teams building manipulators for bin-picking, flexible assembly, or human-robot handover will find limited transfer without additional data covering varied camera angles, object deformability, and contact dynamics. The xArm platform's kinematic structure and workspace dimensions may not generalize to mobile manipulators, dual-arm systems, or soft grippers without retraining action decoders and inverse kinematics modules.
The dataset lacks explicit annotations for failure modes, recoveries, or safety-critical edge cases, as episodes represent successful task completions from the original benchmark collection. Production deployments requiring robustness to occlusions, adversarial instructions, or out-of-distribution objects should augment training with targeted failure data or incorporate online learning loops. Additionally, while the 10 FPS sampling captures gross motion, high-frequency contact events and force profiles are not preserved, limiting applicability for tasks requiring precise torque control or impedance modulation during insertion or fragile object handling.
Related pages
Use these to move from category-level context into specific task, dataset, format, and comparison detail.
FAQ
What is the IPEC-COMMUNITY/language_table_lerobot dataset and what does it contain?
This is a large-scale robotics manipulation dataset containing 442,226 episodes of xArm robot trajectories performing language-conditioned tabletop tasks, totaling over 7 million frames. Each episode pairs natural language instructions with corresponding observation sequences and action trajectories recorded at 10 frames per second, covering 127,605 distinct manipulation objectives derived from the Language Table benchmark. The dataset ships in LeRobot v2.0 format with synchronized video recordings and is structured as 443 Parquet chunks for efficient training pipeline integration. Originally converted from the OpenX RLDS repository, the data captures pick-and-place, sorting, and spatial rearrangement tasks where language instructions specify target object configurations. All episodes represent successful task completions on an xArm manipulator platform, making the dataset suitable for supervised learning of vision-language-action policies, inverse dynamics modeling, and world-model pretraining for tabletop manipulation domains.
What are the licensing terms and can I use this dataset commercially?
The dataset is released under the Apache License 2.0, which explicitly permits commercial use, modification, distribution, and private deployment without royalty payments. Robotics companies can train proprietary VLA models on this data, integrate resulting policies into commercial products, and deploy manipulation systems in revenue-generating applications provided they retain the Apache-2.0 license notice in any redistributed portions of the dataset. There are no field-of-use restrictions or publication requirements attached to the license. The permissive terms cover both trajectory data and video recordings, meaning visual representations learned from the 442,226 episodes carry no downstream licensing encumbrances. Teams building multi-dataset training pipelines should verify compatibility of other data sources, but this IPEC-COMMUNITY release itself imposes no constraints on model commercialization, cloud inference services, or embedded deployment in robotic hardware products.
Which robotics teams should prioritize this dataset for their training pipelines?
Teams developing vision-language-action models for tabletop manipulation with language-conditioned task specifications will find the highest immediate value, particularly those targeting warehouse sorting, kitting operations, collaborative assembly prep, or household decluttering scenarios. The 127,605 language-task pairs provide substantial instruction diversity for training language encoders and grounding modules, while the xArm trajectories offer realistic action distributions for 6-DOF arm control. Startups and research labs building on the LeRobot ecosystem benefit from native format compatibility and established data loading utilities. Organizations requiring large-scale pretraining data for foundation models in manipulation can leverage the 7 million frames to learn visual representations, dynamics models, or inverse kinematics priors before fine-tuning on proprietary task distributions. The Apache-2.0 license and 919,489 download count indicate production-readiness and community validation, reducing procurement risk compared to bespoke data collection for initial VLA prototyping phases.
When is this dataset NOT the right choice for a robotics project?
Teams building systems for contact-rich manipulation, deformable object handling, mobile manipulation, or dual-arm coordination should look elsewhere, as the Language Table scenarios focus exclusively on rigid-body tabletop tasks with fixed-base single-arm kinematics. The dataset lacks force-torque observations, tactile feedback, and high-frequency control signals necessary for precise insertion, assembly, or fragile object manipulation. Applications requiring bin-picking from clutter, human-robot handover, or navigation-manipulation integration will find limited transfer due to the constrained camera viewpoints and workspace geometry. Projects demanding failure recovery data, safety-critical edge case coverage, or adversarial robustness should recognize that episodes represent successful completions without explicit annotations for recoveries, collisions, or out-of-distribution handling. The 10 FPS sampling rate may prove insufficient for dynamic catching, high-speed pick-and-place, or reactive replanning scenarios requiring sub-100ms control loops. Finally, teams deploying on manipulator platforms with significantly different kinematic structures from the xArm will need additional domain adaptation or supplemental data to bridge action space gaps.
How should I integrate this dataset with existing VLA training infrastructure?
The LeRobot v2.0 format enables direct ingestion via the official LeRobot library's dataset loaders, which handle Parquet deserialization, video decoding, and trajectory normalization out of the box. Teams can load episodes by chunk index for distributed training or stream the full 442,226-episode corpus with the library's built-in samplers, leveraging the 1,000-episode chunk granularity for efficient multi-GPU data parallelism. For custom pipelines, the Parquet files are readable via standard Arrow or Pandas tools, with episode structure documented in the meta/info.json manifest file included in the repository. Observation and action space alignment requires validating that your target robot's state representation matches the xArm's joint configuration and end-effector pose encoding. The 10 FPS capture rate aligns with many commercial manipulator control frequencies, but teams running higher-rate policies should implement temporal interpolation or train action chunking modules. When mixing with other datasets for multi-source pretraining, ensure consistent normalization schemes for pixel observations, language embeddings, and action vectors, using the provided splits metadata to maintain clean train-validation boundaries across the 442,226 episodes.
What are the key technical specifications I need to validate before procurement?
Verify that your training infrastructure can handle the full dataset size of 7,045,476 frames distributed across 443 Parquet chunks, with each episode averaging roughly 16 frames at 10 FPS. Storage requirements include both trajectory data and 442,226 video files, so provision sufficient disk capacity and network bandwidth if streaming from HuggingFace Hub during training. The xArm action space consists of joint velocities or end-effector deltas depending on the control mode encoded in the trajectories, requiring inspection of sample episodes to confirm compatibility with your policy architecture's output head. The language task distribution spans 127,605 unique instructions, but teams should analyze the semantic diversity and linguistic complexity relative to their deployment domain's natural language interface requirements. Camera viewpoints in Language Table are predominantly overhead and fixed, so applications requiring viewpoint invariance or ego-centric perception may need augmentation strategies. Finally, confirm that the Apache-2.0 license terms satisfy your organization's legal and compliance requirements for commercial model training and deployment before integrating the dataset into production pipelines.
Need data like IPEC-COMMUNITY/language_table_lerobot Dataset Profile?
If your project needs similar modality, scale, or licensing, truelabel can surface comparable open datasets or match you with capture partners that deliver to spec.
Access dataset on HuggingFace