truelabelRequest dataEarnRequest

Dataset profile

IPEC-COMMUNITY/kuka_lerobot: Large-Scale KUKA iiwa Manipulation Dataset

IPEC-COMMUNITY/kuka_lerobot is an Apache-2.0 licensed robotics dataset containing 209,880 episodes captured from KUKA iiwa manipulators, totaling approximately 2.46 million frames at 10 fps. Built with the LeRobot framework and distributed in Parquet format with synchronized video, this dataset provides dense manipulation trajectories suitable for training vision-language-action models, imitation learning policies, and world models requiring industrial-arm kinematics. Teams building VLA foundation models or teleoperation systems for 7-DOF manipulators can leverage this scale for pre-training or fine-tuning without commercial restrictions.

Updated 2026-08-318 min read
By Truelabel Team
Reviewed by Truelabel Team ·
KUKA iiwa robotics dataset

Quick facts

Scale
209,880 episodes, 2.46M frames
License
Apache-2.0
Format
Parquet + video
Modality
Video
Robot type
KUKA iiwa (7-DOF)
Commercial use
Permitted

Dataset composition and structure

The IPEC-COMMUNITY/kuka_lerobot dataset comprises 209,880 recorded episodes from KUKA iiwa manipulators, each episode stored as synchronized Parquet files containing state and action trajectories alongside corresponding video observations. Captured at 10 frames per second, the collection spans 2.46 million total frames distributed across 210 chunks of approximately 1,000 episodes each, facilitating efficient streaming and partial downloads for teams with bandwidth or storage constraints. All episodes represent a single task category, providing depth rather than breadth for manipulation scenarios involving the 7-degree-of-freedom KUKA iiwa arm.

Built using the LeRobot v2.0 codebase, the dataset follows standardized schemas compatible with OpenX format conventions, enabling drop-in integration with popular imitation learning and policy training frameworks. Each episode includes the data path pattern pointing to chunked Parquet files and a corresponding video path, allowing researchers to load visual observations and proprioceptive state in lockstep. The train split encompasses the full 209,880 episodes, and no validation or test splits are pre-defined, leaving hold-out strategy to the downstream user.

Licensing and commercial deployment

Released under the Apache-2.0 license, IPEC-COMMUNITY/kuka_lerobot permits commercial use, modification, and redistribution with minimal restrictions, making it suitable for both academic research and production robotics applications. Teams developing proprietary VLA models, shipping teleoperation products, or offering manipulation-as-a-service can train on this data without negotiating bespoke agreements or paying per-sample fees. The permissive license also allows derivative dataset creation, so organizations may augment, re-label, or blend these episodes with proprietary collections while retaining the right to deploy resulting models in commercial settings.

Apache-2.0 requires preservation of copyright notices and disclaimers but does not mandate open-sourcing trained weights or disclosing training recipes, offering flexibility for teams balancing open collaboration with competitive differentiation. Procurement and legal teams should verify that downstream model licenses remain compatible if the trained policy will be embedded in hardware sold under separate terms, though the dataset license itself imposes no reach-through obligations on model artifacts.

Procurement and integration notes

With over 607,000 downloads recorded on Hugging Face, IPEC-COMMUNITY/kuka_lerobot has seen significant adoption, indicating stable hosting and an active user base that can provide community support and integration examples. The dataset is accessible via the Hugging Face datasets library with a single API call, and the chunked Parquet structure supports lazy loading, so teams can prototype on subsets before committing to a full download. Each chunk is independently addressable, enabling parallel fetch operations and reducing time-to-first-batch when experimenting on distributed training infrastructure.

Organizations operating on-premises or in air-gapped environments should budget approximately 250 GB for the full dataset, including video files, and confirm that their data pipeline can parse LeRobot v2.0 metadata schemas. For teams sourcing through procurement platforms, verifying the dataset's continued availability on Hugging Face and establishing a local mirror or backup ensures continuity if upstream hosting policies change. The single-task focus means teams requiring multi-task generalization will need to combine this collection with complementary datasets, while those targeting KUKA iiwa deployments specifically will benefit from the robot-type homogeneity.

Known limitations and scope boundaries

Because the dataset documents a single task across all 209,880 episodes, diversity in manipulation skills, object geometries, and environmental conditions may be limited compared to multi-task collections like Open X-Embodiment or DROID. Teams training foundation models that generalize across object categories, grasp strategies, or task semantics should treat this dataset as a specialist source for KUKA iiwa kinematics rather than a broad-coverage manipulation corpus. The 10 fps capture rate is adequate for many tabletop tasks but may undersample fast contact-rich motions, so policies trained here may require temporal up-sampling or sim-to-real transfer techniques when deploying on higher-rate control loops.

No explicit quality-control metadata, success labels, or human annotations are surfaced in the schema, meaning teams will need to implement their own filtering heuristics if they wish to exclude failed trajectories or low-confidence episodes. The lack of pre-split validation and test sets requires users to carve out hold-out data manually, and the absence of language annotations or task descriptors limits applicability to vision-language-action architectures that rely on natural-language conditioning. Organizations requiring multi-modal grounding, semantic task labels, or fine-grained success metrics should plan to augment or re-annotate subsets before training.

Use these to move from category-level context into specific task, dataset, format, and comparison detail.

FAQ

What is the IPEC-COMMUNITY/kuka_lerobot dataset and who created it?

IPEC-COMMUNITY/kuka_lerobot is a large-scale robotics manipulation dataset containing 209,880 episodes captured from KUKA iiwa 7-DOF industrial arms, totaling approximately 2.46 million frames. Created using the LeRobot v2.0 framework and hosted on Hugging Face, the dataset provides synchronized video observations and state-action trajectories stored in Parquet format, designed for training imitation learning policies, vision-language-action models, and manipulation-focused world models. The IPEC-COMMUNITY organization published the dataset under the Apache-2.0 license, enabling both academic and commercial use without royalty obligations. While the dataset description does not detail the specific task or collection methodology, the scale and standardized format indicate it was produced through systematic teleoperation or scripted data-collection campaigns on KUKA hardware.

What license governs this dataset and can I use it commercially?

The dataset is released under the Apache-2.0 license, one of the most permissive open-source licenses available. This means you may use, modify, and distribute the data commercially without paying fees, obtaining special permissions, or open-sourcing your trained models. Apache-2.0 requires only that you retain copyright notices and include a copy of the license with any redistribution, but it does not impose copyleft obligations or restrict proprietary applications. For robotics teams shipping products, offering manipulation services, or deploying VLA models in commercial settings, Apache-2.0 provides the legal clarity needed to integrate this dataset into proprietary training pipelines. Ensure your organization's legal review confirms compatibility with any additional licenses governing model weights or downstream software components.

Which robotics teams should prioritize using this dataset?

Teams developing manipulation policies for KUKA iiwa arms or 7-DOF manipulators with similar kinematics will benefit most, as the entire 209,880-episode collection focuses on this single robot type. The dataset is particularly valuable for pre-training vision-language-action models that require large-scale trajectory data, fine-tuning imitation learning policies on industrial arms, or building world models that predict contact-rich manipulation outcomes. Organizations with existing KUKA deployments can use this data to bootstrap teleoperation systems or train adaptive controllers without collecting proprietary episodes from scratch. Research labs investigating sample efficiency, sim-to-real transfer, or multi-task generalization can leverage the dataset's scale as a baseline corpus, though the single-task scope means it should be combined with more diverse collections for broader coverage. Teams operating in regulated industries may appreciate the permissive license, which simplifies compliance compared to datasets with academic-use-only or share-alike restrictions.

When is this dataset NOT the right choice for my project?

If your project requires multi-task diversity, semantic task labels, or language-conditioned policies, IPEC-COMMUNITY/kuka_lerobot's single-task focus and lack of natural-language annotations make it a poor standalone choice. The dataset does not include explicit success labels, quality scores, or failure annotations, so teams needing curated positive examples or reward signals must implement their own filtering or re-labeling workflows. The 10 fps frame rate may also be insufficient for high-speed manipulation, dynamic contact modeling, or control loops running above 20 Hz. Organizations deploying on non-KUKA platforms or arms with fewer than seven degrees of freedom should consider whether the kinematic mismatch will require expensive domain adaptation, as the trajectories and state representations are tightly coupled to the iiwa's joint configuration. For foundation model builders seeking maximum task and embodiment diversity, datasets like Open X-Embodiment or DROID offer broader coverage at the cost of smaller per-robot sample counts.

How do I download and integrate this dataset into my training pipeline?

The dataset is hosted on Hugging Face and can be fetched using the datasets library with a single line of Python code referencing the identifier IPEC-COMMUNITY/kuka_lerobot. Because the data is chunked into 210 segments of approximately 1,000 episodes each, you can stream subsets for prototyping or download the full collection in parallel to minimize wall-clock time. Each episode is stored as a Parquet file with a corresponding video, following the LeRobot v2.0 schema, so standard dataloader patterns for trajectory data apply. Once downloaded, confirm your pipeline can parse the metadata in meta/info.json, which specifies frame counts, fps, splits, and path templates. For distributed training, leverage the chunk structure to shard episodes across workers, and consider caching the video frames in a preprocessed format if repeated decoding becomes a bottleneck. Teams using PyTorch or TensorFlow can adapt existing LeRobot or OpenX data loaders with minimal modification.

What are the main technical limitations I should plan around?

The dataset's single-task composition limits generalization to novel manipulation scenarios, so teams building multi-task policies must augment with additional data sources or accept reduced zero-shot performance on out-of-distribution objects and goals. The absence of pre-split validation and test sets means you will need to manually partition episodes, and without published success labels, establishing ground-truth performance benchmarks requires either heuristic filtering or human review of a sample subset. The 10 fps capture rate may undersample fast dynamics, requiring temporal interpolation or hybrid approaches that blend low-frequency visual observations with higher-rate proprioceptive feedback. Finally, the lack of language annotations or semantic task descriptors prevents direct use in vision-language-action architectures that condition on natural-language instructions, so teams pursuing that paradigm should budget for annotation or pair the dataset with a language-labeled collection.

Need data like IPEC-COMMUNITY/kuka_lerobot: Large-Scale KUKA iiwa Manipulation Dataset?

If your project needs similar modality, scale, or licensing, truelabel can surface comparable open datasets or match you with capture partners that deliver to spec.

Access dataset on Hugging Face