truelabelRequest dataEarnRequest

Physical AI Model Profile

GR-2: ByteDance's Video-Language-Action Model for Robot Manipulation

GR-2 is a generative video-language-action model from ByteDance Research that first generates future video frames from a language instruction, then infers the robot action trajectory from that predicted video. It is pretrained on 38 million Internet video clips totaling over 50 billion tokens, then fine-tuned on roughly 40,000 teleoperation trajectories spanning 105 table-top tasks across 8 skills. On its own real-robot multi-task benchmark GR-2 reaches a 97.7% average success rate across those 105 tasks, and on the simulated CALVIN long-horizon benchmark it completes a chain of five instructions in a row 85.9% of the time. Unlike vision-language-action models that compress each observation into a fixed embedding, GR-2 encodes video frames as discrete VQGAN tokens, letting large-scale video pretraining and the manipulation policy share one representation.

Updated 2026-07-1413 min read
By Truelabel Team
Reviewed by Truelabel Team ·
GR-2 robot model

Quick facts

Topic
GR2
Audience
Procurement leads, ML ops, robotics engineers
Deliverable
Buyer-facing reference + procurement guidance

What GR-2 is and why the video-first design matters

GR-2 is a generative transformer from ByteDance Research that predicts video frames and robot actions in one autoregressive stream[1]. It does not split perception from control. Each camera frame becomes a short run of discrete VQGAN tokens, and the 6-DoF action deltas that drive the arm are interleaved with those tokens in the same sequence. From there the model does what a language model does: predict the next token, whether that token is a patch of future video or a slice of motion.

That interleaving is the whole argument for GR-2. Older vision-language-action models such as RT-2 and OpenVLA compress each observation into a fixed embedding before a separate head emits an action, so visual pretraining and the control policy never share a representation. GR-2 keeps both in one token space, which lets it pretrain on raw human video at scale and carry that prior straight into manipulation. Whether the prior is worth the cost of collecting video is the real procurement question, and the benchmark numbers below answer it.

The design dictates the data. GR-2's own robot setup pairs a 7-DoF Kinova Gen3 arm and a Robotiq 2F-85 gripper with two synchronized cameras, a static head view over the workspace and a wrist camera, alongside 6-DoF end-effector trajectories, a binary gripper state, and a natural-language task string, all at resolutions the tokenizer can encode. Miss the calibration or the timestamp alignment and the interleaved sequence pairs the wrong action with the wrong frame, which is why synchronized teleoperation is harder to source than plain video.

Architecture: VQGAN video tokens and generate-then-act

Start with how GR-2 sees. A frozen VQGAN, trained on Internet images plus in-domain robot footage, encodes each camera frame into a grid of discrete tokens drawn from a fixed codebook, and a frozen text encoder turns the instruction into language tokens. The robot's own state and its end-effector actions enter the same sequence, so vision, language, state, and motion share one token space. Because the frame tokens dominate that sequence, frame resolution and camera count are the levers that decide whether training fits in memory.

The backbone is a causal transformer that GR-2 trains at four sizes: 30M, 95M, 312M, and 719M trainable parameters, with the default model carrying 230M total parameters of which 95M are trainable. Two objectives run together, one for generating future video frames and one for predicting the action trajectory, and video-generation validation loss falls cleanly as the model grows, which is the paper's evidence that the recipe scales. Language conditions generation through the frozen text encoder rather than being trained end to end, and the robot state is fed in alongside the video tokens.

The reason for discrete video tokens is reuse. Discrete codes let GR-2 borrow language-model machinery directly, the same next-token prediction that holds up at scale, and the fixed codebook acts as a bottleneck that pushes the model toward object-level structure instead of pixel memorization. The cost is that VQGAN is lossy, so anything the codebook cannot represent, fine textures, thin cables, the policy cannot act on either.

What ties it together is that GR-2 generates video first and then infers actions from that predicted visual trajectory, so the model effectively plans in image space before it moves. On the real system, its predicted trajectory is handed to a low-level controller that runs at 200 Hz and folds in collision and manipulability constraints. Because the policy outputs end-effector motion rather than joint commands, the same checkpoint can in principle drive a different arm given an IK solver and the right normalization statistics, the kind of separation RT-1 and later manipulation stacks also rely on.

Training data: 38M pretraining videos, then ~40K robot trajectories

GR-2's first stage never sees a robot. It runs next-token video prediction over 38 million Internet clips, more than 50 billion VQGAN tokens, spanning everyday human activity across households, workplaces, and outdoor scenes[1]. The objective is deliberately simple: given a text description and a starting frame, predict the frames that follow. Learning to continue human video at this scale is what gives GR-2 its priors for temporal dynamics and hand-object contact, the priors that later transfer into manipulation.

Stage two fine-tunes on the robot. ByteDance collected about 40,000 teleoperation trajectories across 105 table-top tasks covering eight skills, picking, placing, uncapping, capping, opening, closing, pressing, and pouring, at roughly 400 demonstrations per task[1]. The same run trains both objectives, video generation and action prediction, so the model keeps its visual-prediction ability while learning to act. Data augmentation, inserting new objects and swapping backgrounds during fine-tuning, is what lifts generalization to scenes the robot never saw.

The scarcity result is the number to keep in mind when scoping a data budget. Trained on roughly one-eighth of the set, about 50 demonstrations per task, GR-2 still reaches 73.9% average success on the multi-task benchmark and beats GR-1 in every generalization setting[1]. Task breadth and the video-pretraining prior carry more weight than raw trajectory count, because the video stage already supplies general visual reasoning and the robot data mainly has to teach contact and force. That asymmetry is the reverse of RT-2, which leans on Internet-scale vision-language pretraining and pairs it with far less bespoke robot data per task.

  1. 01

    Preserve the motion

    Sample at rates that keep frame-to-frame motion intact; heavily subsampled classification clips lose the dynamics a video model learns from.

  2. 02

    Keep hands on objects

    Favor clips where hands or end-effectors are actually manipulating objects; contact-rich footage transfers to robot manipulation far better than passive scenes.

  3. 03

    Spread objects and tasks

    Cover a wide range of objects and task types so the model learns manipulation primitives instead of memorizing a handful of shapes.

  4. 04

    Deliver at tokenizer resolution

    Ship frames at the tokenizer's native resolution so preprocessing does not bake rescaling artifacts into the discrete codes.

Benchmarks: 97.7% real-robot multi-task, 85.9% on CALVIN

The headline is the real-robot multi-task result: across 105 table-top tasks in the standard setting, GR-2 averages 97.7% success[1]. That is a single-policy number, one model handling every task from a language prompt, not a per-task best. Generalization is where the video prior shows: GR-2 reaches 71.4% on unseen backgrounds and 71.7% in unseen environments, roughly double GR-1, and with data augmentation climbs to 87.0% in unseen environments and 74.7% averaged across all three generalization settings. Unseen manipulation, novel objects and shapes, is hardest at 55.8%.

CALVIN is the simulated long-horizon test, the ABCD-D split of 34 tasks with over 20,000 demonstrations, scored by chaining five instructions in a row across 1,000 sequences[1]. GR-2 sets a new state of the art there: 98.6% on the first task and 85.9% on all five in a row, lifting the average completed-sequence length from GR-1's 4.21 to 4.64. Because one failure ends the chain, the five-in-a-row figure is a genuine long-horizon result rather than a per-step average, and it beats the RT-1, MT-ACT, HULC, RoboFlamingo, and GR-1 baselines the paper compares against.

Read the two benchmarks as measuring different things. The 97.7% is a real Kinova arm under Simple conditions; CALVIN runs on clean simulated physics with full state observability, so its numbers flatter every policy on it, GR-2 included. The honest signal for a deployment is the spread between in-distribution strength and out-of-distribution scores, 97.7% in the Simple setting against 55.8% on unseen manipulation. Budget around the out-of-distribution numbers and the failure modes the paper actually names, failing to grasp unfamiliar shapes and selecting the wrong object when the target is novel, not around the in-distribution best case.

GR-2 vs RT-2 vs OpenVLA

The three models split on what the vision front end produces and how much bespoke robot data the policy then needs. GR-2 tokenizes video into discrete VQGAN codes and generates video and actions autoregressively. RT-2 co-fine-tunes a large vision-language model on robot data and emits actions as discretized tokens in the model's own vocabulary. OpenVLA follows the same discretized-action-token approach on an open vision-language backbone, and both RT-2 and OpenVLA pretrain on Web image-text rather than video, which is the design choice GR-2 is arguing against.

The data profiles differ as much as the architectures. GR-2 fine-tunes on about 40,000 teleoperation trajectories across 105 tasks[1] and puts most of its scale into 38 million pretraining videos. OpenVLA is trained on roughly 970,000 robot trajectories drawn from the Open X-Embodiment collection[2][3], while RT-2 is built by co-fine-tuning its vision-language backbone on robot demonstrations from the RT-1 program[4][5]. GR-2 is the outlier in trading robot episodes for Web-scale video, the more expensive modality per useful sample, but the one that supplies the temporal prior the other two get from static image-text.

So the choice is task-shaped. Long-horizon, multi-step, temporally coupled work is where GR-2's video pretraining earns its collection cost. Single-step or knowledge-heavy tasks play to RT-2's semantic breadth. Broad cross-embodiment generalization is OpenVLA's pitch, at the price of the largest robot dataset of the three.

GR-2RT-2OpenVLA
Vision front endDiscrete VQGAN video tokensVision-language model (continuous)Vision-language model (continuous)
Action outputGenerated video + end-effector trajectoryDiscretized action tokensDiscretized action tokens
Pretraining signal38M Internet video clipsWeb image-text (VLM)Web image-text (VLM)
Robot fine-tuning data~40,000 trajectories (105 tasks)RT-1 robot demonstrations~970,000 trajectories (Open X-Embodiment)
GR-2, RT-2, and OpenVLA on the axes that shape sourcing cost

Sourcing and integrating GR-2 training data

Most GR-2-style pipelines store data as RLDS or HDF5, and the choice is real. RLDS gives you batching, shuffling, and prefetching for free across TensorFlow, PyTorch, and JAX, which matters once a fine-tuning set runs to tens of thousands of trajectories. HDF5 is more flexible for a custom schema but leaves batching to you. Either way, keep raw frames on disk and run VQGAN encoding on the fly rather than caching tokens, which keeps storage down, and LeRobot's tooling makes the on-the-fly path cheap enough to keep inline.

The hard part of sourcing is not raw volume, it is synchronization. Frames, joint states, and gripper signals must share a clock tight enough that the training sequence lines up the right action with the right observation, and a few milliseconds of drift quietly corrupts the policy. That takes hardware-triggered capture and per-trajectory calibration, not a folder of MP4s. Even GR-2's comparatively lean ~40,000-trajectory fine-tuning set is dozens of robot-weeks of teleoperation before cleaning and formatting, which is why many teams buy the fine-tuning data rather than build it.

Truelabel's physical AI marketplace is built for that spec: multi-view teleoperation and egocentric video from around 10,000 consented collectors across 100 countries, delivered in RLDS or LeRobot format to S3, GCS, or Azure with camera calibration, temporal alignment, and per-trajectory provenance. Sample packets ship before scale so the calibration and QA gate run on a real batch first. For pretraining, the same network captures egocentric manipulation clips to a task distribution instead of reselling research-licensed public sets that cannot train a deployed model. Truelabel's data provenance system records collection context and consent per episode, so a buyer can audit quality and trace a bad batch to its source.

Use these to move from category-level context into specific task, dataset, format, and comparison detail.

External references and source context

  1. GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

    GR-2 technical report (ByteDance, arXiv 2410.06158): 38M pretraining video clips, over 50B tokens; ~40,000 teleoperation trajectories across 105 tasks (8 skills, ~400/task); 97.7% average real-robot multi-task success; CALVIN ABCD-D 85.9% five-task and 98.6% one-task; 230M-719M parameter sizes; Kinova Gen3 + Robotiq 2F-85, 200 Hz controller; frozen VQGAN + text encoder

    arXiv ↩
  2. OpenVLA: An Open-Source Vision-Language-Action Model

    OpenVLA open-source vision-language-action model and Open X-Embodiment dataset

    arXiv ↩
  3. Open X-Embodiment: Robotic Learning Datasets and RT-X Models

    Open X-Embodiment dataset: 970,000 robot episodes across multiple platforms

    arXiv ↩
  4. RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

    RT-2 vision-language-action model architecture and performance benchmarks

    arXiv ↩
  5. RT-1: Robotics Transformer for Real-World Control at Scale

    RT-1 robotics transformer architecture and real-world deployment

    arXiv ↩
  6. Dataset page

    Something-Something-v2 dataset: 220,000 clips of human hand-object interactions

    developer.qualcomm.com
  7. EPIC-KITCHENS-100 dataset page

    EPIC-KITCHENS-100 dataset: 700 hours of egocentric kitchen activities

    epic-kitchens.github.io
  8. Kinetics dataset repository

    Kinetics-400 dataset: 300,000 clips of human actions across 400 categories

    GitHub
  9. Diffusion Policy training example

    Diffusion policy training and high-frequency action upsampling

    GitHub
  10. Scale AI: Expanding Our Data Engine for Physical AI

    Scale AI physical AI data engine for autonomous data collection

    scale.com
  11. NVIDIA Cosmos World Foundation Models

    NVIDIA Cosmos multi-modal world foundation models with depth and tactile

    NVIDIA Developer
  12. RoboNet: Large-Scale Multi-Robot Learning

    RoboNet large-scale multi-robot learning dataset

    arXiv
  13. CALVIN paper

    CALVIN benchmark for long-horizon manipulation tasks

    arXiv
  14. BridgeData V2: A Dataset for Robot Learning at Scale

    BridgeData V2 large-scale robot learning dataset

    arXiv
  15. DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset

    DROID large-scale in-the-wild robot manipulation dataset

    arXiv
  16. RoboCat: A Self-Improving Generalist Agent for Robotic Manipulation

    RoboCat self-improving generalist agent for robotic manipulation

    arXiv
  17. Do As I Can, Not As I Say: Grounding Language in Robotic Affordances

    SayCan language grounding in robotic affordances

    arXiv
  18. Introduction to HDF5

    HDF5 hierarchical data format for robot manipulation datasets

    The HDF Group
  19. MCAP guides

    MCAP container format for multi-modal sensor data

    MCAP
  20. labelbox

    Labelbox data annotation platform for computer vision

    labelbox.com
  21. scale.com scale ai universal robots physical ai

    Scale AI partnership with Universal Robots for physical AI data

    scale.com
  22. Encord Series C announcement

    Encord $60M Series C for computer vision annotation platform

    encord.com
  23. kognic.com platform

    Kognic annotation platform for autonomous systems and robotics

    kognic.com
  24. segments.ai the 8 best point cloud labeling tools

    Point cloud labeling tools comparison for 3D robotics data

    segments.ai

FAQ

What is the difference between GR-2 and RT-2 for robot manipulation?

GR-2 uses discrete VQGAN video tokenization and generates video and actions autoregressively, while RT-2 co-fine-tunes a vision-language model and emits actions as discretized tokens. GR-2 pretrains on 38 million Internet videos to capture temporal dynamics, whereas RT-2 pretrains on Web image-text. GR-2 is fine-tuned on about 40,000 teleoperation trajectories across 105 tasks and reaches 97.7% average success on its real-robot multi-task benchmark; on the simulated CALVIN five-task benchmark it completes all five instructions in a row 85.9% of the time. RT-2's strength is semantic generalization from Web-scale vision-language pretraining with comparatively little bespoke robot data.

How much training data does GR-2 require for fine-tuning?

GR-2 was fine-tuned on about 40,000 teleoperation trajectories spanning 105 table-top tasks and 8 skills, roughly 400 demonstrations per task, on top of pretraining on 38 million Internet video clips (over 50 billion tokens). Each trajectory pairs multi-view RGB from a static head camera and a wrist camera, 6-DoF end-effector motion, gripper state, and a language instruction. In a data-scarcity test using only about 50 demonstrations per task, GR-2 still reached 73.9% average success on the multi-task benchmark, evidence that the video-pretraining prior does much of the heavy lifting.

What video formats and resolutions does GR-2 require?

GR-2 takes multi-view RGB video, which a frozen VQGAN converts into discrete tokens, together with the language instruction and the robot's state. Training data is typically stored in RLDS or HDF5, with tokenization done offline or on the fly. The practical requirements are tight temporal synchronization across cameras and sensors and accurate camera calibration, so each action lines up with the right frame in the sequence. Truelabel delivers teleoperation data with verified camera calibration, temporal alignment, and per-trajectory provenance.

What robot does GR-2 run on, and can it move to another arm?

ByteDance's GR-2 experiments run on a 7-DoF Kinova Gen3 arm with a Robotiq 2F-85 gripper and two cameras, a static head view and a wrist camera, with a low-level controller executing the predicted trajectory at 200 Hz under collision and manipulability constraints. Because GR-2 outputs end-effector motion rather than joint-space commands, the same policy can in principle drive a different arm given an inverse-kinematics solver and robot-specific action normalization; moving to another arm means matching camera viewpoints to the training setup and re-deriving those normalization statistics.

What are the main failure modes of GR-2 in real-world deployment?

In GR-2's own evaluation, the hardest setting is unseen manipulation, where success falls to 55.8% and the typical failures are failing to grasp unfamiliar object shapes and selecting the wrong object when instructed to pick a novel one. In-distribution multi-task performance is much higher at 97.7%, and generalization to unseen backgrounds and environments lands in between, about 71% each, or up to 87% in unseen environments with data augmentation. The paper does not publish a fine-grained breakdown of collision versus slip versus timeout rates, so plan around the gap between in-distribution and out-of-distribution success rather than a fixed failure taxonomy.

How does GR-2's video pretraining improve manipulation performance?

GR-2's pretraining on 38 million Internet videos teaches temporal dynamics and hand-object interaction that transfer to manipulation, and the paper frames video generation as an implicit planner: the model predicts the future visual trajectory, then infers the actions to realize it. The clearest evidence of the prior's value is data efficiency and generalization, GR-2 still reaches 73.9% average multi-task success with only about 50 demonstrations per task, and roughly doubles GR-1's success in unseen backgrounds and environments. The benefit is largest on long-horizon and out-of-distribution settings.

Looking for GR-2 robot model?

Specify modality, task, environment, requested rights posture, and delivery format. Truelabel routes the request to candidate capture partners and helps scope consent/provenance artifacts and commercial licensing requirements for buyer review before delivery.

Source GR-2 Training Data