Physical AI Implementation Guide
How to Fine-Tune a VLA Model with a Versioned Data Mixture
Fine-tune a VLA from a versioned robot data mixture, not an untracked folder of episodes. Pin the named config and component weights, preserve unknown dataset versions and control schemas as unknown, run automated structural checks separately from human task and rights review, then compare mixture ablations on a held-out target embodiment.
Quick facts
- Topic
- HOW TO Fine Tune A VLA Model
- Audience
- Procurement leads, ML ops, robotics engineers
- Deliverable
- Operational playbook with sample workflow + accept-rule criteria
VLA data mixtures: which slice your recipe is missing
A VLA data mixture is the versioned set of datasets and per-dataset sampling weights used for a run. The recipe below is copied from OpenVLA’s named OXE config at one reviewed commit. Its values are the config’s own relative sampling weights, not normalized probabilities or a promised optimum; versions, filters, normalization, action schema/rate, success/failure composition, splits, and rights stay explicitly unknown when the source does not report them. Octo and Open X-Embodiment provide robotics-primary context, but they do not fill those gaps.
| Field | Value |
|---|---|
| id | MIXTURE-OPENVLA-OXE-MAGIC-SOUP-PLUS |
| model_recipe_id | openvla-7b / oxe_magic_soup_plus |
| source_reported_version | OpenVLA commit c8f03f48af69 |
| normalized_field | named mixture |
| normalized_value | oxe_magic_soup_plus |
| unit | relative sampling weight |
| primary_source_url | https://github.com/openvla/openvla/blob/c8f03f48af69/prismatic/vla/datasets/rlds/oxe/mixtures.py |
| source_type | project |
| source_id | project-github-com-openvla-openvla-mixtures |
| exact_locator | OXE_NAMED_MIXTURES['oxe_magic_soup_plus'] |
| checked_date | 2026-07-22 |
| retrieval_hash | git-commit:c8f03f48af69 |
| confidence | high |
| status | human |
| evidence_basis | source-reported |
| filter | Active tuple entries only; commented broken or omitted entries are excluded by the source config. |
| normalization | unknown — mixtures.py points to separate dataset transforms/configs and does not define one shared normalization rule |
| action_schema_rate | unknown — heterogeneous component action schemas and control rates are not reported in mixtures.py |
| success_failure_ratio | unknown — mixtures.py does not report success/failure composition |
| train_eval_separation | unknown in source config — define a target-embodiment holdout before training |
| eval_split | not specified by mixtures.py — project-specific held-out split required |
| license_compatibility | unknown — code availability does not establish compatibility of every dataset, consent term, or derived-model right |
| limitation | source-reported, not independently validated; weights are relative sampling weights, not normalized probabilities or a universally optimal recipe. |
OpenVLA OXE mixture components and relative weights
Use these source-reported values as a reproducible starting config. Before training, pin every component version and normalization transform in your own manifest; the source mixture file does not do that for most components.
| Recipe / model ID | Dataset | Version | Weight | Unit | Filter | Normalization | Action schema / rate | Success / failure ratio | Train / eval separation | License compatibility | Config locator | Source (type) | Checked | Confidence | Status |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| openvla-7b / oxe_magic_soup_plus | fractal20220817_data | unknown — mixtures.py does not pin a component dataset version | 0.54087122203 | relative sampling weight | Active tuple entries only; commented broken or omitted entries are excluded by the source config. | unknown — mixtures.py points to separate dataset transforms/configs and does not define one shared normalization rule | unknown — heterogeneous component action schemas and control rates are not reported in mixtures.py | unknown — mixtures.py does not report success/failure composition | unknown in source config — define a target-embodiment holdout before training | unknown — code availability does not establish compatibility of every dataset, consent term, or derived-model right | OXE_NAMED_MIXTURES['oxe_magic_soup_plus'] | https://github.com/openvla/openvla/blob/c8f03f48af69/prismatic/vla/datasets/rlds/oxe/mixtures.py (project) | 2026-07-22 | high | human |
| openvla-7b / oxe_magic_soup_plus | kuka | unknown — mixtures.py does not pin a component dataset version | 0.8341046294 | relative sampling weight | Active tuple entries only; commented broken or omitted entries are excluded by the source config. | unknown — mixtures.py points to separate dataset transforms/configs and does not define one shared normalization rule | unknown — heterogeneous component action schemas and control rates are not reported in mixtures.py | unknown — mixtures.py does not report success/failure composition | unknown in source config — define a target-embodiment holdout before training | unknown — code availability does not establish compatibility of every dataset, consent term, or derived-model right | OXE_NAMED_MIXTURES['oxe_magic_soup_plus'] | https://github.com/openvla/openvla/blob/c8f03f48af69/prismatic/vla/datasets/rlds/oxe/mixtures.py (project) | 2026-07-22 | high | human |
| openvla-7b / oxe_magic_soup_plus | bridge_orig | unknown — mixtures.py does not pin a component dataset version | 1.0 | relative sampling weight | Active tuple entries only; commented broken or omitted entries are excluded by the source config. | unknown — mixtures.py points to separate dataset transforms/configs and does not define one shared normalization rule | unknown — heterogeneous component action schemas and control rates are not reported in mixtures.py | unknown — mixtures.py does not report success/failure composition | unknown in source config — define a target-embodiment holdout before training | unknown — code availability does not establish compatibility of every dataset, consent term, or derived-model right | OXE_NAMED_MIXTURES['oxe_magic_soup_plus'] | https://github.com/openvla/openvla/blob/c8f03f48af69/prismatic/vla/datasets/rlds/oxe/mixtures.py (project) | 2026-07-22 | high | human |
| openvla-7b / oxe_magic_soup_plus | taco_play | unknown — mixtures.py does not pin a component dataset version | 2.0 | relative sampling weight | Active tuple entries only; commented broken or omitted entries are excluded by the source config. | unknown — mixtures.py points to separate dataset transforms/configs and does not define one shared normalization rule | unknown — heterogeneous component action schemas and control rates are not reported in mixtures.py | unknown — mixtures.py does not report success/failure composition | unknown in source config — define a target-embodiment holdout before training | unknown — code availability does not establish compatibility of every dataset, consent term, or derived-model right | OXE_NAMED_MIXTURES['oxe_magic_soup_plus'] | https://github.com/openvla/openvla/blob/c8f03f48af69/prismatic/vla/datasets/rlds/oxe/mixtures.py (project) | 2026-07-22 | high | human |
| openvla-7b / oxe_magic_soup_plus | jaco_play | unknown — mixtures.py does not pin a component dataset version | 1.0 | relative sampling weight | Active tuple entries only; commented broken or omitted entries are excluded by the source config. | unknown — mixtures.py points to separate dataset transforms/configs and does not define one shared normalization rule | unknown — heterogeneous component action schemas and control rates are not reported in mixtures.py | unknown — mixtures.py does not report success/failure composition | unknown in source config — define a target-embodiment holdout before training | unknown — code availability does not establish compatibility of every dataset, consent term, or derived-model right | OXE_NAMED_MIXTURES['oxe_magic_soup_plus'] | https://github.com/openvla/openvla/blob/c8f03f48af69/prismatic/vla/datasets/rlds/oxe/mixtures.py (project) | 2026-07-22 | high | human |
| openvla-7b / oxe_magic_soup_plus | berkeley_cable_routing | unknown — mixtures.py does not pin a component dataset version | 1.0 | relative sampling weight | Active tuple entries only; commented broken or omitted entries are excluded by the source config. | unknown — mixtures.py points to separate dataset transforms/configs and does not define one shared normalization rule | unknown — heterogeneous component action schemas and control rates are not reported in mixtures.py | unknown — mixtures.py does not report success/failure composition | unknown in source config — define a target-embodiment holdout before training | unknown — code availability does not establish compatibility of every dataset, consent term, or derived-model right | OXE_NAMED_MIXTURES['oxe_magic_soup_plus'] | https://github.com/openvla/openvla/blob/c8f03f48af69/prismatic/vla/datasets/rlds/oxe/mixtures.py (project) | 2026-07-22 | high | human |
| openvla-7b / oxe_magic_soup_plus | roboturk | unknown — mixtures.py does not pin a component dataset version | 2.0 | relative sampling weight | Active tuple entries only; commented broken or omitted entries are excluded by the source config. | unknown — mixtures.py points to separate dataset transforms/configs and does not define one shared normalization rule | unknown — heterogeneous component action schemas and control rates are not reported in mixtures.py | unknown — mixtures.py does not report success/failure composition | unknown in source config — define a target-embodiment holdout before training | unknown — code availability does not establish compatibility of every dataset, consent term, or derived-model right | OXE_NAMED_MIXTURES['oxe_magic_soup_plus'] | https://github.com/openvla/openvla/blob/c8f03f48af69/prismatic/vla/datasets/rlds/oxe/mixtures.py (project) | 2026-07-22 | high | human |
| openvla-7b / oxe_magic_soup_plus | viola | unknown — mixtures.py does not pin a component dataset version | 2.0 | relative sampling weight | Active tuple entries only; commented broken or omitted entries are excluded by the source config. | unknown — mixtures.py points to separate dataset transforms/configs and does not define one shared normalization rule | unknown — heterogeneous component action schemas and control rates are not reported in mixtures.py | unknown — mixtures.py does not report success/failure composition | unknown in source config — define a target-embodiment holdout before training | unknown — code availability does not establish compatibility of every dataset, consent term, or derived-model right | OXE_NAMED_MIXTURES['oxe_magic_soup_plus'] | https://github.com/openvla/openvla/blob/c8f03f48af69/prismatic/vla/datasets/rlds/oxe/mixtures.py (project) | 2026-07-22 | high | human |
| openvla-7b / oxe_magic_soup_plus | berkeley_autolab_ur5 | unknown — mixtures.py does not pin a component dataset version | 2.0 | relative sampling weight | Active tuple entries only; commented broken or omitted entries are excluded by the source config. | unknown — mixtures.py points to separate dataset transforms/configs and does not define one shared normalization rule | unknown — heterogeneous component action schemas and control rates are not reported in mixtures.py | unknown — mixtures.py does not report success/failure composition | unknown in source config — define a target-embodiment holdout before training | unknown — code availability does not establish compatibility of every dataset, consent term, or derived-model right | OXE_NAMED_MIXTURES['oxe_magic_soup_plus'] | https://github.com/openvla/openvla/blob/c8f03f48af69/prismatic/vla/datasets/rlds/oxe/mixtures.py (project) | 2026-07-22 | high | human |
| openvla-7b / oxe_magic_soup_plus | toto | unknown — mixtures.py does not pin a component dataset version | 1.0 | relative sampling weight | Active tuple entries only; commented broken or omitted entries are excluded by the source config. | unknown — mixtures.py points to separate dataset transforms/configs and does not define one shared normalization rule | unknown — heterogeneous component action schemas and control rates are not reported in mixtures.py | unknown — mixtures.py does not report success/failure composition | unknown in source config — define a target-embodiment holdout before training | unknown — code availability does not establish compatibility of every dataset, consent term, or derived-model right | OXE_NAMED_MIXTURES['oxe_magic_soup_plus'] | https://github.com/openvla/openvla/blob/c8f03f48af69/prismatic/vla/datasets/rlds/oxe/mixtures.py (project) | 2026-07-22 | high | human |
| openvla-7b / oxe_magic_soup_plus | language_table | unknown — mixtures.py does not pin a component dataset version | 0.1 | relative sampling weight | Active tuple entries only; commented broken or omitted entries are excluded by the source config. | unknown — mixtures.py points to separate dataset transforms/configs and does not define one shared normalization rule | unknown — heterogeneous component action schemas and control rates are not reported in mixtures.py | unknown — mixtures.py does not report success/failure composition | unknown in source config — define a target-embodiment holdout before training | unknown — code availability does not establish compatibility of every dataset, consent term, or derived-model right | OXE_NAMED_MIXTURES['oxe_magic_soup_plus'] | https://github.com/openvla/openvla/blob/c8f03f48af69/prismatic/vla/datasets/rlds/oxe/mixtures.py (project) | 2026-07-22 | high | human |
| openvla-7b / oxe_magic_soup_plus | stanford_hydra_dataset_converted_externally_to_rlds | unknown — mixtures.py does not pin a component dataset version | 2.0 | relative sampling weight | Active tuple entries only; commented broken or omitted entries are excluded by the source config. | unknown — mixtures.py points to separate dataset transforms/configs and does not define one shared normalization rule | unknown — heterogeneous component action schemas and control rates are not reported in mixtures.py | unknown — mixtures.py does not report success/failure composition | unknown in source config — define a target-embodiment holdout before training | unknown — code availability does not establish compatibility of every dataset, consent term, or derived-model right | OXE_NAMED_MIXTURES['oxe_magic_soup_plus'] | https://github.com/openvla/openvla/blob/c8f03f48af69/prismatic/vla/datasets/rlds/oxe/mixtures.py (project) | 2026-07-22 | high | human |
| openvla-7b / oxe_magic_soup_plus | austin_buds_dataset_converted_externally_to_rlds | unknown — mixtures.py does not pin a component dataset version | 1.0 | relative sampling weight | Active tuple entries only; commented broken or omitted entries are excluded by the source config. | unknown — mixtures.py points to separate dataset transforms/configs and does not define one shared normalization rule | unknown — heterogeneous component action schemas and control rates are not reported in mixtures.py | unknown — mixtures.py does not report success/failure composition | unknown in source config — define a target-embodiment holdout before training | unknown — code availability does not establish compatibility of every dataset, consent term, or derived-model right | OXE_NAMED_MIXTURES['oxe_magic_soup_plus'] | https://github.com/openvla/openvla/blob/c8f03f48af69/prismatic/vla/datasets/rlds/oxe/mixtures.py (project) | 2026-07-22 | high | human |
| openvla-7b / oxe_magic_soup_plus | nyu_franka_play_dataset_converted_externally_to_rlds | unknown — mixtures.py does not pin a component dataset version | 3.0 | relative sampling weight | Active tuple entries only; commented broken or omitted entries are excluded by the source config. | unknown — mixtures.py points to separate dataset transforms/configs and does not define one shared normalization rule | unknown — heterogeneous component action schemas and control rates are not reported in mixtures.py | unknown — mixtures.py does not report success/failure composition | unknown in source config — define a target-embodiment holdout before training | unknown — code availability does not establish compatibility of every dataset, consent term, or derived-model right | OXE_NAMED_MIXTURES['oxe_magic_soup_plus'] | https://github.com/openvla/openvla/blob/c8f03f48af69/prismatic/vla/datasets/rlds/oxe/mixtures.py (project) | 2026-07-22 | high | human |
| openvla-7b / oxe_magic_soup_plus | furniture_bench_dataset_converted_externally_to_rlds | unknown — mixtures.py does not pin a component dataset version | 0.1 | relative sampling weight | Active tuple entries only; commented broken or omitted entries are excluded by the source config. | unknown — mixtures.py points to separate dataset transforms/configs and does not define one shared normalization rule | unknown — heterogeneous component action schemas and control rates are not reported in mixtures.py | unknown — mixtures.py does not report success/failure composition | unknown in source config — define a target-embodiment holdout before training | unknown — code availability does not establish compatibility of every dataset, consent term, or derived-model right | OXE_NAMED_MIXTURES['oxe_magic_soup_plus'] | https://github.com/openvla/openvla/blob/c8f03f48af69/prismatic/vla/datasets/rlds/oxe/mixtures.py (project) | 2026-07-22 | high | human |
| openvla-7b / oxe_magic_soup_plus | ucsd_kitchen_dataset_converted_externally_to_rlds | unknown — mixtures.py does not pin a component dataset version | 2.0 | relative sampling weight | Active tuple entries only; commented broken or omitted entries are excluded by the source config. | unknown — mixtures.py points to separate dataset transforms/configs and does not define one shared normalization rule | unknown — heterogeneous component action schemas and control rates are not reported in mixtures.py | unknown — mixtures.py does not report success/failure composition | unknown in source config — define a target-embodiment holdout before training | unknown — code availability does not establish compatibility of every dataset, consent term, or derived-model right | OXE_NAMED_MIXTURES['oxe_magic_soup_plus'] | https://github.com/openvla/openvla/blob/c8f03f48af69/prismatic/vla/datasets/rlds/oxe/mixtures.py (project) | 2026-07-22 | high | human |
| openvla-7b / oxe_magic_soup_plus | austin_sailor_dataset_converted_externally_to_rlds | unknown — mixtures.py does not pin a component dataset version | 1.0 | relative sampling weight | Active tuple entries only; commented broken or omitted entries are excluded by the source config. | unknown — mixtures.py points to separate dataset transforms/configs and does not define one shared normalization rule | unknown — heterogeneous component action schemas and control rates are not reported in mixtures.py | unknown — mixtures.py does not report success/failure composition | unknown in source config — define a target-embodiment holdout before training | unknown — code availability does not establish compatibility of every dataset, consent term, or derived-model right | OXE_NAMED_MIXTURES['oxe_magic_soup_plus'] | https://github.com/openvla/openvla/blob/c8f03f48af69/prismatic/vla/datasets/rlds/oxe/mixtures.py (project) | 2026-07-22 | high | human |
| openvla-7b / oxe_magic_soup_plus | austin_sirius_dataset_converted_externally_to_rlds | unknown — mixtures.py does not pin a component dataset version | 1.0 | relative sampling weight | Active tuple entries only; commented broken or omitted entries are excluded by the source config. | unknown — mixtures.py points to separate dataset transforms/configs and does not define one shared normalization rule | unknown — heterogeneous component action schemas and control rates are not reported in mixtures.py | unknown — mixtures.py does not report success/failure composition | unknown in source config — define a target-embodiment holdout before training | unknown — code availability does not establish compatibility of every dataset, consent term, or derived-model right | OXE_NAMED_MIXTURES['oxe_magic_soup_plus'] | https://github.com/openvla/openvla/blob/c8f03f48af69/prismatic/vla/datasets/rlds/oxe/mixtures.py (project) | 2026-07-22 | high | human |
| openvla-7b / oxe_magic_soup_plus | dlr_edan_shared_control_converted_externally_to_rlds | unknown — mixtures.py does not pin a component dataset version | 1.0 | relative sampling weight | Active tuple entries only; commented broken or omitted entries are excluded by the source config. | unknown — mixtures.py points to separate dataset transforms/configs and does not define one shared normalization rule | unknown — heterogeneous component action schemas and control rates are not reported in mixtures.py | unknown — mixtures.py does not report success/failure composition | unknown in source config — define a target-embodiment holdout before training | unknown — code availability does not establish compatibility of every dataset, consent term, or derived-model right | OXE_NAMED_MIXTURES['oxe_magic_soup_plus'] | https://github.com/openvla/openvla/blob/c8f03f48af69/prismatic/vla/datasets/rlds/oxe/mixtures.py (project) | 2026-07-22 | high | human |
| openvla-7b / oxe_magic_soup_plus | iamlab_cmu_pickup_insert_converted_externally_to_rlds | unknown — mixtures.py does not pin a component dataset version | 1.0 | relative sampling weight | Active tuple entries only; commented broken or omitted entries are excluded by the source config. | unknown — mixtures.py points to separate dataset transforms/configs and does not define one shared normalization rule | unknown — heterogeneous component action schemas and control rates are not reported in mixtures.py | unknown — mixtures.py does not report success/failure composition | unknown in source config — define a target-embodiment holdout before training | unknown — code availability does not establish compatibility of every dataset, consent term, or derived-model right | OXE_NAMED_MIXTURES['oxe_magic_soup_plus'] | https://github.com/openvla/openvla/blob/c8f03f48af69/prismatic/vla/datasets/rlds/oxe/mixtures.py (project) | 2026-07-22 | high | human |
| openvla-7b / oxe_magic_soup_plus | utaustin_mutex | unknown — mixtures.py does not pin a component dataset version | 1.0 | relative sampling weight | Active tuple entries only; commented broken or omitted entries are excluded by the source config. | unknown — mixtures.py points to separate dataset transforms/configs and does not define one shared normalization rule | unknown — heterogeneous component action schemas and control rates are not reported in mixtures.py | unknown — mixtures.py does not report success/failure composition | unknown in source config — define a target-embodiment holdout before training | unknown — code availability does not establish compatibility of every dataset, consent term, or derived-model right | OXE_NAMED_MIXTURES['oxe_magic_soup_plus'] | https://github.com/openvla/openvla/blob/c8f03f48af69/prismatic/vla/datasets/rlds/oxe/mixtures.py (project) | 2026-07-22 | high | human |
| openvla-7b / oxe_magic_soup_plus | berkeley_fanuc_manipulation | unknown — mixtures.py does not pin a component dataset version | 2.0 | relative sampling weight | Active tuple entries only; commented broken or omitted entries are excluded by the source config. | unknown — mixtures.py points to separate dataset transforms/configs and does not define one shared normalization rule | unknown — heterogeneous component action schemas and control rates are not reported in mixtures.py | unknown — mixtures.py does not report success/failure composition | unknown in source config — define a target-embodiment holdout before training | unknown — code availability does not establish compatibility of every dataset, consent term, or derived-model right | OXE_NAMED_MIXTURES['oxe_magic_soup_plus'] | https://github.com/openvla/openvla/blob/c8f03f48af69/prismatic/vla/datasets/rlds/oxe/mixtures.py (project) | 2026-07-22 | high | human |
| openvla-7b / oxe_magic_soup_plus | cmu_stretch | unknown — mixtures.py does not pin a component dataset version | 1.0 | relative sampling weight | Active tuple entries only; commented broken or omitted entries are excluded by the source config. | unknown — mixtures.py points to separate dataset transforms/configs and does not define one shared normalization rule | unknown — heterogeneous component action schemas and control rates are not reported in mixtures.py | unknown — mixtures.py does not report success/failure composition | unknown in source config — define a target-embodiment holdout before training | unknown — code availability does not establish compatibility of every dataset, consent term, or derived-model right | OXE_NAMED_MIXTURES['oxe_magic_soup_plus'] | https://github.com/openvla/openvla/blob/c8f03f48af69/prismatic/vla/datasets/rlds/oxe/mixtures.py (project) | 2026-07-22 | high | human |
| openvla-7b / oxe_magic_soup_plus | bc_z | v0.1.0 (source comment) | 0.2 | relative sampling weight | Active tuple entries only; commented broken or omitted entries are excluded by the source config. | unknown — mixtures.py points to separate dataset transforms/configs and does not define one shared normalization rule | unknown — heterogeneous component action schemas and control rates are not reported in mixtures.py | unknown — mixtures.py does not report success/failure composition | unknown in source config — define a target-embodiment holdout before training | unknown — code availability does not establish compatibility of every dataset, consent term, or derived-model right | OXE_NAMED_MIXTURES['oxe_magic_soup_plus'] | https://github.com/openvla/openvla/blob/c8f03f48af69/prismatic/vla/datasets/rlds/oxe/mixtures.py (project) | 2026-07-22 | high | human |
| openvla-7b / oxe_magic_soup_plus | fmb_dataset | unknown — mixtures.py does not pin a component dataset version | 1.0 | relative sampling weight | Active tuple entries only; commented broken or omitted entries are excluded by the source config. | unknown — mixtures.py points to separate dataset transforms/configs and does not define one shared normalization rule | unknown — heterogeneous component action schemas and control rates are not reported in mixtures.py | unknown — mixtures.py does not report success/failure composition | unknown in source config — define a target-embodiment holdout before training | unknown — code availability does not establish compatibility of every dataset, consent term, or derived-model right | OXE_NAMED_MIXTURES['oxe_magic_soup_plus'] | https://github.com/openvla/openvla/blob/c8f03f48af69/prismatic/vla/datasets/rlds/oxe/mixtures.py (project) | 2026-07-22 | high | human |
| openvla-7b / oxe_magic_soup_plus | dobbe | unknown — mixtures.py does not pin a component dataset version | 0.2 | relative sampling weight | Active tuple entries only; commented broken or omitted entries are excluded by the source config. | unknown — mixtures.py points to separate dataset transforms/configs and does not define one shared normalization rule | unknown — heterogeneous component action schemas and control rates are not reported in mixtures.py | unknown — mixtures.py does not report success/failure composition | unknown in source config — define a target-embodiment holdout before training | unknown — code availability does not establish compatibility of every dataset, consent term, or derived-model right | OXE_NAMED_MIXTURES['oxe_magic_soup_plus'] | https://github.com/openvla/openvla/blob/c8f03f48af69/prismatic/vla/datasets/rlds/oxe/mixtures.py (project) | 2026-07-22 | high | human |
| openvla-7b / oxe_magic_soup_plus | droid | unknown — mixtures.py does not pin a component dataset version | 0.06 | relative sampling weight | Active tuple entries only; commented broken or omitted entries are excluded by the source config. | unknown — mixtures.py points to separate dataset transforms/configs and does not define one shared normalization rule | unknown — heterogeneous component action schemas and control rates are not reported in mixtures.py | unknown — mixtures.py does not report success/failure composition | unknown in source config — define a target-embodiment holdout before training | unknown — code availability does not establish compatibility of every dataset, consent term, or derived-model right | OXE_NAMED_MIXTURES['oxe_magic_soup_plus'] | https://github.com/openvla/openvla/blob/c8f03f48af69/prismatic/vla/datasets/rlds/oxe/mixtures.py (project) | 2026-07-22 | high | human |
Gate quality before reweighting
Run structural checks first, then preserve human task and rights judgments as separate records. Influence-based curation and mutual-information curation are robotics-primary ranking signals; neither replaces load validation, task review, a rights decision, or target-embodiment evaluation.
| ID | Layer | Signal | Decision rule | Limitation |
|---|---|---|---|---|
| QUALITY-AUTO-LOAD-SCHEMA | automated fact | load result, missing keys, unexpected keys, dtype and shape changes | accept conforming episodes; quarantine recoverable schema drift; reject unreadable episodes | Operational gate derived for auditability; the cited paper does not prescribe these parser decisions. |
| QUALITY-AUTO-TIMING-COMPLETENESS | automated fact | timestamp order, NaN/Inf counts, stream-length deltas, terminal markers | accept complete aligned streams; quarantine repairable gaps; reject non-reconstructable timing or value corruption | A structural pass does not prove that an action was intentional, safe, or useful. |
| QUALITY-AUTO-DISTRIBUTION | automated fact | counts and coverage by declared target-domain slice | report the distribution without one global score; quarantine a recipe when a required target slice is absent | source-reported scaling behavior is task-specific and does not supply universal mixture weights or thresholds. |
| QUALITY-AUTO-DUPLICATION-LEAKAGE | automated fact | exact/near duplicate groups and identity overlap between training and evaluation | accept disjoint splits; quarantine ambiguous provenance; reject confirmed evaluation leakage | No single fingerprint detects every semantic duplicate or hidden upstream overlap. |
| QUALITY-AUTO-RIGHTS-PROVENANCE | human judgment | presence and review status of each distinct rights/provenance artifact | accept only reviewed compatible terms; quarantine missing or ambiguous terms; reject known incompatible use | The cited curation paper does not provide legal guidance; compatibility requires qualified human review. |
| QUALITY-HUMAN-TASK-VALIDITY | human judgment | reviewer decision with reason code and task-specific rubric | accept desired valid behavior; quarantine uncertain or recoverable behavior; reject invalid, unsafe, or out-of-scope behavior | Human judgments can disagree; retain reviewer identity, rubric version, and disagreement rather than collapsing them into one score. |
| QUALITY-HUMAN-CONTRIBUTION | human judgment | influence estimate plus trajectory-level coverage review | compare ablations over multiple retained-set sizes; do not publish a universal cutoff | source-reported, not independently validated; rankings depend on the model, validation set, estimator, and target behavior. |
Hold out the target embodiment and log ablations
A recipe is a hypothesis. Keep target-embodiment evaluation outside training, record split identities, and compare the published starting weights with at least uniform and target-domain-reweighted alternatives. The methodology page documents automated facts and human decisions; the ledger preserves the exact recipe and explicit unknowns.
- 01
Pin inputs
Record dataset version, transform, action schema, rate, filter, rights status, and split identity for every component.
- 02
Validate before sampling
Reject unreadable or leaked episodes, quarantine ambiguous records, and retain human review reasons separately from automated facts.
- 03
Ablate and report
Compare alternative weights on the same held-out target-embodiment evaluation and retain the ablation log. Do not generalize one result into a universal recipe.
Limitations
This ledger proves that the published recipe is reproducible at one pinned config locator and that the grading rationale is transparent. It does not prove that the recipe improves a model, that the weights transfer to another target, or that the methodology is field-canonical. Those questions require target-specific experiments and review.
Related pages
Use these to move from category-level context into specific task, dataset, format, and comparison detail.
External references and source context
- Data Scaling Laws in Imitation Learning for Robotic Manipulation
Robotics-primary scaling evidence separates environment/object diversity from raw demonstration count in its reported protocol.
arXiv
FAQ
Are the OpenVLA mixture weights normalized probabilities?
No. The pinned config calls them sampling weights. Preserve the exact source values and document any normalization performed by the training loader.
How does OpenVLA represent actions?
OpenVLA represents actions as discretized action tokens. Do not transfer the architecture of a different policy into this recipe.
How should demonstrations be graded?
Keep parser, timing, completeness, distribution, and leakage facts separate from human judgments about task validity, strategy quality, and rights. Do not collapse them into one opaque score.
Turn the mixture gap into a sample brief
Start with a pinned recipe, preserve unknowns, and request the target-domain slice only after compatibility, leakage, and evaluation boundaries are explicit.
Source the underrepresented slice in your mixture