Not provided
The API record was retrieved, but the requested field was absent. This does not prove that no license, paper, or size statement exists elsewhere.
DATASET METHODOLOGY
One method covers the Hugging Face dataset profiles: use explicit source metadata, preserve missing fields, and label failed fetches.
DIRECT ANSWER
Each profile reads a committed snapshot of the dataset's own Hugging Face API record. The snapshot records license, size category, downloads, creator namespace, and linked arXiv ids. A missing value is shown as not provided. A failed fetch is shown as unverified. Neither state is replaced with a guessed license.
SOURCE ORDER
The refresh job requests the per-dataset API endpoint. It reads an explicit cardData.license value first, then an explicitlicense: tag. Some gated cards put the license in their access agreement; in that case the job accepts only a named license or an official license link in that text. Size comes fromcardData.size_categories or its matching tag. Downloads and creator come from the API record. Papers come only from explicitarxiv: tags.
The generated JSON is committed with the application. Page builds do not call Hugging Face, so a network outage cannot change a build or silently replace source facts. A later refresh can update the snapshot and make the resulting diff reviewable.
MISSING VS UNVERIFIED
The API record was retrieved, but the requested field was absent. This does not prove that no license, paper, or size statement exists elsewhere.
The per-dataset request failed or was access-restricted. The profile labels the source facts unverified and links to the live record for review.
HOW TO USE THE FACTS
A license name is reported as source metadata, not as a complete rights opinion. A license marked non-commercial is labeled as a commercial-use restriction. Other licenses still need review of the linked terms, upstream sources, and the intended use. When no license is provided, commercial use is not established.
Size categories and download counts are discovery facts, not quality scores. Before using a dataset, inspect the current card and files, confirm the relevant modalities and formats, and test a representative sample. Related pages are alternatives for source review; shared tags do not prove that two datasets are equivalent.
DATA QUALITY FOR ROBOT LEARNING
Robot demonstration quality is not one score. This methodology keeps automated facts—loadability, schema drift, timing, missing/NaN values, stream lengths, distribution, duplication, and leakage—separate from human judgment about task validity, strategy, safety, and rights. Influence functions and mutual-information estimators are cited robotics-primary curation signals; neither replaces structural validation or a human decision.
TrueLabel analysis is limited to the operational gates and the automated-versus-human split. Source-reported methods retain their exact scope and limitations. This is a transparent demonstration curation record for high-quality robot data review, not a field-canonical grade or a promised model result.
Read the versioned mixture recipe or download the mixture and quality ledger as JSON.
CHECKED 2026-07-22
| id | source_reported_version | normalized_field | normalized_value | unit | primary_source_url | source_type | source_id | exact_locator | checked_date | retrieval_hash | confidence | status | grading_layer | evidence_basis | signal | decision_rule | limitation |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| QUALITY-AUTO-LOAD-SCHEMA | TrueLabel methodology v1.0.0 | loadability and schema drift | parse every episode; compare required observation/action keys and dtypes | episode | https://arxiv.org/abs/2603.09056 | paper | paper-arxiv-org-abs-2603-09056 | Sections IV-A–IV-B (validation-relative quality) | 2026-07-22 | unknown — no retrieval snapshot is committed | medium | automatic | automated fact | TrueLabel analysis | load result, missing keys, unexpected keys, dtype and shape changes | accept conforming episodes; quarantine recoverable schema drift; reject unreadable episodes | Operational gate derived for auditability; the cited paper does not prescribe these parser decisions. |
| QUALITY-AUTO-TIMING-COMPLETENESS | TrueLabel methodology v1.0.0 | timing, missing values, and stream completeness | check monotonic timestamps, finite values, aligned stream lengths, and episode boundaries | stream | https://arxiv.org/abs/2603.09056 | paper | paper-arxiv-org-abs-2603-09056 | Section IV-B (trajectory-wise curation) | 2026-07-22 | unknown — no retrieval snapshot is committed | medium | automatic | automated fact | TrueLabel analysis | timestamp order, NaN/Inf counts, stream-length deltas, terminal markers | accept complete aligned streams; quarantine repairable gaps; reject non-reconstructable timing or value corruption | A structural pass does not prove that an action was intentional, safe, or useful. |
| QUALITY-AUTO-DISTRIBUTION | arXiv:2410.18647v1 | environment and object distribution coverage | report coverage by held-out target environment, object, task, and embodiment | slice | https://arxiv.org/html/2410.18647v1 | paper | paper-arxiv-org-abs-2410-18647 | Abstract and evaluation protocol | 2026-07-22 | unknown — no retrieval snapshot is committed | high | needs-review | automated fact | TrueLabel analysis | counts and coverage by declared target-domain slice | report the distribution without one global score; quarantine a recipe when a required target slice is absent | source-reported scaling behavior is task-specific and does not supply universal mixture weights or thresholds. |
| QUALITY-AUTO-DUPLICATION-LEAKAGE | TrueLabel methodology v1.0.0 | duplication and train/eval leakage | compare content, episode, task, environment, and embodiment identities across splits | pairwise match | https://arxiv.org/abs/2603.09056 | paper | paper-arxiv-org-abs-2603-09056 | Section IV-B (coverage-preserving trajectory selection) | 2026-07-22 | unknown — no retrieval snapshot is committed | medium | automatic | automated fact | TrueLabel analysis | exact/near duplicate groups and identity overlap between training and evaluation | accept disjoint splits; quarantine ambiguous provenance; reject confirmed evaluation leakage | No single fingerprint detects every semantic duplicate or hidden upstream overlap. |
| QUALITY-AUTO-RIGHTS-PROVENANCE | TrueLabel methodology v1.0.0 | rights and provenance completeness | record code license, dataset terms, consent basis, provenance, and derived-model terms separately | dataset component | https://www.roboticsproceedings.org/rss21/p023.html | paper | paper-rss21-p023 | Paper scope and limitations (quality signal, not rights review) | 2026-07-22 | unknown — no retrieval snapshot is committed | medium | needs-review | human judgment | TrueLabel analysis | presence and review status of each distinct rights/provenance artifact | accept only reviewed compatible terms; quarantine missing or ambiguous terms; reject known incompatible use | The cited curation paper does not provide legal guidance; compatibility requires qualified human review. |
| QUALITY-HUMAN-TASK-VALIDITY | RSS 2025 paper 23 | demonstration validity and strategy quality | review task completion, recoveries, unsafe shortcuts, and whether behavior represents the desired policy | trajectory | https://www.roboticsproceedings.org/rss21/p023.html | paper | paper-rss21-p023 | Abstract and human-expert quality comparison | 2026-07-22 | unknown — no retrieval snapshot is committed | high | human | human judgment | TrueLabel analysis | reviewer decision with reason code and task-specific rubric | accept desired valid behavior; quarantine uncertain or recoverable behavior; reject invalid, unsafe, or out-of-scope behavior | Human judgments can disagree; retain reviewer identity, rubric version, and disagreement rather than collapsing them into one score. |
| QUALITY-HUMAN-CONTRIBUTION | arXiv:2603.09056 | contribution to held-out desired behavior | estimate trajectory contribution relative to a declared validation set, then review selection coverage | trajectory | https://arxiv.org/abs/2603.09056 | paper | paper-arxiv-org-abs-2603-09056 | Sections IV-A–IV-B | 2026-07-22 | unknown — no retrieval snapshot is committed | high | needs-review | human judgment | source-reported | influence estimate plus trajectory-level coverage review | compare ablations over multiple retained-set sizes; do not publish a universal cutoff | source-reported, not independently validated; rankings depend on the model, validation set, estimator, and target behavior. |
LIMITATIONS
A structural pass does not prove task usefulness. A curation score depends on the model, estimator, and validation set. A repository license does not establish dataset consent, redistribution, or derived-model rights. Keep those unknowns visible, retain reviewer reasons, and evaluate alternate retained sets on the same held-out target embodiment.