truelabelstaging

DATASET METHODOLOGY

How dataset source facts are checked

One method covers the Hugging Face dataset profiles: use explicit source metadata, preserve missing fields, and label failed fetches.

DIRECT ANSWER

Each profile reads a committed snapshot of the dataset's own Hugging Face API record. The snapshot records license, size category, downloads, creator namespace, and linked arXiv ids. A missing value is shown as not provided. A failed fetch is shown as unverified. Neither state is replaced with a guessed license.

SOURCE ORDER

What is copied into a profile

The refresh job requests the per-dataset API endpoint. It reads an explicit cardData.license value first, then an explicitlicense: tag. Some gated cards put the license in their access agreement; in that case the job accepts only a named license or an official license link in that text. Size comes fromcardData.size_categories or its matching tag. Downloads and creator come from the API record. Papers come only from explicitarxiv: tags.

The generated JSON is committed with the application. Page builds do not call Hugging Face, so a network outage cannot change a build or silently replace source facts. A later refresh can update the snapshot and make the resulting diff reviewable.

MISSING VS UNVERIFIED

Two states that must stay distinct

Fetch succeeded

Not provided

The API record was retrieved, but the requested field was absent. This does not prove that no license, paper, or size statement exists elsewhere.

Fetch failed

Unverified

The per-dataset request failed or was access-restricted. The profile labels the source facts unverified and links to the live record for review.

HOW TO USE THE FACTS

License, size, downloads, and commercial use

A license name is reported as source metadata, not as a complete rights opinion. A license marked non-commercial is labeled as a commercial-use restriction. Other licenses still need review of the linked terms, upstream sources, and the intended use. When no license is provided, commercial use is not established.

Size categories and download counts are discovery facts, not quality scores. Before using a dataset, inspect the current card and files, confirm the relevant modalities and formats, and test a representative sample. Related pages are alternatives for source review; shared tags do not prove that two datasets are equivalent.

DATA QUALITY FOR ROBOT LEARNING

How demonstrations are graded: automated facts vs human judgment

Robot demonstration quality is not one score. This methodology keeps automated facts—loadability, schema drift, timing, missing/NaN values, stream lengths, distribution, duplication, and leakage—separate from human judgment about task validity, strategy, safety, and rights. Influence functions and mutual-information estimators are cited robotics-primary curation signals; neither replaces structural validation or a human decision.

TrueLabel analysis is limited to the operational gates and the automated-versus-human split. Source-reported methods retain their exact scope and limitations. This is a transparent demonstration curation record for high-quality robot data review, not a field-canonical grade or a promised model result.

Read the versioned mixture recipe or download the mixture and quality ledger as JSON.

CHECKED 2026-07-22

Quality and curation ledger

Automated facts and human judgments from the canonical JSON record set
idsource_reported_versionnormalized_fieldnormalized_valueunitprimary_source_urlsource_typesource_idexact_locatorchecked_dateretrieval_hashconfidencestatusgrading_layerevidence_basissignaldecision_rulelimitation
QUALITY-AUTO-LOAD-SCHEMATrueLabel methodology v1.0.0loadability and schema driftparse every episode; compare required observation/action keys and dtypesepisodehttps://arxiv.org/abs/2603.09056paperpaper-arxiv-org-abs-2603-09056Sections IV-A–IV-B (validation-relative quality)2026-07-22unknown — no retrieval snapshot is committedmediumautomaticautomated factTrueLabel analysisload result, missing keys, unexpected keys, dtype and shape changesaccept conforming episodes; quarantine recoverable schema drift; reject unreadable episodesOperational gate derived for auditability; the cited paper does not prescribe these parser decisions.
QUALITY-AUTO-TIMING-COMPLETENESSTrueLabel methodology v1.0.0timing, missing values, and stream completenesscheck monotonic timestamps, finite values, aligned stream lengths, and episode boundariesstreamhttps://arxiv.org/abs/2603.09056paperpaper-arxiv-org-abs-2603-09056Section IV-B (trajectory-wise curation)2026-07-22unknown — no retrieval snapshot is committedmediumautomaticautomated factTrueLabel analysistimestamp order, NaN/Inf counts, stream-length deltas, terminal markersaccept complete aligned streams; quarantine repairable gaps; reject non-reconstructable timing or value corruptionA structural pass does not prove that an action was intentional, safe, or useful.
QUALITY-AUTO-DISTRIBUTIONarXiv:2410.18647v1environment and object distribution coveragereport coverage by held-out target environment, object, task, and embodimentslicehttps://arxiv.org/html/2410.18647v1paperpaper-arxiv-org-abs-2410-18647Abstract and evaluation protocol2026-07-22unknown — no retrieval snapshot is committedhighneeds-reviewautomated factTrueLabel analysiscounts and coverage by declared target-domain slicereport the distribution without one global score; quarantine a recipe when a required target slice is absentsource-reported scaling behavior is task-specific and does not supply universal mixture weights or thresholds.
QUALITY-AUTO-DUPLICATION-LEAKAGETrueLabel methodology v1.0.0duplication and train/eval leakagecompare content, episode, task, environment, and embodiment identities across splitspairwise matchhttps://arxiv.org/abs/2603.09056paperpaper-arxiv-org-abs-2603-09056Section IV-B (coverage-preserving trajectory selection)2026-07-22unknown — no retrieval snapshot is committedmediumautomaticautomated factTrueLabel analysisexact/near duplicate groups and identity overlap between training and evaluationaccept disjoint splits; quarantine ambiguous provenance; reject confirmed evaluation leakageNo single fingerprint detects every semantic duplicate or hidden upstream overlap.
QUALITY-AUTO-RIGHTS-PROVENANCETrueLabel methodology v1.0.0rights and provenance completenessrecord code license, dataset terms, consent basis, provenance, and derived-model terms separatelydataset componenthttps://www.roboticsproceedings.org/rss21/p023.htmlpaperpaper-rss21-p023Paper scope and limitations (quality signal, not rights review)2026-07-22unknown — no retrieval snapshot is committedmediumneeds-reviewhuman judgmentTrueLabel analysispresence and review status of each distinct rights/provenance artifactaccept only reviewed compatible terms; quarantine missing or ambiguous terms; reject known incompatible useThe cited curation paper does not provide legal guidance; compatibility requires qualified human review.
QUALITY-HUMAN-TASK-VALIDITYRSS 2025 paper 23demonstration validity and strategy qualityreview task completion, recoveries, unsafe shortcuts, and whether behavior represents the desired policytrajectoryhttps://www.roboticsproceedings.org/rss21/p023.htmlpaperpaper-rss21-p023Abstract and human-expert quality comparison2026-07-22unknown — no retrieval snapshot is committedhighhumanhuman judgmentTrueLabel analysisreviewer decision with reason code and task-specific rubricaccept desired valid behavior; quarantine uncertain or recoverable behavior; reject invalid, unsafe, or out-of-scope behaviorHuman judgments can disagree; retain reviewer identity, rubric version, and disagreement rather than collapsing them into one score.
QUALITY-HUMAN-CONTRIBUTIONarXiv:2603.09056contribution to held-out desired behaviorestimate trajectory contribution relative to a declared validation set, then review selection coveragetrajectoryhttps://arxiv.org/abs/2603.09056paperpaper-arxiv-org-abs-2603-09056Sections IV-A–IV-B2026-07-22unknown — no retrieval snapshot is committedhighneeds-reviewhuman judgmentsource-reportedinfluence estimate plus trajectory-level coverage reviewcompare ablations over multiple retained-set sizes; do not publish a universal cutoffsource-reported, not independently validated; rankings depend on the model, validation set, estimator, and target behavior.

LIMITATIONS

What these checks cannot establish

A structural pass does not prove task usefulness. A curation score depends on the model, estimator, and validation set. A repository license does not establish dataset consent, redistribution, or derived-model rights. Keep those unknowns visible, retain reviewer reasons, and evaluate alternate retained sets on the same held-out target embodiment.

Where to go next

Sources