npx skills add ...
npx skills add nvidia/skills --skill i4h-workflow-dataset-convert
Convert workflow HDF5 recordings to LeRobot datasets for training or browser inspection. Use for conversion; do not use for replay, augmentation, or raw-data repair.
npx skills add nvidia/skills --skill i4h-workflow-dataset-convert
Preserve recorded actions, state, cameras, task text, and embodiment labels in a local LeRobot dataset.
Treat the resolver above as part of the skill contract: a hosted copy may run outside the base repository, so never assume the current checkout contains workflows/i4h_workflows. I4H_WORKFLOWS_REPO_URL selects the clone source. When I4H_WORKFLOWS is unset, derive the fallback directory from that URL; set I4H_WORKFLOWS only to reuse or choose a specific destination. Never replace an existing checkout.
Use the explicit/current-chain recording. Resolve its workflow and Scene from recording metadata/context, then read the Scene manifest for the embodiment and instruction. Use the embodiment manifest for labels. Do not assume state width equals action width; the converter derives both from the recording.
Use --fps or --skip-frames only when the user requests it or source metadata justifies it. Keep the default H.264 video codec for compatibility with GR00T's fast decord loader; select another --video-codec only when the target consumer requires it.
Conversion writes aggregate meta/stats.json for downstream policy loaders. Native G1 rule-based WBC recordings already contain 43-D state and 50-D action; the converter recognizes that contract and writes GR00T's required semantic meta/modality.json automatically. For a G1 recording made through the legacy 23-D Pink/keyboard contract and destined for a 50-D G1 WBC policy Task, add --g1-wbc-policy-actions. That explicit mapping combines the measured 43-joint state with the recorded navigation, base-height, and torso commands; require source action width 23 and state width 43.
Require:
meta/info.jsonmeta/stats.jsonmeta/modality.json when the target trainer requires semantic modality slicesFor G1, require modality metadata for both supported paths: native state=43/action=50, or explicitly mapped state=43/source-action=23/output-action=50. Treat a native 50-D dataset without meta/modality.json as incomplete.
Treat missing inputs or zero converted episodes as failure. If conversion leaves a partial destination, quarantine or remove that exact incomplete directory before retrying; never report it as usable.
On dimension errors, resolve the source workflow and embodiment again. On missing videos, confirm frames existed before conversion.
Require a readable HDF5 recording and its matching Scene plus embodiment manifests.
Conversion cannot reconstruct missing cameras, actions, state, task text, or successful episodes.
Convert my scissor pick-and-place recording into a LeRobot dataset. → resolve so101, convert successful episodes, and verify metadata, parquet, and both camera videos.Report source HDF5/workflow, embodiment, task text, source/converted/skipped counts, action/state widths, output directory/repo id, aggregate-stats/modality/parquet/video checks, and any missing modality.