← RoboSteer homepage

Dataset guide

RoboSteer contains 773,716 task definitions across three levels. The complete release is distributed as archives on Hugging Face. Its Dataset Viewer shows 22 selected examples for browsing; those examples are not the full benchmark or a training/test split.

Open the RoboSteer dataset on Hugging Face ↗

The Hugging Face dataset card is the source for the current archive inventory, checksum list, extraction procedure, task map, and license status.

Download the complete release

The release contains 23 independent TAR archives, about 355.35 GiB in total. The extracted data is about 352.05 GiB. Keeping both copies requires roughly 707.40 GiB, plus filesystem and download-cache overhead. Download every archive group: a task can reference assets under Data/Shared/ regardless of its level.

Install the Hugging Face client, then download the archives and checksum file into a download folder:

python -m pip install --upgrade huggingface_hub
hf download YanCORANV/RoboSteer --repo-type dataset --include "archives/**" "SHA256SUMS" --local-dir RoboSteer-download

If the download is interrupted, repeat the same command with the same local directory. For the verified, resumable PowerShell extraction procedure, follow Step 2 on the Hugging Face dataset card. Its script checks each archive against SHA256SUMS before extracting and records completed archives so an interrupted extraction can continue.

Extract all archives into one separate dataset root. The download folder contains TAR files; it is not the dataset root. Keep the archives until extraction succeeds. Do not extract each TAR into its own subfolder.

Directory structure and paths

dataset-root/
├── Data/
│   ├── Level1/    Audio, Image, Motion, Spatial, Video
│   ├── Level2/    Audio, Order, Times, Trajectory
│   ├── Level3/    Audio, Image, Video
│   └── Shared/    Metadata, Motion, Video/Human, Video/Skeleton
└── Tasks/
    ├── Level1/
    │   ├── full_conditioning_reproduction/
    │   ├── spatial_completion/
    │   └── temporal_completion/
    ├── Level2/    Amplitude, BodyRestrain, Direction,
    │              Order, Speed, Times, Trajectory
    └── Level3/    task JSON files

Tasks/ contains task-definition JSON files. Data/ contains the assets they reference. After extraction, there should be one Data/ and one Tasks/ directly under the dataset root. Resolve paths in task JSONs against that root, not against the JSON file's directory. For example, Data/Shared/Motion/<shard>/<sample> refers to <dataset-root>/Data/Shared/Motion/<shard>/<sample>.

Read a task

Start with the working Python example on Hugging Face. It opens a real Level 1 task JSON and reads the first frame of its referenced G1 motion. The task and folder map lists all task directories, while the JSON field reference explains nested fields and task-specific semantics.

FieldMeaning
metadataTask identity, level or family, duration, and source references.
input.promptsGeneral instruction and task-specific modifier.
input.modalitiesText, audio, video, image, or spatial conditions for Levels 1 and 2.
input.interleave.sequenceOrdered and timed multimodal conditions for Level 3.
ground_truthReference or target motion and task-specific constraints; it is not an extra conditioning input.

A literal "none" in ground_truth.rendered_video means that optional video is absent; it is not a filename. Reference media are not model predictions.

Motion formats and reference semantics

Shared G1 motion packages use six synchronized CSV files: joint_pos.csv, joint_vel.csv, body_pos.csv, body_quat.csv, body_lin_vel.csv, and body_ang_vel.csv. The appendix describes 50 Hz output and 29 joint degrees of freedom. Read each package's metadata for timing and coordinate conventions; a task's duration alone does not determine its stored frame count. Order and Times instead reference AMASS/BABEL-linked PKL motion packages, and Rotation-to-Pose conditions can also use PKL files.

For Level 2 Amplitude, Speed, Direction, and Body Restrain, the reference motion is the unmodified source action. The instruction specifies the change the model should make. For Order, preserve the order of both input clips; for Times, the recorded source video may show a single action even when the reference PKL encodes repetition. Temporal Completion references point to the target segment; do not crop them again solely from the recorded time window. Level 3 input components must retain their recorded ordering and timestamps.

Order, Times, and data rights

The full archives include the existing Level 2 Order and Times materials. These tasks use AMASS motion data and BABEL annotations; no separate local construction is required to reproduce the files in this release. The repository does not assign one blanket license to all contents. The team's own-data license and redistribution scope for third-party-derived materials remain under review. Availability for download does not grant redistribution or commercial-use rights. Review the dataset card's data-source and licensing status and the applicable AMASS and BABEL terms.

Paper and citation

Minghe Gao et al., Benchmarking Behavioral Steerability in Behavior Foundation Models, arXiv:2610.10198 (2026). The dataset card provides BibTeX. Cite AMASS and BABEL separately where applicable.

Return to the project