Conditional Steering
12 task typesGenerates behaviors matching input conditions with varying completeness: full conditioning reproduction, temporal completion, and spatial limb completion across text, video, audio, rhythm, and pose.
Benchmarking Behavioral Steerability
in Behavior Foundation Models
Team and affiliations to be announced
Behavior Foundation Models (BFMs) are emerging as a new paradigm for translating diverse human intentions into executable humanoid behaviors. As these models evolve beyond behavior generation toward general-purpose behavioral systems, a fundamental question arises: can they be reliably steered according to user intentions?
In this paper, we introduce the concept of behavioral steerability, defined as the ability of BFMs to faithfully generate behaviors that satisfy user-specified behavioral intentions. To systematically study this capability, we present RoboSteer, the first benchmark for behavioral steerability in BFMs. RoboSteer organizes behavioral steerability into a three-level hierarchy—Conditional Steering, Constraint Steering, and Compositional Steering—and establishes a unified evaluation framework supported by a large-scale multimodal motion corpus.
Using RoboSteer, we conduct the first large-scale empirical study of behavioral steerability across existing BFMs. Our results reveal that strong behavior generation capability does not necessarily translate into strong behavioral steerability, with the performance gap widening as steering complexity increases. We hope RoboSteer will establish behavioral steerability as a fundamental capability for future BFMs and facilitate the development of more steerable general-purpose behavioral systems.
RoboSteer organizes behavioral steerability into a three-level hierarchy across multimodal conditions, explicit constraints, and compositional sequences.
Generates behaviors matching input conditions with varying completeness: full conditioning reproduction, temporal completion, and spatial limb completion across text, video, audio, rhythm, and pose.
Executes behaviors under explicit physical constraints: controlling movement speed, amplitude, direction, action order, repetition times, spatial trajectory, or restraining specific limbs.
Interleaved multi-source steering that sequentially coordinates diverse conditions (video, text, audio, and keyframe boundaries) within a continuous behavior.
A unified evaluation formulation decoupling basic generation quality from intention realization across steering levels.
Measures fundamental motion quality, fidelity, and diversity, evaluating baseline generation capacity independently of intention realization.
Quantifies intention fulfillment across levels: condition overlap (IR₁), constraint fulfillment (IR₂), and weighted sequence composition (IR₃).
BS = BG × IR
Couples generation quality with intention realization—high steerability strictly requires both high-fidelity behaviors and faithful intention adherence.
RoboSteer contains 38,522 source motions, 77,044 human and rendered motion sequences spanning 91.58 hours, and 773,716 behavioral task instances across three levels of behavioral steerability.
Open the full benchmark, evaluator weights, and three supporting RoboSteer repositories directly.
Complete task definitions and multimodal motion assets across Levels 1–3. Start here to download the benchmark.
Evaluator model weights and registry files for RoboSteer metrics. Download one evaluator or the complete collection.
Multimodal evaluation assets organized by level, modality, and method or task unit.
Model documentation and resource manifests. Runnable model packages are not yet published.
Reusable Level 1 processed text instructions and image-conditioned static videos.
Choose one constraint and evaluate a single output from your model. Your Task ID lets the service find its benchmark reference.
Choose a constraint, enter a Task ID, and upload all three CSV files.
Choose Order or Times, enter a Task ID, upload video, and provide your VLM credentials.
Coming soon.
Coming soon.