RoboSteerOverview

RoboSteer

Benchmarking Behavioral Steerability
in Behavior Foundation Models

Benchmarking Behavioral Steerability
in Behavior Foundation Models

Authors to be announced

Team and affiliations to be announced

* Work done in Joint Laboratory of Embodied Intelligence, ZJU & Unitree Robotics.
† Corresponding author.

Abstract

Evolution towards steerable behavior foundation models and comparison between steerable and non-steerable behaviors.
Figure 1. (a) Foundation models evolve from single-task generation toward steerable generalists. (b) Steerable behavior foundation models reliably execute multimodal user intentions.

Behavior Foundation Models (BFMs) are emerging as a new paradigm for translating diverse human intentions into executable humanoid behaviors. As these models evolve beyond behavior generation toward general-purpose behavioral systems, a fundamental question arises: can they be reliably steered according to user intentions?

In this paper, we introduce the concept of behavioral steerability, defined as the ability of BFMs to faithfully generate behaviors that satisfy user-specified behavioral intentions. To systematically study this capability, we present RoboSteer, the first benchmark for behavioral steerability in BFMs. RoboSteer organizes behavioral steerability into a three-level hierarchy—Conditional Steering, Constraint Steering, and Compositional Steering—and establishes a unified evaluation framework supported by a large-scale multimodal motion corpus.

Using RoboSteer, we conduct the first large-scale empirical study of behavioral steerability across existing BFMs. Our results reveal that strong behavior generation capability does not necessarily translate into strong behavioral steerability, with the performance gap widening as steering complexity increases. We hope RoboSteer will establish behavioral steerability as a fundamental capability for future BFMs and facilitate the development of more steerable general-purpose behavioral systems.

Benchmark Overview

RoboSteer organizes behavioral steerability into a three-level hierarchy across multimodal conditions, explicit constraints, and compositional sequences.

Overview of the three-level behavioral steerability hierarchy in RoboSteer.
Figure 4. Three-level behavioral steerability hierarchy in RoboSteer.
Level 1

Conditional Steering

12 task types

Generates behaviors matching input conditions with varying completeness: full conditioning reproduction, temporal completion, and spatial limb completion across text, video, audio, rhythm, and pose.

Level 2

Constraint Steering

7 constraint types

Executes behaviors under explicit physical constraints: controlling movement speed, amplitude, direction, action order, repetition times, spatial trajectory, or restraining specific limbs.

Level 3

Compositional Steering

1 task type

Interleaved multi-source steering that sequentially coordinates diverse conditions (video, text, audio, and keyframe boundaries) within a continuous behavior.

Evaluation Framework

A unified evaluation formulation decoupling basic generation quality from intention realization across steering levels.

Progressive computation of Intention Realization across three steering levels.
Figure 5. Progressive computation of Intention Realization (IR) and Behavioral Steerability (BS).
BG

Behavior Generation

Measures fundamental motion quality, fidelity, and diversity, evaluating baseline generation capacity independently of intention realization.

IR

Intention Realization

Quantifies intention fulfillment across levels: condition overlap (IR₁), constraint fulfillment (IR₂), and weighted sequence composition (IR₃).

BS

Behavioral Steerability

BS = BG × IR

Couples generation quality with intention realization—high steerability strictly requires both high-fidelity behaviors and faithful intention adherence.

Dataset & Resources

RoboSteer contains 38,522 source motions, 77,044 human and rendered motion sequences spanning 91.58 hours, and 773,716 behavioral task instances across three levels of behavioral steerability.

Level2 Intention Realization Evaluator

Choose one constraint and evaluate a single output from your model. Your Task ID lets the service find its benchmark reference.

Loading the official example…

📋

Evaluation Output

Choose a constraint, enter a Task ID, and upload all three CSV files.

Citation

Coming soon.

Contributors

Coming soon.