HAI: Hierarchical Anchored Interaction for Multi-View Bimanual World Models

Structured conditioning for long-horizon multi-view bimanual rollout prediction

Anonymous Authors
Anonymous Institution

★Structure-Aligned Conditioning
Align control pathways with the physical structure of bimanual interaction.

HAI routes action and view information according to actuator ownership and camera observation structure, instead of collapsing both arms into a view-agnostic control signal.

★Failure-Mode-Targeted Design
Address wrong-arm response, cross-view inconsistency, and long-horizon drift with explicit modules.

Each HAI component is tied to a concrete failure mode, making the architecture problem-driven rather than a loose collection of generic world-modeling blocks.

★Policy-Relevant World Modeling
Generate rollouts that are useful not only visually, but also for downstream policy improvement.

By preserving scene persistence and synchronized multi-view interaction cues, HAI provides stronger synthetic rollout signals for downstream VLA policies.

HAI introduction figure

Abstract

Multi-view bimanual robot world models must predict future observations while preserving scene identity, following two synchronized but non-exchangeable arms, and maintaining consistency between global and wrist-local views. Existing conditioning schemes often collapse left- and right-arm actions into a view-agnostic control signal, making long-horizon rollout prone to scene drift, wrong-arm responses, and cross-view inconsistency. We propose HAI, a Hierarchical Anchored Interaction architecture for controllable multi-view bimanual world modeling. HAI organizes prediction as a structured information flow. First, hierarchical action-view conditioning routes the bimanual action chunk to camera streams, injects coarse chunk-level intent, performs structured multi-view modeling on the coarse-modulated features, and then injects fine per-frame action control. Second, anchored dynamic generation fuses the fine action-conditioned descriptors with persistent per-view scene anchors before decoding future observations. This design aligns actuator ownership, camera-specific prediction, cross-view evidence, and scene persistence within a single conditioning architecture. Experiments on AgiBot and DROID show that HAI improves long-horizon rollout quality over the baselines. Ablations confirm gains in scene stability, wrist-view controllability, and cross-view consistency, and policy-in-the-loop diagnostics indicate more useful synthetic rollout signals for downstream VLA policies.

HAI Pipeline

HAI organizes multi-view rollout as hierarchical action-view conditioning together with anchored dynamic generation inside a rectified-flow DiT backbone.

Method Overview
HAI pipeline architecture with hierarchical action-view interaction and anchored dynamic generation

Hierarchical action-view interaction and anchored dynamic generation. The pipeline routes action chunks to camera-specific streams, performs structured multi-view modeling inside DiT blocks, and fuses persistent anchor information before decoding future frames.

1. Structured conditioning formulation

We formulate bimanual multiview rollout as a structured conditioning problem and instantiate it as HAI, separating Hierarchical Action Alignment, Anchored Static-Dynamic Decoupling, and Interaction-Structured Multi-View Communication.

2. Rectified-flow DiT implementation

We implement HAI in a rectified-flow DiT with camera-contract masks for view communication and actuator exposure, preserving explicit structure across views and control channels.

3. AgiBot evaluation and diagnostics

On AgiBot, we compare HAI against Base and Ctrl-World on the same validation clips and matched generated trajectory lengths, and we report learned-world-model rollout diagnostics for downstream VLA policies.

Featured Long-Horizon Rollout

Long-horizon rollout cases are presented separately from short supplementary clips, while preserving the same structured comparison layout across views.

1 / 1
Case 01: Three-view bimanual rollout comparison

Held-out AgiBot validation clip rendered as a synchronized 2x3 panel for direct comparison between HAI rollout prediction and ground truth.

AgiBot validation Three-view rollout Prediction vs GT
Top row = HAI prediction
Bottom row = ground truth

Synchronized multi-view rollout on a held-out validation clip across third, left, and right cameras.

Short Highlights

Short supplementary clips are shown in a separate viewer so they do not mix with the long-horizon evidence, while keeping the same GT-versus-prediction presentation style.

1 / 1
Case 01: Short supplementary rollout clip

Compact supplementary clip shown in the same viewer format for direct comparison, with ground truth on top and HAI prediction on bottom.

Short clip Three-view rollout GT vs prediction
Top row = ground truth
Bottom row = HAI prediction

Short supplementary rollout on a separate clip across left, third, and right cameras.

Evaluation Lenses

Hierarchical Action Alignment

Check whether view-routed action tokens expose actuator information consistently with visible hand motion across left, third, and right views.

Anchored Static-Dynamic Decoupling

Inspect whether anchor retrieval preserves stable scene context while the history stream focuses on dynamic contact and object interaction.

Cross-view Communication

Evaluate whether camera-contract masks allow synchronized rollout prediction without washing out view-specific evidence.

Baseline and Policy Diagnostics

Use matched rollouts to compare HAI with Base and Ctrl-World, then study which learned-world-model diagnostics matter for downstream VLA policies.

Quantitative results table for HAI, Base model, Symmetric View Conditioning, and Ctrl-World on AgiBot validation clips

Quantitative results on AgiBot validation clips. We evaluate 10-second generated trajectories on the same 256 validation clips. Each method uses recorded validation actions and its own implementation-specific conditioning, without ground-truth visual-frame feedback after the initial context.

HAI vs Ctrl-World

1 / 3
Case 01
Comparison figure between Ground Truth, Ctrl-World, and HAI across rollout steps

HAI vs Ctrl-World. Compared with Ctrl-World, HAI maintains stronger long-horizon visual fidelity and interaction consistency under the same rollout setting.

AgiBot WorldArena Metrics

These additional AgiBot metrics show where HAI improves third-view and wrist-view behavior beyond the main rollout table, while preserving a camera-role-specific view of geometric fidelity, temporal smoothness, and semantic consistency.

12 Best or Tied-Best
HAI attains the strongest value in most camera-role entries.

Across third-view and wrist-view reporting, HAI is best or tied-best in 12 metric entries, indicating broad gains rather than a narrow single-metric advantage.

6 / 8 Wrist-View Leads
The wrist-view improvements are especially consistent.

HAI leads on six wrist-view metrics, including semantic alignment, depth accuracy, aesthetic quality, background consistency, motion smoothness, and subject consistency.

Depth + Motion
The strongest gains appear in geometry and temporal coherence.

HAI is strongest on depth accuracy and motion smoothness for both camera roles, supporting the qualitative impression of more stable multi-view rollouts.

Additional WorldArena-style evaluation metrics on AgiBot for Ctrl-World and HAI

Additional WorldArena-style evaluation metrics on AgiBot. Scores are normalized to a 0-100 scale, where higher is better. Third-view columns report the third camera, and wrist-view columns average the wrist cameras for each method. Best results are bolded within each camera role.

DROID Validation
Quantitative comparison between Ctrl-World and HAI on the DROID validation set

Quantitative results on the DROID validation set. Results are averaged within each camera type before reporting. HAI improves most metrics over Ctrl-World, with the clearest gains on external-view rollouts.

Policy-in-the-Loop Evaluation

Policy-in-the-loop world-model evaluation across five downstream VLA tasks

Policy-in-the-loop world-model evaluation on five downstream VLA tasks. Bars report per-task percentages of policy-generated rollout clips judged successful under the same rollout budget, averaged over pi0, pi0.5, and GO-1. This is a clip-level learned-world evaluation rather than an end-to-end physical robot success rate.

Real-Robot Experiments

We fine-tune pi0.5 on synthetic rollout videos generated by either HAI or Cosmos Predict 2.5, then deploy the resulting policies on two physical robot platforms. These trials test whether improved world-model rollouts translate into stronger task execution and recovery behavior.

HAI rollout fine-tuning Cosmos Predict 2.5 rollout fine-tuning Successful or corrective state Failure state

Franka Real-Robot Evaluation

Two long-horizon manipulation tasks compare pi0.5 policies fine-tuned from HAI and Cosmos Predict 2.5 rollout videos under the same physical setup.

Franka
Franka real-robot comparison of pi0.5 fine-tuned on HAI and Cosmos Predict 2.5 rollout videos across whiteboard wiping and bread transfer tasks
Franka Task 01

Wipe the whiteboard using a blackboard eraser

Observed result. The policy fine-tuned on HAI rollouts clears the writing, while the Cosmos Predict 2.5 baseline fails to complete the wipe.

Failure analysis. Successful wiping requires the rollout to preserve the identities and relative geometry of the eraser, whiteboard plane, writing region, and gripper across multiple views while maintaining the eraser's orientation and contact trajectory over time. HAI's structured multi-view conditioning produces more consistent object boundaries and higher-fidelity interaction frames, giving pi0.5 clearer supervision for approaching the board, establishing contact, choosing a sweep direction, and covering the written region. With weaker cross-view object modeling, small errors in the eraser pose or board contact location can accumulate through the synthetic rollout, leading the baseline policy to form unstable contact or miss part of the target area.

Franka Task 02

Transfer bread from the bowl into the pot and place the lid on top

Observed result. When the bread lands off-center, the HAI-trained policy corrects the placement, lays the bread flat, and continues the task. The Cosmos Predict 2.5 baseline does not recover.

Failure analysis. We attribute this recovery behavior to HAI's mixed positive and negative training samples. In addition to nominal transfers, the model observes perturbed states in which the bread is dropped off-center or remains tilted, together with the corrective re-contact, repositioning, and flattening motions needed before the lid can be placed. These rollouts teach pi0.5 that an imperfect placement is a recoverable intermediate state rather than the end of the action sequence. Without equally coherent recovery supervision, a small placement error can compound: the policy proceeds from an invalid bread pose, fails to restore a stable configuration, and cannot reliably complete the later lid-placement stage.

AgileX Real-Robot Evaluation

Two contact-rich manipulation tasks compare pi0.5 policies fine-tuned from HAI and Cosmos Predict 2.5 rollout videos on the AgileX dual-arm platform.

AgileX
AgileX dual-arm real-robot comparison of pi0.5 fine-tuned on HAI and Cosmos Predict 2.5 rollout videos
AgileX Task 01

Pour coffee beans into the ceramic cup

Observed result. The policy fine-tuned on HAI rollouts grasps the transparent cup containing the coffee beans, lifts it, and pours the beans into the ceramic cup. The Cosmos Predict 2.5 baseline fails to secure the initial grasp on the transparent cup.

Failure analysis. Transparent containers provide weak, background-dependent boundaries and ambiguous depth cues because their appearance is dominated by refraction and the scene behind them. HAI better preserves the cup's boundary, pose, and graspable geometry across views, so its synthetic rollouts provide pi0.5 with a clearer sequence of approach, grasp, lift, and pouring motions. A weaker transparent-object representation in the baseline makes the initial grasp pose less reliable, preventing the policy from reaching the later pouring stage.

AgileX Task 02

Open the black pen

Observed result. The policy fine-tuned on HAI rollouts stabilizes the pen and removes its cap. The Cosmos Predict 2.5 baseline struggles to grasp and detach the cap, leaving the pen unopened.

Failure analysis. Opening the pen requires the policy to preserve the identities of the pen body and cap while they partially occlude one another, track their contact state, and coordinate opposing stabilization and pulling motions. HAI's multi-view interaction modeling better maintains these part-level relationships through occlusion, yielding rollout supervision with more accurate grasp points and relative motion. When the cap and pen body are not modeled distinctly, small pose errors compound at contact and make the fine-tuned baseline policy less likely to generate enough controlled separation to remove the cap.

HAI Ablation Study

AgiBot Quantitative Ablation
Quantitative ablation results for HAI on AgiBot with third-view and wrist-view reporting

Ablation results on AgiBot. Third-view and wrist-view metrics are reported under the same training and inference budget. WorldArena-style evaluation metrics are reported on a 0-100 scale.