Dream4ACT

A Shared Visual Action Interface
for Multi-Embodiment Video-Action Modeling

Explore the research
Xiangyu Zhu1,2Jin Xu1Yue Guo1,3 Xin Wu1,4Yifan Sun1,5Xiancong Ren1 Jianxin Sun1Yong Dai1,†Xiaozhu Ju1,✉
1 Beijing Humanoid Robot Innovation Center 2 Beijing Institute of Technology 3 Harbin Institute of Technology, Shenzhen 4 The University of Hong Kong 5 China University of Mining & Technology, Beijing

† Project leader · ✉ Corresponding author

01 / VIDEO

02 / OVERVIEW

Different robots.
Shared visual
actions.

What if robot actions could be modeled in the same visual space as the world they change?

Dream4ACT renders target joint configurations as four action views using URDF-based forward kinematics. The views preserve full-arm articulation and gripper geometry, while giving different embodiments a fixed-shape visual interface.

A shared video model learns to predict observations, infer actions, or generate both. Training-free multiview matching then recovers executable joint targets from the predicted action views.

Read the full abstract

Video generation models (VGMs) offer strong spatiotemporal priors for embodied observation–action modeling. However, joint-space action vectors lack explicit image-space structure and vary in dimensionality and semantics across embodiments, making it challenging to directly leverage the rich spatiotemporal priors of VGMs. End-effector visualizations provide an alternative but do not specify the full articulated configuration needed for robot execution. We present Dream4ACT, a world model built for joint video-action modeling across embodiments. To unify action representations across embodiments, we introduce a shared visual action interface, called action views, which render target joint configurations from four prescribed virtual cameras using URDF-based forward kinematics. This shared visual representation preserves embodiment-specific articulated geometry while allowing observation and action sequences to share a video autoencoder and diffusion transformer. Through masked flow-matching, our model supports forward dynamics, inverse dynamics, and joint observation–action generation within a single jointly trained model by varying which future sequences are corrupted. To recover executable action sequences from predicted action views, we propose a training-free, URDF-constrained multiview recovery mechanism, without a learned embodiment-specific decoder. Dream4ACT achieves an average success rate of 88.98% on RoboTwin 2.0 and an overall score of 65.66 on TriWorldBench, supporting effective closed-loop manipulation and competitive action-conditioned multiview prediction through the visual action interface.

Overview of Dream4ACT. Top: A shared action-view interface supports forward dynamics, inverse dynamics, and joint generation, with training-free recovery of joint targets. Bottom left: Examples of multi embodiments real-world task execution. Bottom right: Success rates and scores show strong performance in both robotic manipulation and action-conditioned video generation.

03 / METHOD

Actions as views.
Views into actions.

01

Render action views

Use the robot’s URDF to render articulated configurations from four prescribed virtual cameras, independent of the observation cameras.

02

Model video and action

A shared video VAE and diffusion transformer process RGB observations and action views. A frozen VLM and trainable semantic adapter provide instruction-grounded scene features; masked flow matching selects the prediction mode.

03

Recover joint targets

Match predicted views against URDF-based renderings across all four cameras to recover actions without a learned embodiment-specific decoder.

Architecture of Dream4ACT. (a) URDF-based rendering produces four action views. (b) A shared video VAE and diffusion transformer model RGB observations and action views with VLM-derived semantic conditioning, here we show the joint generation mode. (c) Training-free multiview action solver recovers joint targets from predicted views.

One model. Three modes.

GIVENFuture action views
Dream4ACT
GENERATEDFuture observations

Predict how the scene evolves under a given action trajectory.

04 / RESULTS

One visual interface.
From prediction to control.

SIMULATIONRoboTwin 2.0

88.98%

Average task success

Clean90.50
Randomized87.46

50 bimanual manipulation tasks

WORLD MODELINGTriWorldBench

65.66

Overall TWB-Score

View consistency81.63
Task alignment84.22

500-episode official test set

REAL WORLDShared weights

4 platforms

One jointly trained checkpoint

Aloha-AgilexTienYi2.5 ProFranka Research 3UR5e

Single-arm and bimanual manipulation

Explore the benchmark results
Success rates (%) across 50 RoboTwin 2.0 tasks, with 100 trials per task and setting. Baseline results are those reported in the paper.
MethodCleanRandomizedAverage
π0.582.7476.7679.80
Motus88.6687.0287.80
LingBot-VA92.9091.5092.20
Fast-WAM91.8891.7891.80
Dream4ACT90.5087.4688.98

The full 50-task evaluation uses a separately trained checkpoint; the five-embodiment evaluation below uses its own jointly trained checkpoint.

05 / PERFORMANCE

Across bodies.
Into the real world.

Real-world evaluation details
Success rates (%) for one jointly trained checkpoint over 20 trials per embodiment–task pair, with randomized backgrounds and distractor objects. Common averages place block and wipe plate; All tasks reports the paper’s four-task averages for the bimanual platforms. “–” denotes a task not evaluated on that platform.
PlatformPlace blockWipe plateStack blocksStorage itemCommonAll tasks
Aloha-Agilex85.090.020.050.087.561.3
TienYi2.5 Pro90.080.030.065.085.066.3
Franka Research 390.085.0––87.5–
UR5e85.095.0––90.0–

SIMULATION / ROBOTWIN 2.0

Manipulation in simulation.

CITATION

@misc{zhu2026dream4actsharedvisualaction,
      title={Dream4ACT: A Shared Visual Action Interface for Multi-Embodiment Video-Action Modeling}, 
      author={Xiangyu Zhu and Jin Xu and Yue Guo and Xin Wu and Yifan Sun and Xiancong Ren and Jianxin Sun and Yong Dai and Xiaozhu Ju},
      year={2026},
      eprint={2609.40153},
      archivePrefix={arXiv},
      primaryClass={cs.RO},
      url={https://arxiv.org/abs/2609.40153}, 
}

Dream4ACT · Paper figure