Render action views
Use the robot’s URDF to render articulated configurations from four prescribed virtual cameras, independent of the observation cameras.
01 / VIDEO
02 / OVERVIEW
What if robot actions could be modeled in the same visual space as the world they change?
Dream4ACT renders target joint configurations as four action views using URDF-based forward kinematics. The views preserve full-arm articulation and gripper geometry, while giving different embodiments a fixed-shape visual interface.
A shared video model learns to predict observations, infer actions, or generate both. Training-free multiview matching then recovers executable joint targets from the predicted action views.
Video generation models (VGMs) offer strong spatiotemporal priors for embodied observation–action modeling. However, joint-space action vectors lack explicit image-space structure and vary in dimensionality and semantics across embodiments, making it challenging to directly leverage the rich spatiotemporal priors of VGMs. End-effector visualizations provide an alternative but do not specify the full articulated configuration needed for robot execution. We present Dream4ACT, a world model built for joint video-action modeling across embodiments. To unify action representations across embodiments, we introduce a shared visual action interface, called action views, which render target joint configurations from four prescribed virtual cameras using URDF-based forward kinematics. This shared visual representation preserves embodiment-specific articulated geometry while allowing observation and action sequences to share a video autoencoder and diffusion transformer. Through masked flow-matching, our model supports forward dynamics, inverse dynamics, and joint observation–action generation within a single jointly trained model by varying which future sequences are corrupted. To recover executable action sequences from predicted action views, we propose a training-free, URDF-constrained multiview recovery mechanism, without a learned embodiment-specific decoder. Dream4ACT achieves an average success rate of 88.98% on RoboTwin 2.0 and an overall score of 65.66 on TriWorldBench, supporting effective closed-loop manipulation and competitive action-conditioned multiview prediction through the visual action interface.
03 / METHOD
Use the robot’s URDF to render articulated configurations from four prescribed virtual cameras, independent of the observation cameras.
A shared video VAE and diffusion transformer process RGB observations and action views. A frozen VLM and trainable semantic adapter provide instruction-grounded scene features; masked flow matching selects the prediction mode.
Match predicted views against URDF-based renderings across all four cameras to recover actions without a learned embodiment-specific decoder.
Predict how the scene evolves under a given action trajectory.
Infer the actions consistent with observed or desired future outcomes.
Generate future observations and action views together from the current context.
04 / RESULTS
88.98%
Average task success
50 bimanual manipulation tasks
65.66
Overall TWB-Score
500-episode official test set
4 platforms
One jointly trained checkpoint
Single-arm and bimanual manipulation
| Method | Clean | Randomized | Average |
|---|---|---|---|
| π0.5 | 82.74 | 76.76 | 79.80 |
| Motus | 88.66 | 87.02 | 87.80 |
| LingBot-VA | 92.90 | 91.50 | 92.20 |
| Fast-WAM | 91.88 | 91.78 | 91.80 |
| Dream4ACT | 90.50 | 87.46 | 88.98 |
The full 50-task evaluation uses a separately trained checkpoint; the five-embodiment evaluation below uses its own jointly trained checkpoint.
| Method | TVC | TA | P3D | MQ | TC | VQ | Overall |
|---|---|---|---|---|---|---|---|
| Ctrl-World | 57.42 | 43.72 | 34.67 | 29.28 | 46.17 | 16.64 | 38.98 |
| Motus | 66.70 | 49.69 | 34.60 | 24.63 | 26.56 | 16.26 | 42.35 |
| Genie Envisioner | 62.39 | 33.17 | 54.00 | 20.17 | 32.18 | 17.46 | 40.73 |
| DreamDojo | 69.63 | 56.24 | 43.84 | 27.96 | 60.84 | 21.02 | 51.72 |
| BWM | 81.87 | 86.05 | 60.40 | 41.29 | 62.81 | 31.42 | 65.54 |
| Dream4ACT | 81.63 | 84.22 | 61.30 | 41.66 | 64.88 | 31.42 | 65.66 |
TVC: tri-view consistency · TA: task alignment · P3D: physical and 3D coherence · MQ: motion quality · TC: temporal consistency · VQ: visual quality.
| Embodiment | Clean (%) | Randomized (%) | Position error (mm) ↓ | Rotation error (°) ↓ | Average (%) |
|---|---|---|---|---|---|
| Aloha-Agilex | 81.94 | 82.45 | 1.39 | 1.57 | 82.19 |
| Piper | 81.55 | 82.32 | 1.66 | 1.71 | 81.94 |
| ARX-X5 | 82.19 | 80.26 | 1.22 | 1.95 | 81.23 |
| Franka-Panda | 64.52 | 62.45 | 4.11 | 9.45 | 63.48 |
| UR5-Xsg | 31.87 | 30.06 | 11.85 | 21.39 | 30.97 |
This setting is separate from the full 50-task RoboTwin 2.0 evaluation. End-effector recovery errors are measured from ground-truth action views using five held-out trajectories per task and 40 future frames per window, with equal task weighting.
05 / PERFORMANCE
Real-world demo
Video coming soonReal-world demo
Video coming soonReal-world demo
Video coming soon| Platform | Place block | Wipe plate | Stack blocks | Storage item | Common | All tasks |
|---|---|---|---|---|---|---|
| Aloha-Agilex | 85.0 | 90.0 | 20.0 | 50.0 | 87.5 | 61.3 |
| TienYi2.5 Pro | 90.0 | 80.0 | 30.0 | 65.0 | 85.0 | 66.3 |
| Franka Research 3 | 90.0 | 85.0 | – | – | 87.5 | – |
| UR5e | 85.0 | 95.0 | – | – | 90.0 | – |
SIMULATION / ROBOTWIN 2.0
@misc{zhu2026dream4actsharedvisualaction,
title={Dream4ACT: A Shared Visual Action Interface for Multi-Embodiment Video-Action Modeling},
author={Xiangyu Zhu and Jin Xu and Yue Guo and Xin Wu and Yifan Sun and Xiancong Ren and Jianxin Sun and Yong Dai and Xiaozhu Ju},
year={2026},
eprint={2609.40153},
archivePrefix={arXiv},
primaryClass={cs.RO},
url={https://arxiv.org/abs/2609.40153},
}