arXiv 2609.10506

DUET-DINO: Simultaneous Cross-View World Modeling for Latent Planning in Robot Manipulation

By Nisarga Nilavadi, Ralf Römer, et al.

Published 2026-09-09

Citation lineage

Review the prior work and downstream research connected to this paper.

Action-conditioned latent world models predict future visual representations, enabling zero-shot goal-conditioned robot planning and control. However, their predictions for fine-grained spatial and rotational actions are unreliable for full 7-DoF end-effector control. To address this gap, we introduce DUET-DINO, a simultaneous cross-view latent world model that jointly learns action-conditioned predictions from stat…

View the original paper on arXiv