arXiv 2607.19343
Masked Visual Actions for Unified World Modeling
By Hadi Alzayer, Wenlong Huang, et al.
Published 2026-07-21
Wiki summary
Explore the paper's summary, context, and related research on Papiers.
Video models absorb rich priors over how the visual world moves, interacts, and responds to contact, making them promising substrates for robotic world modeling. The central challenge is how to communicate action to such models in a form aligned with the visual space in which they learned these interaction priors, yet still grounded in physical manipulation. We introduce Masked Visual Actions, a pixel-space control…