arXiv 2607.19343

Masked Visual Actions for Unified World Modeling

By Hadi Alzayer, Wenlong Huang, et al.

Published 2026-07-21

Wiki summary

Explore the paper's summary, context, and related research on Papiers.

Video models absorb rich priors over how the visual world moves, interacts, and responds to contact, making them promising substrates for robotic world modeling. The central challenge is how to communicate action to such models in a form aligned with the visual space in which they learned these interaction priors, yet still grounded in physical manipulation. We introduce Masked Visual Actions, a pixel-space control…

View the original paper on arXiv