Masked Visual Actions for Unified World Modeling
By Hadi Alzayer, Wenlong Huang, Haonan Chen, Christopher Luey, Lvmin Zhang, Maneesh Agrawala, Gordon Wetzstein, Li Fei-Fei, Yilun Du, Jiajun Wu, Jia-Bin Huang
Masked Visual Actions (MVA) introduces a pixel-space control interface for video world models, expressing action as a partially revealed trajectory of an arbitrary entity. Revealing robot motion makes the model predict scene response (forward dynamics); revealing desired object motion makes it recover robot behavior (inverse dynamics). Fine-tuned with only 15 hours of manipulation data, it unifies forward/inverse modeling.