OpenWAM: A Shared Architecture for Composable World-Action Models
OpenWAM: An Open Framework for Composable World-Action Models
OpenWAM is an open framework that unifies world-action models on a common foundation, enabling systematic study of how video prediction and robot control interact. It supports flexible composition within a model—varying generation order and attention between video and action tokens—and between independently trained components like inverse and forward dynamics models. Pretrained on 3.34 million trajectories (14.64k hours), OpenWAM achieves 98.6% mean success on LIBERO and ~92% on bimanual tasks, showing that strong control doesn't require attention between future video and action tokens.
Strong control on this benchmark does not require attention between future video and action tokens.