DreamZero is a World Action Model that learns robot control by jointly predicting future world states and actions using video diffusion. It achieves over 2× better generalization to new tasks and environments compared to Vision-Language-Action models, and can adapt to new robot embodiments with just 30 minutes of play data while maintaining zero-shot generalization capabilities.
Dream Machines shares empirical results from fine-tuning Physical Intelligence's π0.5 vision-language-action model on a real manufacturing task: transferring actuators from cardboard boxes into trays using a bimanual robot. The study documents findings from tuning on data collected from a German manufacturer, providing practical recommendations for parameter adjustment and data collection to help others deploy the model effectively.