Rho is a foundation model for vision-language-action robots that separates adaptation into two stages: first learning a specific robot's embodiment, then adapting to particular tasks. This two-stage approach reduces the finetuning data needed by half compared to baseline models, with midtrained variants matching or outperforming competing systems on three physical robots.