HomeBody is a humanoid robot that uses a vision-language model to reason about tasks and select skills for physical execution based on spatial targets. It coordinates multiple manipulation and navigation skills—picking, placing, and drawer opening—by passing structured commands that handle perception, planning, and motion execution without requiring the VLM to manage low-level implementation details.