Robot foundation models learn reusable representations of perception, language, motion, and physical interaction from mixtures of robot demonstrations, human demonstrations, video, or simulation. Some produce actions directly, while others predict future states or provide features to a separate controller. The label does not guarantee cross-embodiment deployment: teams must still verify which bodies, tasks, sensors, and post-training stages a specific model actually supports.