A team from Stanford and Caltech has wired a frontier language model directly into a walking humanoid robot, and the result cleaned up a kitchen it had never seen before. The project, called HomeBody and published on Saturday, is notable less for what the robot did than for what the researchers left out.

No policy in the middle

Most robots that run on AI use a two-part stack. A high-level model decides what to do, and a separately trained control policy, learned from mountains of demonstration data, translates that intent into motor commands. HomeBody drops the second part. It gives OpenAI's GPT Astra direct command of five composable skills built into a Unitree G1 humanoid, and lets the model call them itself. There is no learned action policy sitting between the model's decision and the robot's joints.

The robot starts by wandering an unfamiliar room. As it moves, it builds what the team calls a Real2Sim digital twin, stitched together from its own camera feed, its sense of where its body is, and the paths it has walked. That twin becomes the robot's spatial memory. Ask it later to fetch something it saw earlier, and it can reason over the map it made rather than starting blind.

In a test kitchen with no environment-specific training, a G1 guided by GPT Astra tidied up across the room and retrieved an object it had noticed earlier from a vague, underspecified request. The point the researchers are making is that a good enough general model, handed the right primitives, can skip a training step that the field had treated as mandatory.

Real limits, honestly stated

To their credit, the authors do not oversell it. They list the friction plainly. Building the Real2Sim twin takes setup time. Every API call to the model costs money, and the bill adds up over a long task. There is noticeable latency between skills while the model thinks. And during extended runs, the robot's finger servos overheated, a very physical reminder that the hard parts of robotics do not disappear just because the brain got smarter.

HomeBody fits a pattern from the past few weeks, in which capable general models keep absorbing jobs that used to need bespoke systems. MIT recently showed an insect-sized flying robot with an AI controller nimble enough to do somersaults, and Alphabet's Intrinsic open-sourced the core of its robot software stack. The common thread is a shift in where the intelligence lives, away from narrow trained controllers and toward large models doing the reasoning.

A tidy kitchen is not a housekeeper, and a research demo is not a product. Servos still overheat, latency still stutters, and the cost of running a frontier model on every step is real. What HomeBody shows is a direction, not an arrival. The direction is worth watching.

Sources

  1. i. tml.stanford.edu
  2. ii. aiweekly.co

Commentarii · 0

Add · a · Comment