Watch the right thirty-second clip and the future looks settled. A humanoid robot folds laundry, loads a dishwasher, hands you a drink. The forecasts around those clips are just as confident. Elon Musk has called Tesla's Optimus "probably the biggest product ever" and suggested it could reach the public by the end of 2027. The investor Marc Andreessen has mused that robotics could become the "biggest industry in the history of the planet." Nvidia's Jensen Huang said this year that humanoids would reach human-level ability. Morgan Stanley has sketched a world with roughly a billion human-like robots by 2050.
The people who build these machines tell a slower story. A recent survey of the field by MIT Technology Review lays out why the recipe that transformed chatbots may not carry over to the physical world, and why a general-purpose robot in the average home is likely still about ten years away.
The progress is real
None of this is a case that robotics is stuck. Google DeepMind's Gemini Robotics, running on a two-armed research rig, can pack a simple lunch. Physical Intelligence has shown its models doing things they were barely trained for, such as loading a sweet potato into an air fryer, an early sign of what researchers call compositional generalisation. The whole field has shifted from hand-coded rules toward vision-language-action models, trained by having humans puppeteer the robot through a task again and again. We have watched a Stanford robot take its orders straight from a language model. The demos are not fake. They are just narrower than they look.
Why the recipe does not transfer
The chatbot boom ran on an ocean of text scraped from the internet. There is no comparable ocean of physical experience for a robot to learn from, and Jonathan Hurst of Agility Robotics calls the belief that more data alone will fix this "a fundamentally flawed premise." Push a vision-language-action robot slightly outside its training and it tends to fail. Edward Johns of Imperial College London says today's models can manage only "a few things here and a few things there." And the bar in the physical world is unforgiving. As Marc Raibert, who founded Boston Dynamics, puts it, "70% success is like it doesn't work." A chatbot that is right 70% of the time is useful. A robot that drops your plate three times in ten is not.
Then there is the matter of the demos themselves. Many of the most jaw-dropping videos are teleoperated or scripted. Real deployments today mostly involve a machine shuttling bins across a warehouse floor. Even Yann LeCun, no pessimist about AI, says current approaches "do not work" for gathering the kind of data robots need, and Fei-Fei Li, whose spatial-intelligence startup was just bought by AMD, calls the field "nascent."
What world models can and cannot do yet
The great hope is a technology called world models, AI trained on video and 3D and sensor data to predict how a physical action will play out. Done well, they could let a robot rehearse in simulation, cheaply and safely, and anticipate what happens before it acts. They are genuinely promising. They are also early-stage research, not a shortcut to a general-purpose robot you can buy next year.
So the myth is not that humanoid robots are coming. They are, in some form, and the economics will keep catching up, as we have written before. The myth is the timeline, the quiet assumption that because machines learned to write and draw so fast, they will learn to move through our cluttered, slippery, unpredictable homes just as quickly. The body, it turns out, is a harder problem than the sentence.
Commentarii · 0