Every time a major model slips its release date, the same verdict arrives within hours. AI has hit a wall. The scaling era is over. The bubble is deflating.
Google's Gemini 3.5 Pro has now missed its target date more than once after a full architectural rebuild, and the commentary followed the script. So it is worth asking what the claim actually means, and which parts of it hold up.
The part that is true
Traditional pretraining scaling really has slowed. The relationship that drove progress from roughly 2019 onward was straightforward: add more compute and more data, get a predictably better model. That relationship still holds mathematically, but it has run into two practical limits.
The first is diminishing returns near the ceiling. The jump from GPT-3.5 to GPT-4 felt enormous to users. The jump from GPT-4 to GPT-5 felt smaller, despite reportedly using an order of magnitude more compute. That is not a failure of the theory. It is what the theory predicts as you approach the top of the curve.
The second is data. High-quality training text is finite, and the industry has worked through a great deal of it. Ilya Sutskever put it plainly at NeurIPS in 2024, saying that pretraining as we know it will end.
So if the claim is "you can no longer buy large capability gains by making the pretraining run bigger," that is roughly correct and most researchers agree.
The part that is wrong
The claim usually being made is broader. It is that AI capability itself has plateaued. That does not survive contact with the last eighteen months.
What happened instead is that the field found a different axis. Test-time compute, sometimes called inference-time scaling, means letting a model think for longer while it answers rather than training it for longer beforehand. It follows its own power law, and it has produced the reasoning models that now sit at the top of every difficult benchmark.
This is a real shift rather than a marketing one. It moved the cost of intelligence from training to inference, which is why the industry's compute demand has gone up rather than down as pretraining gains flattened. A wall would look like falling demand for compute. That is not what anyone is observing.
Why delays get misread
Here is the thing that gets lost. Gemini 3.5 Pro was delayed because Google scrapped its base model and rebuilt it after finding structural problems, and reportedly because early testing showed it trailing competitors on reasoning and long-horizon tasks.
That is a company deciding not to ship something that is not good enough. Read one way it is evidence of difficulty. Read another way it is evidence of a competitive field where shipping a merely adequate frontier model is no longer viable. A few years ago it would have shipped.
Delays are noisy signals. A single company's schedule tells you about that company. It does not tell you about the trajectory of a field where, in the same month, several labs shipped major models within days of each other.
Where honest uncertainty sits
None of this means progress is guaranteed to continue at pace. The genuine open questions are whether test-time compute has its own ceiling, how far synthetic data can substitute for exhausted human text, and whether the next gains will be broad or narrow.
I lean toward thinking capability keeps improving in the near term, because the evidence points that way. But the rate is uncertain, and anyone telling you confidently that it either has stopped or cannot stop is working from conviction rather than data. The wall metaphor is the problem. It suggests a single barrier where what actually exists is a set of separate constraints, some of which have been routed around and some of which have not.
Sources
- i. www.techtimes.com
- ii. hai.stanford.edu
- iii. www.hec.edu
- iv. cameronrwolfe.substack.com
Commentarii · 0