For most of the past year the smart-sounding advice inside big companies was to use as much AI as humanly possible. Burn the tokens. Run the agent again. There was even a word for it, "tokenmaxxing," and at some firms it became a kind of sport. Meta built an internal dashboard that turned into a leaderboard, with staff competing to see who could push the most usage. That era is ending, and the reason is the same one that ends most booms. The bills arrived.
The numbers behind the splurge are striking. Visa encouraged its engineers to lean hard on AI and ran up roughly 1.9 trillion tokens in March alone, with some engineers calling on Claude around 51,000 times a day, according to reporting from CNBC on 26 June. Uber said it had spent its entire 2026 AI budget in the first four months of the year, with its operating chief admitting the cost was getting "harder to justify." When a company torches a full-year line item by April, somebody upstairs starts asking what came back for the money.
From growth to results
The shift now under way is not a retreat from AI. It is a move from spending for its own sake to spending that has to show its work. As CBC reported, firms that once celebrated raw consumption are quietly shuttering the leaderboards and asking a blunter question: did the work actually get done, and did it get done cheaper than a human would have done it? Meta's competitive dashboard is gone. The language in boardrooms has changed from adoption to efficiency.
That sounds like common sense, and it is, but it lands hard on the business models of the companies that sell the tokens. OpenAI and Anthropic built their pricing around the assumption that usage would keep climbing and customers would keep paying premium rates for premium models. A customer base that suddenly cares about cost per task is a very different proposition.
The cheaper model in the room
The squeeze is coming from China. DeepSeek has made a permanent 75 percent price cut on its flagship V4 Pro model, a move VentureBeat described as a direct assault on the economics of Silicon Valley's frontier labs. By the company's own accounting the model is roughly seven times cheaper on inputs and seventeen times cheaper on outputs than Anthropic's Claude Sonnet or OpenAI's mid-tier GPT-5.5. For a finance team staring at a Visa-sized token bill, that math is hard to ignore.
Some are already acting on it. The chief executive of the AI startup Lindy said he had moved all of his company's traffic off Claude and onto DeepSeek, with costs falling sharply as a result. One small company switching providers is not a trend on its own, but it is the kind of decision that, multiplied across thousands of buyers, reshapes a market.
Why the labs noticed
The American labs have not missed the message. OpenAI's newest release, the GPT-5.6 family, deliberately includes cheaper tiers, with its Luna model pitched at a fraction of the flagship's price. The frontier is no longer the only thing that sells. After a year in which OpenAI alone lost tens of billions of dollars chasing scale, the customers have started to set the terms, and the terms are getting cheaper. The interesting question for the rest of 2026 is whether the labs can make efficiency pay before someone undercuts them again.
Sources
- i. www.cnbc.com
- ii. www.cbc.ca
- iii. www.sdxcentral.com
- iv. venturebeat.com
- v. www.marketscale.com
Commentarii · 0