Anthropic released Claude Sonnet 5 on 30 June, and the pitch is straightforward: most of what its flagship Opus model can do, at a fraction of the cost. The company calls it the most agentic Sonnet yet, built to make plans, drive browsers and terminals, and see long multi-step jobs through on its own. Work that needed a larger, pricier model a few months ago now runs on the midsize tier.
The pricing is the headline for anyone building on the API. Sonnet 5 launches at $2 per million input tokens and $10 per million output tokens, an introductory rate that holds through 31 August before settling to the standard $3 and $15. TechCrunch framed it as a cheaper way to run agents, which is the market Anthropic is going after as customers leave models to churn through tasks unattended for hours at a stretch.
How close to Opus, really
Anthropic says Sonnet 5's agentic performance sits close to Opus 4.8, its most capable model, and on at least one knowledge-work benchmark Sonnet edges slightly ahead. Opus still holds the lead on the hardest problems, the subtle judgment calls and deep research where an extra margin of reasoning earns its keep. For the bulk of everyday agent work, though, the gap has narrowed to the point where price becomes the deciding factor.
The improvements over Sonnet 4.6, which shipped in February, cover reasoning, tool use, coding, and knowledge tasks. In Claude Code, Sonnet 5 becomes the default model, and it is now the default on Free and Pro plans as well. Max, Team, and Enterprise users get it too, and developers call it as claude-sonnet-5 on the API. It follows Anthropic's own research into who actually gets results from coding agents, which found the payoff comes less from raw model power than from the person steering it.
Safety built for agents that run alone
Letting a model act on its own raises the stakes on misbehaviour, and Anthropic leaned into that. Sonnet 5 shows a lower rate of what the company calls undesirable behaviours, including cooperation with misuse and deception, than its predecessor. It refuses malicious requests more reliably and holds up better against prompt-injection attempts, the trick where hidden instructions in a webpage or document try to hijack an agent mid-task. Hallucination and sycophancy rates are down, and cyber safeguards are on by default.
That last point lands in a pointed moment. Anthropic spent much of June with its flagship models offline after a government export ban, a saga set off by a jailbreak that turned safety guardrails into a liability. A midsize model that is genuinely harder to misuse, and cheap enough to deploy widely, is the kind of thing the company needs to be seen shipping right now.
The read on Sonnet 5 is less about a benchmark crown and more about where the floor now sits. Capable autonomous agents used to mean paying flagship rates. They no longer do, and that changes the arithmetic for anyone weighing whether to put a model to work while nobody is watching.
Sources
- i. www.anthropic.com
- ii. techcrunch.com
- iii. www.testingcatalog.com
- iv. www.techtimes.com
Commentarii · 0