Anthropic released Claude Haiku 5.5 on October 7th, calling it the cheapest, fastest, and most capable small model it has shipped. The claim is easy to make and hard to verify, so the useful part is the gap between this model and the one it replaces. On the company's own figures, that gap is large.
Haiku is the budget tier, the model you reach for when a job runs millions of times a day and every fraction of a cent matters. Summaries, classification, database lookups, the quiet plumbing of a product rather than its headline feature. Until now that meant accepting a model that was fast and cheap but not especially bright. Haiku 5.5 changes the arithmetic.
On Anthropic's published benchmarks, the new model scores 1620 Elo on the GDPval agentic test against 735 for Haiku 4.5, the version from a year ago. On Terminal-Bench, a measure of how well a model drives a command line, Haiku 4.5 scored a flat zero; Haiku 5.5 scores 39.2 percent. On OSWorld, a test of controlling a computer, it climbs from 15.7 to 72.4 percent. These are not incremental gains. They are the difference between a model that could not do a task at all and one that often can.
A dial for how hard it thinks
The more interesting addition is a setting. Haiku 5.5 is the first model in its class with an adjustable effort control, the same Low-to-Max dial Anthropic offers on its larger Sonnet 5.5. Turn it down and the model answers quickly and cheaply. Turn it up and it spends more time, and more of your money, reasoning through a harder problem. That lets one small model stretch across jobs that used to need two.
Pricing is where the launch earns its billing. Haiku 5.5 starts at ten cents per million input tokens and fifty cents per million output, which Anthropic says works out to roughly 75 percent cheaper than Haiku 4.5 on an average workload, and as much as 90 percent cheaper on shorter requests. The company also quietly halved the cache-read price on Sonnet 5.5, which it says trims most agentic workloads by about a fifth.
Early customers describe the kind of results that matter at scale rather than on a leaderboard. Asana reported task completions running more than 30 percent faster. Box said the model scored eleven points higher than its predecessor at about half the latency. HubSpot put it at 92.8 percent across its customer-relationship test suite.
The safety notes
Anthropic says Haiku 5.5's alignment testing shows fewer misaligned behaviours than the older model and less willingness to go along with obvious misuse. Its cybersecurity safeguards sit tighter than Haiku 4.5's, though looser than Sonnet 5.5's, and the company says they still refuse attacker-oriented work. Its biosecurity protections match the heavier models in the lineup.
The model is available now through the Claude API and on Amazon Web Services, Google Cloud, and Microsoft Azure, which is the same everywhere-at-once distribution that has become routine for these launches. For most developers the headline is simpler than any benchmark: the cheap model is no longer the dumb one. That has been true of the frontier tier for a while. It is newly true at the bottom of the menu, where the bulk of the actual work gets done.
Sources
- i. www.anthropic.com
- ii. platform.claude.com
- iii. llm-stats.com
Commentarii · 0