The code-editor startup Cursor released Composer 2.5 on May 18, a coding model the company says matches the headline scores of Claude Opus 4.7 and OpenAI's GPT-5.5 at roughly one tenth the price per token. The model is built on the same open-source checkpoint as its predecessor, Moonshot's Kimi K2.5, and the release sharpens a question the industry has been circling for a year: how much of the work that frontier labs have priced into the high tiers can be done by a fine-tune of an open base model that anyone can download.
The headline benchmark numbers, reported in TestingCatalog, are 79.8% on SWE-Bench Multilingual and 63.2% on CursorBench v3.1. Those land within a few points of the frontier closed models on the same harness. Cursor lists Composer 2.5 at $0.50 per million input tokens and $2.50 per million output for its standard variant, with a faster default tier at $3.00 in and $15.00 out. Even the more expensive tier sits well below Opus or GPT-5.5 pricing.
How the gap closed
Cursor attributes the jump from Composer 2 to 25x more synthetic task data and what the company calls targeted textual feedback, a reinforcement-style technique that lets trainers correct the model at the exact step in a trajectory where it went off the rails rather than rating only the final answer. Gigazine notes that the result is a model substantially better at sustained long-running tasks than the previous version, which is precisely where coding agents tend to fall apart.
None of this is alchemy. The Kimi K2.5 base remains the work of Moonshot's Beijing team, and Cursor's contribution is at the post-training layer. That is also where the open-source story gets interesting. A product company has now made a fine-tune of a Chinese-trained open-source model competitive with the most expensive American closed models on the workload its users actually care about. The pattern is the same one that played out around Llama derivatives last year, only with more polish at the application layer.
What it means for closed-model pricing
For working developers, the immediate question is whether Composer 2.5 holds up outside the benchmarks. Coding tasks have a notoriously long tail, and SWE-Bench scores have a habit of overestimating real-world reliability. Cursor's own framing, that the model is "more pleasant to collaborate with," gestures at the parts of agent behavior that benchmarks do not capture. Early reviews from Winbuzzer and Kingy AI are cautiously positive on the long-horizon improvements.
The strategic question is louder. If a $2.50-output-token model can hold pace with a $40 one on the workloads that fill most developers' days, the economic case for routing every coding query to frontier infrastructure starts to look thin. Anthropic and OpenAI both have answers, ranging from caching to specialised coding tiers, and the gap will narrow in places. The story to watch is whether Cursor's customers find Composer 2.5 good enough to use as the default, with a frontier model called only for the hard problems. That is the routing pattern that puts real pressure on closed-model pricing.
Cursor also flagged that it is training a larger model from scratch with SpaceXAI, using ten times the compute that went into Composer 2.5. That puts the company on the slow road toward owning the base model rather than the fine-tune. Whether the bet pays off will not be clear for a year or more, but the signal is unmistakable: a tools company is no longer content to sit at the application layer.
Sources
- i. cursor.com
- ii. www.testingcatalog.com
- iii. gigazine.net
- iv. winbuzzer.com
- v. kingy.ai
Commentarii · 0