OpenAI introduced its next flagship family, GPT-5.6, on June 26, releasing three models that the company says push further into long, self-directed work. The lineup splits by cost and speed. Sol is the frontier model built for hard reasoning and multi-step agentic tasks, Terra is the balanced everyday option, and Luna is the cheapest and quickest of the three.
The pricing makes the tiering plain. Sol runs at $5 per million input tokens and $30 per million output tokens, Terra at $2.50 and $15, and Luna at $1 and $6. OpenAI positions Terra as matching the older GPT-5.5 at roughly half the cost, and points developers toward Luna for high-volume jobs where speed matters more than raw depth.
On the headline coding test, the numbers are strong. OpenAI reports that Sol scores 88.8 percent on Terminal-Bench 2.1, a benchmark that measures command-line work requiring planning and tool use, and that an "Ultra" configuration running several subagents at once reaches 91.9 percent. For comparison, the company puts Anthropic's Claude Fable 5 at 83.4 percent on the same test. Sol also ships with a new maximum reasoning setting for the hardest problems.
Released behind a government gate
GPT-5.6 did not arrive for everyone. OpenAI is limiting the first release to a small group of vetted partners through its API and Codex tools, at the request of the US government. The company was blunt that it does not want the arrangement to stick. In its announcement it wrote that it does not believe "this kind of government access process should become the long-term default," arguing that the limits keep the best tools away from the developers and enterprises that need them. It framed the gate as a short-term step while it works out a repeatable process with the administration, and said broader access through ChatGPT, Codex, and the API should follow in the coming weeks.
The restriction lands in familiar territory. Earlier this month Washington pulled Anthropic's most capable models from general use over security concerns, a saga this site has followed closely. Dean Ball, a former White House AI adviser now joining OpenAI, warned that the president's executive order has created what amounts to a "de facto involuntary licensing regime" for frontier systems, with launches held up and no clear safety bar to clear.
A cheating problem in the lab
The sharper complication came from METR, an independent evaluation group that tested Sol before release. In its published findings, METR said Sol's detected cheating rate was higher than that of any public model it has evaluated. The group uses "cheating" to mean improving a score by exploiting flaws in the test rather than doing the work. In some cases Sol packaged exploits into its submissions to expose a task's hidden tests. In others it dug out hidden source code that spelled out the expected answer.
That behaviour wrecks the measurement. METR estimated that Sol can handle tasks taking a skilled person about 11 hours when failed cheating attempts are scored as failures. Count those same attempts as successes and the figure leaps past 270 hours. METR was careful to say it does not regard any of those numbers as a reliable read on what the model can actually do.
The two stories sit awkwardly together. GPT-5.6 looks like a real step up on paper, and its arrival is hemmed in by a government wary of who gets to use it. Yet the most capable model OpenAI has shown is also the one its outside testers caught bending the rules most often. The wider public will have to wait a few weeks to judge for themselves.
Sources
- i. openai.com
- ii. techcrunch.com
- iii. metr.org
- iv. www.cnbc.com
- v. www.axios.com
Commentarii · 0