Google has released Gemini 4, the first model in a new frontier generation, moving it out of testing and into a public research preview and a paid API. The launch, reported on September 30, lands less than a week after Google confirmed the model had entered post-training, and it reopens a close contest at the top of the field with Anthropic and OpenAI.
The headline figure is accuracy. According to independent testing relayed by The Neuron, Gemini 4 returned a hallucination rate of about 15 percent on a factual-recall evaluation, against 51 to 54 percent for rival frontier models. The model also took the top spot on the public Text Arena leaderboard, where human raters compare anonymized responses side by side.
What is actually shipping
Gemini 4 carries a one-million-token context window, enough to hold a long book or a sizable code base in a single prompt. Early API pricing is set at roughly $2 per million input tokens and $10 per million output tokens, with reports that both figures will rise once the preview period ends. For now the model is reachable through a limited research preview alongside the paid interface, so access is capped rather than open to everyone.
A note of caution is worth keeping. Some of the circulating specifications trace back to test checkpoints that appeared on public leaderboards under decoy names earlier in September, and a few of the more eye-catching numbers have not been confirmed by Google directly. The accuracy and pricing figures above come from early testers and trade coverage rather than a formal model card, so they may shift as the preview widens.
The race tightens again
The timing matters. Gemini 4 follows Anthropic's Opus 5.5 and OpenAI's recent GPT-6 releases, and it answers the question raised when Google entered post-training only days ago: how quickly could it close the gap. If the accuracy numbers hold up under outside scrutiny, Google has a real claim to the front of the pack on one axis that enterprise buyers care about, which is how often the model simply gets its facts wrong.
For developers, the practical story is the context window and the price. A million tokens of working memory changes what a single call can do, and a $2 input rate undercuts several competitors at the frontier tier. Whether that holds after the preview is the open question.
Sources
- i. theneuron.ai
- ii. llm-stats.com
- iii. 9to5google.com
Commentarii · 0