The model has landed. On September 2 Google made Gemini 3.8 Flash generally available, confirming the coding-first release we reported earlier this week when it was still a set of leaks and staging pages. It is the company's fourth Flash-tier model in under four months, and it arrives barely three weeks after the last one.

Google is pitching 3.8 Flash as its most capable small model to date, tuned for long-horizon software engineering, autonomous agents and multi-step enterprise reasoning. The headline number is a jump on Terminal-Bench 2.1, a test of whether a model can work through real command-line tasks. Gemini 3.8 Flash scores 90.8 percent there, up from 81.6 percent for 3.7 Flash. On DeepSWE v1.1, a benchmark for sustained coding work, Google says the new Flash model beats most larger frontier systems. Its score on the independent Artificial Analysis Intelligence Index rose three points to 59.

Cheap now, pricier in January

The economics are the real story. Gemini 3.8 Flash ships at 75 cents per million input tokens and 3.75 dollars per million output tokens, the same introductory rate Google used for 3.7 Flash three weeks ago. Both figures are set to double on January 1, to 1.50 dollars and 7.50 dollars. That pattern, a low launch price followed by a scheduled increase once developers have built on the model, has become a familiar move across the industry. It slots the release into a week of aggressive pricing that also brought Anthropic's cheaper Fable 5.1.

The model handles text, images, audio, video and PDF input, with a one million token context window and 64,000 tokens of output. It is available through Google AI Studio and the Gemini Enterprise Agent Platform. Google released it alongside a locked-down variant, 3.8 Flash Cyber, aimed at security work.

Speed as a strategy

What stands out is the cadence. Four Flash models since June is a release schedule that would have seemed frantic a year ago, and it raises a fair question about whether each new version number marks genuine progress or incremental tuning. We looked at exactly that tension in a recent piece on whether a new model every few weeks means anything at all. On the benchmarks Google chose to publish, 3.8 Flash is a real step up from its predecessor. Whether developers feel the difference in daily use, rather than on the leaderboard, is the test that matters now.

Sources

  1. i. artificialanalysis.ai
  2. ii. www.theregister.com
  3. iii. 9to5google.com

Commentarii · 0

Add · a · Comment