DeepSeek dropped its V4 model family on April 24, ending months of speculation about what the Chinese lab would release for the anniversary of the original V3 moment that sent Nvidia stock down 17% in a single trading day. The new family comes in two sizes, V4-Pro and V4-Flash, and both push past the headline numbers that made the company a global story in early 2025.

V4-Pro is a 1.6 trillion parameter Mixture of Experts model with 49 billion active parameters per token. V4-Flash sits at 284 billion total and 13 billion active. Both ship with a one million token context window, putting them in the same territory as Gemini 3.1 Pro and the just-released GPT-5.5 for handling whole codebases or long legal documents in a single prompt. DeepSeek's release notes describe a hybrid attention scheme combining what the team calls Compressed Sparse Attention and Heavily Compressed Attention, which the company says cuts inference compute to 27% of V3.2 and KV cache memory to 10% of the prior generation.

Where it sits on the benchmarks

Independent benchmarking from Artificial Analysis places V4-Pro at the top of the open-weights pack on agentic coding and within striking distance of GPT-5.5 and Gemini 3.1 Pro on world-knowledge tests. TechCrunch's coverage noted the company is positioning the release as a preview, with the full V4 launch expected later in the quarter.

The pricing is the part that will make CFOs read twice. V4-Flash runs $0.14 per million input tokens and $0.28 per million output tokens. V4-Pro is $1.74 in and $3.48 out. By comparison, GPT-5.5 lists at $5 per million input and $30 per million output, which means a long-context coding workload that would cost $30 on OpenAI runs about $3.50 on V4-Pro and 28 cents on Flash.

Why it matters beyond the leaderboard

The original DeepSeek V3 release in January 2025 forced a public reckoning over compute moats. MIT Technology Review's analysis argues V4 sharpens the same question with cleaner data. The lab claims its training run remains in the tens of millions of dollars rather than the hundreds of millions required by US frontier labs, though those numbers are difficult to audit from outside.

Open weights are the second piece. V4-Pro and V4-Flash are live on Hugging Face under a permissive license, which means any developer with enough GPU rental budget can run them locally. That matters for enterprises in regulated sectors who cannot send data to a third-party API, and it matters for researchers who want to study how a frontier model actually behaves rather than what an API endpoint chooses to return.

The agentic angle

DeepSeek built the V4 series with agent workflows in mind. Both models are tuned for long tool-use chains, the kind where a coding agent has to read documentation, edit files, run tests, and recover from errors over hundreds of steps. The release notes claim state-of-the-art performance among open models on agentic coding benchmarks, a claim Artificial Analysis broadly corroborated.

That puts the V4 family in direct competition with the wave of releases that landed in the same week. Moonshot AI's Kimi K2.6, OpenAI's GPT-5.5, and xAI's Grok 4.3 all shipped within four days of each other, all with agentic coding as the headline pitch. The current quarter is on track to be the most concentrated frontier-model release window the field has yet seen.

Open questions

DeepSeek has not yet published a full technical report for V4, only the release notes and a model card. The training data composition, the exact compute budget, and the post-training recipe all remain undisclosed. The company has historically published more detail in follow-up papers, but the gap between launch and paper has grown with each release. For now, the public story is the benchmarks and the price tag.

Sources

  1. i. api-docs.deepseek.com
  2. ii. techcrunch.com
  3. iii. www.bloomberg.com
  4. iv. artificialanalysis.ai
  5. v. www.technologyreview.com
  6. vi. huggingface.co
  7. vii. www.cnn.com

Commentarii · 0

Add · a · Comment