Mistral released Medium 3 on April 9, and the benchmark numbers are worth attention. The model performs at or above 90% of Claude Sonnet 3.7 across standard evaluations, with particular strength in coding and STEM tasks. The price is $0.40 per million input tokens and $2.00 per million output tokens. That is well below what comparable performance costs from OpenAI or Anthropic's commercial offerings.
The context window runs to 131,000 tokens. For most production use cases, that is enough to handle long documents, extended conversations, or complex code analysis without hitting limits.
Where it runs
Medium 3 is an API-first model. It is available on Mistral's La Plateforme and on Amazon SageMaker at launch, with planned availability on IBM WatsonX, NVIDIA NIM, Azure AI Foundry, and Google Cloud Vertex. For teams that want to self-host, the model can run on a four-GPU setup, though Mistral positions it primarily as a hosted inference product.
That is worth clarifying. Mistral has previously released genuinely open-weight models, like Mistral 7B and Mixtral 8x7B, that anyone could download and redistribute freely. Medium 3 is not in that category. You can access it through multiple cloud providers and self-host with sufficient hardware, but it is not a freely redistributable open-weight model. The company's open-source commitments have become more selective over time.
Where it fits in the market
Medium 3 lands in a space that has been getting crowded: the tier below the largest frontier models, where performance is strong enough for most enterprise applications but cost is significantly lower. Google's Gemma 4 and NVIDIA's Nemotron 3 are competing in roughly the same territory, as is Alibaba's Qwen3.6 series, which released an Apache 2.0 open-weight version on April 16.
The claim that Medium 3 matches 90% of Claude Sonnet 3.7 on benchmarks should be taken with the usual caveats. Benchmark performance and real-world production performance are not the same thing, and models that do well on standard evaluations sometimes disappoint on niche tasks. That said, the pricing is aggressive enough that teams looking to reduce inference costs have a concrete reason to run their own tests.
Sources
- i. mistral.ai
- ii. blog.galaxy.ai
Commentarii · 0