NVIDIA is better known for selling the chips that train everyone else's models than for shipping its own. This month it did both. On June 4 the company released the weights for Nemotron 3 Ultra, a 550-billion-parameter model it calls the most capable open weights system built in the United States.

The model was announced June 1 at Jensen Huang's Computex keynote, with the weights following three days later. What arrived was unusually complete. NVIDIA published the model, its training data, and the recipes used to build it under the Linux Foundation's permissive OpenMDW-1.1 license, according to MarkTechPost. That is a fuller release than most open models, which tend to ship weights and little else.

Built for agents that keep going

Nemotron 3 Ultra is a mixture-of-experts model: 550 billion parameters in total, but only 55 billion active for any given token. That design keeps the running cost closer to a 55B model while drawing on the knowledge of a much larger one. It uses a hybrid Mamba-Attention architecture, a combination meant to handle long, drawn-out tasks without the memory blowup that pure attention suffers over long inputs. The context window runs to a million tokens.

The target is clear from the design. Long-running agents, the kind that work through a problem for hours rather than answering a single prompt, need exactly this mix of cheap inference and long memory. NVIDIA reports up to roughly six times the inference throughput of comparable open models at similar accuracy, citing a 5.9x edge over GLM-5.1 in its own tests, per Artificial Analysis.

One number stands out. The model posted the highest non-hallucination score in its comparison group, 78.7 on the AA-Omniscience benchmark, which measures how well a model knows the limits of what it knows. For agent work, where a confident wrong answer can derail an hour of automated effort, that matters as much as raw speed.

Best in America, behind China

The honest framing came from the independent reviewers rather than NVIDIA's marketing. Nemotron 3 Ultra is the strongest open weights model to come out of the United States, but it still trails the best open models from China, including the DeepSeek and GLM families, on several benchmarks. The gap in open AI between the two countries has not closed. It has a new American entry that narrows it.

NVIDIA paired the release with wide availability. The model went live the same day across more than twenty-five platforms, including Hugging Face, OpenRouter, Perplexity, Together AI, NVIDIA's own NIM service, and Amazon SageMaker JumpStart. Anyone wanting to test it against China's open models, or against the recently funded DeepSeek, can do so today.

It also extends a busy stretch for NVIDIA beyond silicon. The Nemotron release lands days after the company opened Cosmos 3, its foundation model for the physical world. Two open models in a week, from the company that sells the hardware to run them, is a strategy worth watching.

Sources

  1. i. www.marktechpost.com
  2. ii. artificialanalysis.ai
  3. iii. chatforest.com
  4. iv. www.digitalapplied.com

Commentarii · 0

Add · a · Comment