NVIDIA released Nemotron 3 Nano Omni on April 28, an open model designed to process video, audio, images, documents, and text through a single unified architecture. At 30 billion total parameters with only 3 billion active per inference pass, it runs on 25 gigabytes of RAM and delivers up to nine times the throughput of comparable open omni models on video and document tasks, according to NVIDIA's published benchmarks.

The architecture is a hybrid Mamba-Transformer Mixture-of-Experts design. The practical upshot is a model that can take a spoken question, a PDF, and a video clip as simultaneous inputs without routing them through separate AI components. NVIDIA is positioning it for agentic AI systems that need real-time multimodal reasoning.

On NVIDIA's benchmark suite, Nemotron 3 Nano Omni tops six leaderboards covering complex document intelligence, video understanding, and audio comprehension. It is available now on Hugging Face, OpenRouter, and build.nvidia.com as an NVIDIA NIM microservice.

The open model angle

The model is open and free to run. That sets it apart from GPT-4o and Gemini 1.5 Pro, which offer similar multimodal capabilities but only through paid API calls. For organizations that want a capable voice-and-vision model on their own hardware, without routing sensitive data through external endpoints, Nemotron 3 Nano Omni is now the most capable publicly available option at this weight class, at least by NVIDIA's own numbers.

Early adopters include Palantir, Foxconn, and H Company. Dell, Oracle, DocuSign, and Infosys are evaluating it. That list suggests interest well beyond pure AI research teams, into industries where on-premise deployment matters for compliance or data residency reasons.

The efficiency argument

Running on 25GB of RAM puts the model within reach of workstation-class hardware. A machine a single engineer can have on their desk. That is a meaningful threshold: hospital imaging departments, legal firms, and defense contractors can now deploy a capable multimodal model without cloud dependency, data egress costs, or the latency that comes with remote inference.

NVIDIA is using this release to argue that open, efficient models are catching up to proprietary frontier systems on tasks that matter in production. The benchmarks support that argument to a degree, though benchmark-to-real-world comparisons always deserve some scrutiny.

Sources

  1. i. blogs.nvidia.com
  2. ii. developer.nvidia.com
  3. iii. huggingface.co
  4. iv. www.hpcwire.com
  5. v. dataconomy.com

Commentarii · 0

Add · a · Comment