For years the processor at the heart of an AI server was almost an afterthought, a traffic controller feeding data to the graphics chips that did the real work. At the Hot Chips 2026 conference on August 24, Nvidia argued that this arrangement is over. The company laid out the internal design of Vera, the custom CPU that will anchor its next generation of AI systems, and the choices behind it say a lot about where the industry thinks computing is heading.

Vera carries 88 cores built on a new in-house design Nvidia calls Olympus, its first fully custom Arm server core rather than a licensed one. The interesting part is what Nvidia optimised for. Instead of chasing the highest possible core count, the team leaned on instructions per cycle and predictability, pairing the cores with a large 164 MB L3 cache and eight 128-bit LPDDR5X memory controllers. The company said it picked that memory because it was the most energy efficient option, a telling priority in a business where power, not silicon, is fast becoming the binding constraint.

Designed for interruption, not just speed

The detail that stands out is how Vera handles multiple threads. Rather than letting several threads share the same execution resources, as most modern CPUs do, Nvidia uses what it calls statically partitioned spatial multithreading, carving out dedicated hardware for each thread. That reduces the noisy contention that shows up when a processor is juggling dozens of small, unpredictable tasks at once.

That is a deliberate bet on how AI is actually being used. A model answering a single question is a fairly orderly workload. An agent booking travel, calling tools, waiting on results and firing off new requests is anything but. Nvidia is designing for the second case. Its SPEC CPU 2026 figures showed Vera running roughly 1.8 times faster on agentic workloads than its predecessor, and in Linux kernel compilation the chip edged out AMD's EPYC 9655P.

The data centre as one machine

Vera does not ship alone. It slots into the Vera Rubin platform, Nvidia's successor to the current Grace Blackwell generation, where it connects to Rubin GPUs over the company's NVLink-C2C link and sits alongside new networking and data-processing chips. Nvidia claims that at high interactivity, the full platform can deliver up to 30 times the throughput of Grace Blackwell, though the real figure depends heavily on the workload.

Strip away the benchmark numbers and the argument underneath is consistent. Nvidia wants to design the CPU, the GPU, the networking and the software together, treating an entire rack as the unit of compute rather than a collection of parts. It is the same logic that has pushed rivals to build ever larger single systems, as Cerebras did by lashing three wafer-scale chips together, and it is why so much attention now falls on Nvidia's own earnings as a read on the whole AI trade.

Whether agents justify this much custom hardware is still an open question. But Nvidia is no longer treating the processor as plumbing, and that shift alone tells you how seriously the company is taking the idea that software will spend most of its time talking to other software.

Sources

  1. i. www.servethehome.com
  2. ii. cryptobriefing.com
  3. iii. www.igorslab.de

Commentarii · 0

Add · a · Comment