Cerebras used its Supernova event in San Francisco on August 19 to launch the CS-4, the first system from the company to wire three of its dinner-plate-sized processors together in a single machine. The pitch is speed. Cerebras says the CS-4 runs inference, the stage where a trained model answers a query, up to 30 times faster than the GPU-based services most companies rent today.
The numbers the company put forward are large. Each WSE-3 Turbo processor delivers 250 petaFLOPs of compute, so three of them reach 750 petaFLOPs, paired with memory bandwidth measured in the hundreds of petabytes per second. On the open GPT-OSS-120B model, Cerebras reported more than 4,400 tokens per second for a single user, against roughly 350 on the fastest GPU inference service it tested. First shipments are due later this quarter.
There is a catch worth stating plainly, and The Next Web pointed it out: the chip inside the CS-4 is not new. The WSE-3 Turbo is the same wafer-scale part Cerebras has shipped before. What is new is the packaging, a server architecture the company calls Nexus that lets three wafers work in parallel as one. That is an engineering result rather than a silicon breakthrough, and it matters mostly because inference, not training, is where the day-to-day cost of running AI now sits.
Why inference is the battleground
Training a model is a one-time expense. Serving it happens every time someone asks it a question, and at the scale of billions of daily queries the bill never stops. That is the market Cerebras is aiming at, and it is Nvidia's to lose. Nvidia's GPUs still handle the overwhelming majority of inference workloads, and the company has been lining up half a trillion dollars in financing to keep that lead.
Cerebras is not the only firm trying to pry the door open. AMD bought a startup to etch models directly into silicon, and a London company raised $312 million to run inference on light rather than electricity. Each is betting that Nvidia's general-purpose hardware leaves room for a specialist that does one job faster or cheaper.
The honest read on the CS-4 is that speed claims from a vendor's own benchmarks are a starting point, not a verdict. Token-per-second figures depend heavily on the model, the batch size, and how the comparison is framed. What the launch does confirm is that the fight over inference has moved from labs to loading docks, and that customers now have real reasons to price out something other than a rack of GPUs. UBS, for one, told clients the launch strengthens the case that Cerebras architecture is genuinely differentiated for fast inference.
Whether that translates into orders is the question the next quarter will answer, when the first CS-4 machines actually ship.
Sources
- i. investors.cerebras.ai
- ii. thenextweb.com
- iii. dataconomy.com
- iv. www.domain-b.com
Commentarii · 0