AMD said on August 6th, after the market closed, that it is acquiring Taalas, a Toronto startup with an unusual idea about how to run a large language model quickly: stop moving the model at all. Rather than streaming billions of weights out of high bandwidth memory on every token, Taalas burns those weights permanently into the transistors of a custom chip. The arithmetic still happens, but the slow part, the constant fetch from memory, largely goes away.
That slow part is the whole game in inference. Modern accelerators spend much of their time and power waiting on memory rather than doing math, which is why a new class of high bandwidth flash has been racing to widen that pipe. Taalas takes the opposite route. If the model never leaves the silicon, the pipe barely matters.
What the chip actually does
Taalas calls its demonstrator the HC1. Built on TSMC's 6nm process, it is a large piece of silicon, roughly 815 square millimetres carrying about 53 billion transistors, and in the company's own testing it served Meta's Llama 3.1 8B model at close to 17,000 tokens per second for a single user. Taalas measured that against Nvidia's H200 and B200, along with Groq, SambaNova and Cerebras, so the figures deserve the usual caution that comes with a vendor's own benchmarks. The trade is blunt: each chip is welded to one model. Change the model in a meaningful way and you need new silicon.
AMD is betting that this trade makes sense for what it calls stable, high volume inference, the workloads where a popular model serves the same kind of request millions of times a day. The company plans to fold the technology into its Helios rack scale systems next to its Instinct GPUs and EPYC processors, with its ROCm software tying the parts together.
The wider race
AMD did not disclose a price, and it expects the deal to close in the fourth quarter. Taalas was founded in 2023 and had raised around 219 million dollars, according to reporting by The Register and ServeTheHome.
The move fits a larger shift in the industry. As the cost of running models has climbed and power has become the binding constraint on data centres, the biggest buyers have started designing hardware that fits their models rather than the other way round. Anthropic has begun building its own silicon team, and now AMD is buying its way into model specific chips. Nvidia still owns the training market. Inference, where most of the money will eventually be spent, is starting to look like a genuine contest.
Sources
- i. www.theregister.com
- ii. www.servethehome.com
- iii. qz.com
Commentarii · 0