On August 6, AMD announced a definitive agreement to acquire Taalas — a small but technically remarkable startup out of Toronto. And the idea behind it is radical enough to be worth a pause.
Models made of transistors, not memory
Most AI chips today work the same way: a model’s weights live in fast memory (HBM) and get shuttled back and forth on every calculation. That constant data traffic is one of the biggest bottlenecks — it costs time and, above all, energy.
Taalas flips this around. Founded in 2023, the company builds what it calls model-specific integrated circuits: instead of keeping the model in memory, Taalas casts the architecture and the trained weights directly into the chip’s transistors. The model becomes the chip. The price is flexibility — a chip like that can only ever run that one model. The reward is speed: a demo chip reached over 16,000 tokens per second per user running Llama 3.1-8B. That’s many times what comparable general-purpose hardware delivers.
The third acquisition in nine months
For AMD, Taalas is already the third AI acquisition in nine months — after MK1 in November and memory startup Mext in June. Terms weren’t disclosed. The direction, though, is unmistakable: AMD wants to gear up for inference, meaning the everyday running of AI models. And that everyday running is becoming the industry’s single biggest cost, because far more people use models than train them.
My take
Why does this matter to us here? Because it’s the same movement we saw at Anthropic just a few days ago — they spun up their own chip team to make Claude run cheaper and faster. AMD and Anthropic are pulling on the same rope from different ends: the era of just throwing everything at expensive general-purpose GPUs is slowly ending.
For us as users, that’s good news, even if it arrives with a delay. Cheaper inference ultimately means more tokens for your money, faster agents, fewer compromises. And the fact that a GPU giant like AMD is betting on “model-in-silicon” of all things is the clearest signal yet that specialized hardware is no longer a niche experiment. The open question is whether Taalas’ approach can survive in a world where new models ship every week — a hardwired model, after all, goes stale faster than a flexible chip.
Sources: AMD press release, SiliconANGLE, The Register