AI Infrastructure

AMD Buys Taalas: Etching AI Models Straight Into Silicon

AMD agreed to buy Toronto startup Taalas, whose chips etch model weights into silicon: Llama 3.1 8B at nearly 17,000 tokens per second on a 6nm test chip.

AMD Buys Taalas: Etching AI Models Straight Into Silicon — article cover

On August 6, AMD announced a definitive agreement to acquire Taalas, a Toronto startup founded in 2023. Its approach is among the most radical in the industry: instead of general-purpose GPUs, it etches the weights of a specific model directly into silicon, trading flexibility for speed and cost by building the hardware around the model. Terms were not disclosed, and the deal is subject to regulatory approval; The Register reports an expected close in Q4. CNBC notes Taalas has raised $219 million in venture funding.

For AMD, it is another tile in a full-stack build-out that already includes Silo AI ($665 million) and ZT Systems ($4.9 billion), and the timing lands months after Nvidia’s roughly $20 billion Groq licensing deal in December 2025. The pattern is deliberate: Silo AI brought software talent, ZT Systems brought rack design, a July partnership pulled Cerebras into the orbit of AMD systems, and Taalas now brings the most specialized silicon of the lot. Specialized inference hardware is becoming the newest arena in the chip giants’ arms race.

Deal at a Glance

  • Target: Taalas, Toronto, founded 2023; AMD calls it “a pioneer in specialized AI inference silicon.”
  • Price undisclosed; the deal is subject to customary closing conditions and regulatory approvals.
  • Vamsi Boppana, SVP of AMD’s AI Group: customers should get “the flexibility to deploy the right compute solutions for every AI workload.”
  • Taalas co-founder and CEO Ljubisa Bajic: “We founded Taalas to rethink AI inference from the ground up by building the hardware around the model.”

Printing a Model Into a Chip: How MSICs Work

Taalas’ chips are model-specific integrated circuits — MSICs. Each has two regions: a mask-ROM recall fabric that holds the weights etched into the hardware, and an SRAM recall fabric for KV caches and fine-tuning adapters. The mask-ROM half is what makes the design extreme: weights are written once at fabrication and cannot be updated, which is why Taalas treats model re-spins as a product cadence rather than a bug. The HC1 test chip, fabricated on TSMC’s 6nm process and revealed in February, served Llama 3.1 8B at 16,960 tokens per second — 48x faster than Nvidia GPUs and 8.5x faster than Cerebras accelerators at the time. The follow-up HC2, due this summer, targets 20 billion parameters per chip; a trillion-parameter model would need about 50 accelerators chained in pipeline parallelism.

The Price of Speed

The trade-off is explicit: any model change beyond LoRA-style adapters requires a chip re-spin, though only two metal layers need to change. Bajic says a new model can be “realized in hardware in only two months.” MSICs fit workloads with stable weights and extreme throughput, not fast-moving frontier training. AMD’s integration idea is hybrid: GPUs handle prompt processing while Taalas silicon generates tokens, packaged with Helios rack-scale systems, EPYC CPUs, and ROCm software. The Register adds that etching weights into silicon reportedly costs about one percent of training a frontier model, and notes that serving the same model on Nvidia’s LPX systems would take a few dozen GPUs plus at least 2,000 Groq LPUs.

What It Means for the Inference Market

Lisa Su put the philosophy plainly: “I’m a big believer that there’s no one-size-fits-all as it comes to chips.” That sentence summarizes where the industry is heading: training versus inference, prompt processing versus generation, general versus specialized, all splitting into different hardware classes. The press release frames inference as one of the fastest-growing AI segments, with workloads growing specialized enough that general-purpose architectures leave compute and memory on the table. AMD also stressed the Canadian angle, promising to retain and grow Taalas’ Toronto engineering base.

For developers buying inference capacity, the near-term impact is limited, but the direction is clear. Fixed-model, high-throughput workloads — recommendations, moderation, agents that run the same pipeline all day — get dedicated silicon options first, and general-purpose GPU pricing gains another source of downward pressure. Worth remembering, too, how the Groq precedent resolved: Nvidia paid for licenses and people rather than a product line, and AMD may see Taalas the same way — a portfolio of designs and a team that knows how to hard-wire models, aimed at whichever inference niches grow big enough to justify fixed-function hardware. The open question is iteration speed: two months from trained weights to hardware is fast for custom silicon, but it is still an eternity next to a GPU that runs a new checkpoint the day it ships.

Sources

AI-assisted summary compiled from the sources above, reviewed by a human before publishing.

SHAREXEMAIL