AMD acquires Taalas to add model-specific silicon to its AI inference stack

The Toronto chipmaker raised $219 million to hardwire AI models into silicon; AMD did not disclose the price and plans to pair the technology with Instinct GPUs.

By · Published

Why it matters

AMD is buying the design flow and engineering team needed to make model-specific inference chips, giving it an owned specialized engine to pair with Instinct GPUs as inference splits into distinct prompt and token-generation workloads.

AMD acquires Taalas to add model-specific silicon to its AI inference stack

AMD agreed Thursday to acquire Taalas, the Toronto AI inference chipmaker co-founded by former Tenstorrent CEO Ljubisa Bajic, giving AMD control of a specialized architecture designed to run individual AI models at high speed and lower power than general-purpose accelerators.

AMD did not disclose the purchase price or an expected closing date. The transaction remains subject to regulatory approval and other customary closing conditions, according to AMD's August 6th announcement.

Lisa Su (@LisaSu) described Taalas (@taalas_inc) as a team working at the edge of AI inference in a post on X. AMD said Taalas will join the artificial intelligence organization led by Vamsi Boppana and that AMD intends to integrate Taalas technology into its accelerator roadmap and systems built around Instinct GPUs.

Bajic founded Taalas in 2023 with Drago Ignjatovic and Lejla Bajic after previously founding AI chip developer Tenstorrent. The three founders had worked across AMD, Nvidia and Tenstorrent on CPUs, GPUs and AI processors, according to Taalas' 2024 funding announcement. The acquisition sends Bajic back to AMD with a team built around engineers who already know its hardware organization.

"We founded Taalas to rethink AI inference from the ground up by building the hardware around the model," Bajic said in AMD's announcement. He said AMD would give the Toronto team greater engineering resources and global reach.

A model hardwired into silicon

Taalas takes specialization further than conventional AI accelerators. Its design places a model's weights and dataflow directly into silicon, stripping away most programmability to reduce the movement of data between memory and compute hardware.

The tradeoff is severe. A Taalas chip optimized for one model cannot simply load a different model through software. New models require new chip masks, making Taalas dependent on a design process that can translate a model into silicon quickly enough to keep pace with AI releases.

Taalas says its automated process can turn a previously unseen model into custom silicon in about two months. That turnaround claim is central to the value AMD is buying: without it, model-specific chips risk becoming obsolete before reaching production.

Taalas publicly demonstrated the approach in February with HC1, a chip hardwired for Meta's Llama 3.1 8B model. Taalas reported that HC1 could generate 17,000 tokens per second per user while using one-tenth the power of competing systems. Those results were produced or compiled by Taalas, rather than established through a broad set of independent production tests.

Taalas also acknowledged that HC1's aggressive mix of 3-bit and 6-bit parameters caused quality degradation compared with GPU benchmarks. The chip's performance came with another basic constraint: HC1 could run Llama 3.1 8B and no other model.

Those limitations explain why AMD plans to use the technology alongside Instinct GPUs rather than present it as a wholesale GPU replacement. GPUs can handle changing models, large prompt-processing workloads and general-purpose computation. Taalas silicon can be assigned the stable, repetitive portions of inference where specialization produces the largest gain.

AMD builds around specialized inference

The deal follows AMD's July 23rd inference partnership with Cerebras, which pairs AMD Helios systems with Cerebras wafer-scale processors. Under that design, AMD hardware processes prompts and long context windows while Cerebras accelerates token generation, or decode.

Taalas gives AMD an architecture it owns for a similar division of work. EE Times reported that Taalas silicon could serve as a decode accelerator beside Instinct GPUs, allowing AMD to supply and control both compute engines. Taalas chips could also run complete inference workloads for smaller models used in power-constrained edge systems.

The acquisition comes less than six months after Taalas raised $169 million, taking its total financing to $219 million. Investors included Quiet Capital, Fidelity and semiconductor investor Pierre Lamond, according to Reuters reporting.

Taalas said in February that 24 employees developed HC1 after spending $30 million of the capital it had raised. AMD is therefore acquiring a compact engineering group, its model-to-silicon tool flow and working hardware before Taalas had to prove that the approach could support a broad commercial product line.

For Bajic, the sale shifts the task from financing repeated chip designs to fitting model-specific hardware inside AMD's Instinct, Helios and ROCm portfolio. AMD gains a technical route around the memory bottlenecks that limit token generation. The remaining test is whether Taalas can preserve its speed advantage while moving beyond a single, heavily quantized model and into systems customers can deploy at scale.

Reader comments

Conversation for this story loads after sign-in.