d-Matrix's Raptor puts stacked DRAM to work on inference bandwidth

An ISCA paper reports up to 100 TB/s per card, while its throughput comparisons remain modeled and Raptor is still pre-production.

By · Published

Primary source: Aligned News - AI Intelligence

Why it matters

Raptor puts d-Matrix's long-running inference thesis into early silicon, but modeled speedups and 100 TB/s bandwidth still have to translate into useful capacity, manufacturable chips and customer deployments.

An unbranded semiconductor sample with stacked DRAM dies sits beside a generic test fixture, illustrating d-Matrix's Raptor bandwidth story.

d-Matrix's Raptor accelerator puts compute logic face-to-face with stacked DRAM to address the data-movement bottleneck in generative inference. A paper presented at ISCA 2026 reports up to 100 TB/s of bandwidth per card and higher throughput than HBM- and SRAM-based configurations under modeled deployment assumptions. Those are research results, not production benchmarks against shipping accelerators.

The figures resurfaced in an October 2nd post on X, but the underlying research was presented at ISCA in Raleigh, North Carolina, during the conference's June 27th through July 1st run. d-Matrix CTO Sudeep Bhoja also presented the technology at Hot Chips in August. The work describes early Raptor silicon; d-Matrix says the commercial product is expected in the fourth quarter of 2027.

For co-founder and CEO Sid Sheth, Raptor extends a bet made before generative AI turned inference into a mainstream infrastructure priority. Sheth wrote that he and co-founder Sudeep Bhoja left Inphi in 2019 after seeing cloud customers use its data-center interconnect products in distributed AI systems. They concluded that running trained models at scale would become a major computing market, even while investors were still focused on training.

A memory bet from the start

Sheth's background is in the movement of data through networks and data centers. Before founding d-Matrix, he ran Inphi's networking business; Bhoja, now d-Matrix's CTO, also brought experience in networking and interconnects. They based their inference thesis on data-center AI workloads and the infrastructure built to serve them, before generative AI became a mainstream infrastructure priority.

The bet has taken time to reach product silicon. In a 2024 account of d-Matrix's early years, Sheth recounted an early architecture pivot and the difficulty of keeping hardware and software development aligned. Raptor carries that original memory-focused premise into its physical design: d-Matrix says it places a TSMC 4-nanometer logic die face-to-face with a custom DRAM die.

That design targets the repeated movement of model weights and key-value cache data during autoregressive decoding. The Raptor paper reports stream-blocking to map KV-cache streams across configurable DRAM channels, alongside techniques for power, thermal management and error correction. Its workloads include Llama 3.1 70B, DeepSeek-V3, Kimi K2, GPT-OSS, Whisper and Canary.

Diagram of Raptor's logic die positioned face-to-face with a custom DRAM die, with model weights and KV-cache data used in autoregressive decoding and KV-cache streams mapped across configurable DRAM channels.
The Raptor paper describes face-to-face logic and DRAM dies and stream-blocking for mapping KV-cache streams across configurable DRAM channels - AI explanatory diagram, not documentary evidence. RuntimeWire - AI-generated diagram.

The paper reports 4.71 times the throughput of an HBM configuration and 2.44 times that of an SRAM configuration across those workloads. The comparison is modeled, so it does not establish that a Raptor card outperforms a specific Nvidia or AMD product in a customer deployment. The up-to-100-TB/s bandwidth figure is also only part of the system tradeoff: d-Matrix's reported Raptor configuration has 32GB of memory, far less capacity than some HBM-based accelerators.

From chip architecture to a rack strategy

Raptor's commercial case now depends on more than the memory stack. In September, d-Matrix announced that it would integrate its next-generation accelerators into Nvidia's MGX rack architecture using NVLink Fusion. The partnership announcement positions Raptor alongside Nvidia GPUs in a system that could assign compute-intensive prefill work to GPUs and latency-sensitive decoding to d-Matrix accelerators. That gives d-Matrix a route into an established rack ecosystem, while making integration, software and deployment economics part of the test alongside chip performance.

D-Matrix says Raptor is expected to tape out before the end of 2026, with initial availability in Nvidia MGX racks expected in the fourth quarter of 2027. Those milestones leave manufacturing yield, production volume, pricing and customer performance to be demonstrated. The design's memory capacity also raises a question: how workloads divide between Raptor's high-bandwidth local memory and other memory in a rack.

In November 2025, d-Matrix said it closed a $275 million Series C at a $2 billion valuation, bringing its disclosed funding total to $450 million. The round was co-led by BullhoundCapital, Triatomic Capital and Temasek, with participation from Qatar Investment Authority, EDBI, Microsoft's M12, Nautilus Venture Partners, Industry Ventures and Mirae Asset, according to d-Matrix's funding announcement.

The money backs the thesis Sheth has pursued since 2019: inference hardware should be designed around the cost of moving data, not simply the volume of arithmetic. Raptor's early silicon and research comparisons give that thesis a more concrete form. Its commercial test will be whether the bandwidth advantage survives the constraints that matter in a data center: enough capacity for useful workloads, reliable production, software compatibility and a system customers can deploy.

Reader comments

Conversation for this story loads after sign-in.