d-Matrix stacks DRAM directly under logic to hit 20x HBM bandwidth density

D-Matrix Raptor 3D-DRAM Accelerator for Generative Inference at Hot Chips 2026

d-Matrix stacks DRAM directly under logic to hit 20x HBM bandwidth density

At Hot Chips 2026, d-Matrix presented Raptor, a 3D-DRAM accelerator for generative inference that stacks a TSMC N4 logic die face-to-face on DRAM. The company claims roughly 20 times the bandwidth per square millimeter of HBM4 and 13.5 times better power efficiency per GB/s, enough to serve a 3-trillion-parameter model with 1M context at about 1,000 tokens per second per user. The talk details solutions to bank mapping, I/O power, and thermal reliability challenges.

Raptor posts about 32.6 GB/s per mm2 compared with roughly 1.5 GB/s for the HBM parts, around 20 times the bandwidth per square millimeter, and 2.96 mW per GB/s against 40 mW, a 13.5x improvement.
  1. bix6

    Can anyone explain this article in English?

  2. unixhero

    What is Hot Chips?

More from this day

2026-09-14