d-Matrix stacks DRAM directly under logic to hit 20x HBM bandwidth density
D-Matrix Raptor 3D-DRAM Accelerator for Generative Inference at Hot Chips 2026

At Hot Chips 2026, d-Matrix presented Raptor, a 3D-DRAM accelerator for generative inference that stacks a TSMC N4 logic die face-to-face on DRAM. The company claims roughly 20 times the bandwidth per square millimeter of HBM4 and 13.5 times better power efficiency per GB/s, enough to serve a 3-trillion-parameter model with 1M context at about 1,000 tokens per second per user. The talk details solutions to bank mapping, I/O power, and thermal reliability challenges.
Raptor posts about 32.6 GB/s per mm2 compared with roughly 1.5 GB/s for the HBM parts, around 20 times the bandwidth per square millimeter, and 2.96 mW per GB/s against 40 mW, a 13.5x improvement.