CaSA - Ternary LLM inference inside COTS DRAM via charge-sharing

Show HN: Running PrismML's Bonsai inside DRAM by breaking DDR4 timing rules

CaSA is a groundbreaking architecture that enables ternary LLM inference directly within standard COTS DRAM, bypassing the memory bus entirely. While software quantization like PrismML's Bonsai reduces model size, CaSA solves the critical hardware bottleneck of memory bandwidth and thermal throttling on edge devices. By executing AI natively inside memory through charge-sharing, it eliminates the energy drain of transferring gigabytes of data across the SoC bus. This innovation paves the way for running massive 27B+ parameter models on smartphones with true on-device privacy and efficiency.

Why are we still transferring data to the compute? Why not execute AI inference natively within the memory?

More from this day

2026-07-27