Persistent State Machines: LLM Attention with INT4 In-Memory Cells
I present a formal framework for Large Language Model attention using Persistent State Machines with stationary in-memory cells. My work includes complete mathematical proofs for quantization error bounds and deterministic finite-automaton equivalence. I validated this architecture on Zynq-7000 and UltraScale+ FPGAs, demonstrating ultra-low power consumption and successful system-level integration with PCIe Gen3 bridges while maintaining bit-exact accuracy.
The dynamic power of the core logic is estimated below 1.0 mW, yielding a normalized dynamic energy of 3.81 × 10^-5 pJ/op.