FIBER: A New GPU Architecture That Decouples Threads from Registers for Faster Tensor Compute
A Thread-Register Decoupled GPU Execution Model for Efficient Tensor Computation

Modern GPUs rely on Tensor Cores for AI workloads, but fixed parallelism and coarse-grained scheduling limit efficiency when non-GEMM ops are interleaved with GEMM. Researchers propose FIBER, an architecture that decouples execution from private register ownership, enabling dynamic parallelism scaling and fine-grained dataflow scheduling. In mixed-precision LLM serving, FIBER achieves up to 2.25x end-to-end speedup on Ampere, with 1.8x and 2.09x on Hopper and Blackwell, and kernel-level gains up to 2.49x.
FIBER achieves a 2.25x end-to-end speedup on Ampere (1.15x for the original FP16 computation), with 1.8x and 2.09x on Hopper and Blackwell respectively, and kernel-level gains up to 2.49x.