A Thread-Register Decoupled GPU Execution Model for Efficient Tensor Computation
FIBER decouples GPU execution fibers from register ownership to enable finer-grained scheduling and more flexible tensor execution. The paper reports up to 2.25× end-to-end gains for mixed-precision LLM serving across Ampere, Hopper, and Blackwell GPUs.