At-the-Roofline Sparse Tensor Contractions on Vector Processors for Transformer Inference
The problem is that fine-grained sparsity in Transformer inference cannot be efficiently exploited on vector processors because existing RVV architectures lack native support for Gustavson's dataflow, forcing reliance on software index decoding and L1-backed indexed memory operations that keep sparse tensor contractions far below the roofline bound. The method introduces Ventaglio, a runtime-configurable sparse execution unit with RVV ISA extensions that provides indexed gather-accumulate-scatter support to drive sparse tensor contractions toward their roofline performance. Experimental evidence from a 12nm FinFET implementation shows Ventaglio accelerates sparse tensor contraction kernels by 6.9–7.4× over optimized RVV baselines with only 3.1% area overhead, and on a DuoGPT-pruned LLaMA-3-8B model with 40–60% dual sparsity achieves 2.40–5.25× and 2.06–3.16× speedup over dense baselines during prefill and autoregressive decoding, respectively. This matters because it demonstrates that hardware-software co-design can close the roofline gap for sparse tensor contractions, enabling practical speedups for Transformer inference on vector processors without prohibitive area cost.