A Flexible Sparsity-Aware FPGA Accelerator with Column-Wise Compression for Efficient CNN Inference

Amirhossein Zarei, Shervin Vakili 2026-07-25

Problem: Efficient CNN inference on resource-constrained FPGAs is challenged by the irregularity of sparsity patterns and hardware overhead. Method: SparHiXcel-v2 introduces a column-wise kernel compression technique and a hardware-algorithm co-design framework with ordering optimization and multi-phase structured pruning. Finding: On a cost-effective AMD Kintex UltraScale+ FPGA, SparHiXcel-v2 achieves over 2.5 TOPS and 210 GOP/s/W for VGG16, and over 1.1 TOPS and 72 GOP/s/W for ResNet18 with modest accuracy degradation. Why it matters: This work provides a flexible and efficient FPGA accelerator that balances sparsity flexibility and hardware efficiency for resource-constrained CNN inference.

PDF