Filtered by: Accelerator × Clear all

A Flexible Sparsity-Aware FPGA Accelerator with Column-Wise Compression for Efficient CNN Inference

Amirhossein Zarei, Shervin Vakili 2026-07-24

Problem: Efficient CNN acceleration on resource-constrained FPGAs is challenged by the irregularity of sparsity patterns and associated hardware overhead. Method: SparHiXcel-v2 introduces a column-wise kernel compression technique within a scalable 2D MAC array and a hardware-algorithm co-design framework with ordering optimization and multi-phase structured pruning. Finding: On a cost-effective AMD Kintex UltraScale+ FPGA, SparHiXcel-v2 achieves over 2.5 TOPS and 210 GOP/s/W for VGG16 and over 1.1 TOPS and 72 GOP/s/W for ResNet18 with modest accuracy degradation. Why it matters: This work demonstrates a practical balance between sparsity flexibility and hardware efficiency, enabling high-throughput, energy-efficient CNN inference on resource-constrained platforms.

PDF