DGNA: Dissecting GPU NUMA Architecture through Microbenchmarking and Data Analysis

Changxi Liu, Yun Chen, Trevor E. Carlson 2026-07-26

The problem is that modern GPU memory architectures, particularly NUMA mechanisms within L2 and DRAM, remain black-boxes, hindering optimization and simulation. DGNA introduces a methodology using microbenchmarking and Gaussian mixture models to measure L2 and DRAM latency without intrinsic instructions. Applied to NVIDIA A100 and H100 GPUs, it reveals NUMA node architecture, SM-NUMA relationships, and NUMA-aware memory allocation strategies for cache coherence. This matters because it is the first detailed dissection of GPU memory subsystem NUMA, enabling better application optimization and architectural design.

PDF

A Flexible Sparsity-Aware FPGA Accelerator with Column-Wise Compression for Efficient CNN Inference

Amirhossein Zarei, Shervin Vakili 2026-07-26

The problem is that unstructured sparsity in CNNs causes hardware inefficiency, while structured sparsity sacrifices flexibility. SparHiXcel-v2, a flexible FPGA accelerator, uses a column-wise kernel compression technique and a scalable 2D MAC array to handle irregular sparsity with minimal overhead. On a cost-effective AMD Kintex UltraScale+ FPGA, it achieves over 2.5 TOPS and 210 GOP/s/W for VGG16 in structured sparsity mode with modest accuracy loss. This matters because it enables efficient CNN inference on resource-constrained platforms by balancing sparsity flexibility and hardware efficiency.

PDF