Filtered by: Simulation × Clear all

DGNA: Dissecting GPU NUMA Architecture through Microbenchmarking and Data Analysis

Changxi Liu, Yun Chen, Trevor E. Carlson 2026-07-24

The problem is that modern GPU memory architectures, particularly NUMA mechanisms within L2 and DRAM, remain black-boxes, hindering optimization and simulation. DGNA introduces a methodology using microbenchmarking and Gaussian mixture models to measure L2 and DRAM latency without relying on intrinsic instructions. Applied to NVIDIA A100 and H100 GPUs, it reveals NUMA node architecture, SM-NUMA relationships, and coherence-aware allocation strategies. This matters because it is the first work to detail GPU memory subsystem NUMA, enabling better application tuning and architectural design.

PDF

A Flexible Sparsity-Aware FPGA Accelerator with Column-Wise Compression for Efficient CNN Inference

Amirhossein Zarei, Shervin Vakili 2026-07-24

Problem: Efficient CNN acceleration on resource-constrained FPGAs is challenged by the irregularity of sparsity patterns and associated hardware overhead. Method: SparHiXcel-v2 introduces a column-wise kernel compression technique within a scalable 2D MAC array and a hardware-algorithm co-design framework with ordering optimization and multi-phase structured pruning. Finding: On a cost-effective AMD Kintex UltraScale+ FPGA, SparHiXcel-v2 achieves over 2.5 TOPS and 210 GOP/s/W for VGG16 and over 1.1 TOPS and 72 GOP/s/W for ResNet18 with modest accuracy degradation. Why it matters: This work demonstrates a practical balance between sparsity flexibility and hardware efficiency, enabling high-throughput, energy-efficient CNN inference on resource-constrained platforms.

PDF