Filtered by: Cache × Numa × Clear all

DGNA: Dissecting GPU NUMA Architecture through Microbenchmarking and Data Analysis

Changxi Liu, Yun Chen, Trevor E. Carlson 2026-07-23

The problem is that modern GPU memory architectures, particularly NUMA mechanisms within L2 and DRAM, remain a black-box, hindering optimization and simulation. DGNA introduces a methodology using microbenchmarking and Gaussian mixture models to measure L2 and DRAM latency without relying on intrinsic instructions. Applied to NVIDIA's A100 and H100 GPUs, it reveals NUMA node architecture, SM-NUMA relationships, and NUMA-aware memory allocation strategies for cache coherence. This matters because it is the first work to detail GPU memory subsystem NUMA architecture, enabling better application optimization and architectural design.

PDF