NELSSA: A GPU-PNM Heterogeneous System for Mixed-Length LLM Serving via Length-based Request Placement
NELSSA addresses the problem of GPU-centric LLM serving systems suffering from throughput and latency inefficiencies when handling highly heterogeneous mixed-length workloads. The method introduces a GPU-PNM heterogeneous system that uses length-based request placement, routing short-context requests to GPUs and long-context requests to PNM, with runtime migration for dynamic context growth. Experimental evidence shows that NELSSA improves decode throughput by up to 5.5x in tokens/sec and reduces P99 latency by up to 15x compared to GPU-only baselines across mixed-length LLM workloads. This matters because it demonstrates that integrated GPU-PNM serving, enabled by CXL-based disaggregation, is a promising paradigm for building scalable and flexible LLM infrastructures that can efficiently support evolving workloads.