Filtered by: GPU × LLM × Clear all

Harness Engineering for LLM-Driven GPU Kernel Generation

Yue Shui, Chenyu Ma, Hangfei Xu, Shengzhao Wen 2026-07-22

LLM-driven GPU kernel generation is unreliable without a system to constrain, validate, and select candidate code. This paper introduces a harness-centered system that separates an evaluation harness from a profile-backed optimization controller to enforce compilation, correctness, and timing. Across five operator definitions, the system achieved mean-latency speedups over FlashInfer baselines ranging from 1.12x to 29.68x on NVIDIA Blackwell B200 GPUs. The findings demonstrate that expert-provided optimization directions and workload context remain critical for reliable AI-driven kernel optimization.

PDF

ExpertPlex: A High-Goodput Disaggregated Serving System for MoE LLMs with Adaptive Persistent Kernels

Bingyang Wu, Chao Jin, Zili Zhang, Xinming Wei 2026-07-22

ExpertPlex addresses the problem of inefficient resource utilization in serving large Mixture-of-Experts (MoE) LLMs, where existing prefill-decode disaggregation wastes resources due to coarse allocation and colocation suffers from phase interference and poor load tracking. The method shares massive MoE experts across phases while disaggregating attention modules, and introduces adaptive persistent kernels, attention-initiated MoE communication, and a tile-to-cluster model to optimize execution. Experiments serving MiniMax-M2.7 and GLM-5.1-FP8 show ExpertPlex improves goodput by up to 2.01× over instance-level disaggregation and 1.66× over colocation. This matters because it enables efficient, high-throughput serving of increasingly large MoE models without overprovisioning or phase interference.

PDF