Filtered by: Serving × Clear all

FlashRT: Agent Harness for Guiding Agents to Deploy Real-Time Multimodal Applications

Krish Agarwal, Zhuoming Chen, Yanyuan Qin, Zhenyu Gu 2026-07-22

FlashRT addresses the problem of manually optimizing real-time multimodal application deployments across heterogeneous models and hardware. It introduces an agent harness that uses a chain-of-program paradigm to transform simple reference implementations into optimized multi-GPU deployments via iterative IR construction, validation, and measurement-gated optimization. On NVIDIA B200 GPUs, FlashRT achieves up to ~70x latency reduction and 2.8x throughput improvement, while on AMD MI355X GPUs it matches peak latency reduction and increases throughput improvement to 3.6x, outperforming expert implementations like vLLM-Omni by 65% in response latency. This matters because agent-driven optimization enables scalable, high-performance deployment across diverse hardware platforms without requiring hand-crafted expert tuning.

PDF

ExpertPlex: A High-Goodput Disaggregated Serving System for MoE LLMs with Adaptive Persistent Kernels

Bingyang Wu, Chao Jin, Zili Zhang, Xinming Wei 2026-07-22

ExpertPlex addresses the problem of inefficient resource utilization in serving large Mixture-of-Experts (MoE) LLMs, where existing prefill-decode disaggregation wastes resources due to coarse allocation and colocation suffers from phase interference and poor load tracking. The method shares massive MoE experts across phases while disaggregating attention modules, and introduces adaptive persistent kernels, attention-initiated MoE communication, and a tile-to-cluster model to optimize execution. Experiments serving MiniMax-M2.7 and GLM-5.1-FP8 show ExpertPlex improves goodput by up to 2.01× over instance-level disaggregation and 1.66× over colocation. This matters because it enables efficient, high-throughput serving of increasingly large MoE models without overprovisioning or phase interference.

PDF

HyMCache: A KV Cache Framework for Multi-Turn LLM Serving with CXL-Hybrid Memory

Hakbeom Jang, Inho Song, Sam H. Noh, Jongryool Kim 2026-07-22

HyMCache addresses the problem of high memory costs in multi-turn LLM serving by proposing a KV-cache framework that uses CXL-hybrid memory (CXL-HM), combining a small in-device DRAM with large SSD-backed capacity. The method exploits the read-dominant, predictable, and append-only nature of multi-turn KV-cache access, using request-level prefix prefetching and opportunistic write buffering to stage latency-critical reads in device DRAM. Experimental evidence on a real CXL-HM prototype shows that under the same DRAM budget, HyMCache outperforms local LMCache by 3.0x in single-node serving and 1.45x in PD-disaggregated serving, and compared to 1 TB distributed-DRAM Mooncake, it incurs about 30% lower performance but uses 16x less DRAM. This matters because it enables TB-scale SSD-backed KV reuse at DRAM-scale efficiency and SSD-level cost, significantly reducing memory expenses for long-context and agentic LLM workloads.

PDF