HyMCache: A KV Cache Framework for Multi-Turn LLM Serving with CXL-Hybrid Memory

Hakbeom Jang, Inho Song, Sam H. Noh, Jongryool Kim 2026-07-23

HyMCache addresses the problem of high memory costs in multi-turn LLM serving by proposing a KV-cache framework that uses CXL-hybrid memory (CXL-HM), combining a small in-device DRAM with large SSD-backed capacity. The method employs request-level prefix prefetching and opportunistic write buffering to stage latency-critical reads in device DRAM, exploiting the read-dominant and predictable access patterns of multi-turn KV-cache. Experimental evidence on a real CXL-HM prototype shows HyMCache outperforms local LMCache by 3.0x in single-node and 1.45x in PD-disaggregated serving under the same DRAM budget, and incurs about 30% lower performance than 1 TB distributed-DRAM Mooncake while using 16x less DRAM. This matters because it enables TB-scale shared KV-cache reuse at SSD-level cost, significantly reducing DRAM requirements for scalable multi-turn LLM serving.

PDF