Automated Tensor Scheduling for Hybrid CPU-GPU LLM Inference on Consumer Devices

Yangyijian Liu, Hongyi Ye, Mingyang Li, Wu-jun Li 2026-07-14

Running large language models on consumer devices such as laptops and desktops is challenging because model weights often exceed GPU memory capacity, making offloading inference necessary to extend effective model capacity with CPU memory. Existing offloading systems, however, typically rely on.

PDF