Yixun Hong’s Homepage Yixun Hong
Home Publications Daily Gallery Resume
Use the up and down arrow keys to navigate search results.
Home Publications Daily Gallery Resume
Filtered by: GPU × Scheduling × Clear all

Automated Tensor Scheduling for Hybrid CPU-GPU LLM Inference on Consumer Devices

Yangyijian Liu, Hongyi Ye, Mingyang Li, Wu-jun Li 2026-07-14
Scheduling × GPU × LLM Inference Hardware Data

Running large language models on consumer devices such as laptops and desktops is challenging because model weights often exceed GPU memory capacity, making offloading inference necessary to extend effective model capacity with CPU memory. Existing offloading systems, however, typically rely on.

PDF

GPU-Tile-Sim: A Tile-Centric GPU Simulation Framework for LLM Hardware-Software Co-Design

Yitong Ding, Jiawei Huang, Renyang Guan, Yangjie Zhou 2026-07-14
GPU × Simulation LLM Hardware Design Workload Scheduling × Architecture

Modern LLM (large language model) workloads increasingly rely on optimized GPU kernels through hardware-software co-design. These kernels achieve high-performance through fine-grained dependency scheduling and computation-memory overlap.

PDF

AI Keywords

All
Agentic (2) Architecture (3) LLM (3) Agent (1) Scheduling (2) × Design (1) GPU (2) × Hardware (2) Simulation (1) Inference (2) Data (2) Workload (1)

Profile

Gem5 Simulator Architectural Modeling Simulation Design Memory Ramulator
© Yixun Hong. Powered by Yixun’s Homepage.
浙ICP备2023041507号-1