Filtered by: Cache × GPU × Clear all

kvcache-ai/ktransformers

kvcache-ai 2026-07-20
kvcache-ai/ktransformers Python
A Flexible Framework for Experiencing Heterogeneous LLM Inference/Fine-tune Optimizations
18,506 1,457

KTransformers implements a flexible framework for CPU-GPU heterogeneous LLM inference and fine-tuning, with a focus on MoE models and quantized kernels. It relates to agentic architecture and hardware/software co-design by enabling efficient deployment of large models on consumer hardware through heterogeneous expert placement and CPU-optimized AMX/AVX operations. The repository has a strong upward star trend, currently at 18,506 total stars with 360 stars today. The original paper is published in the ACM Digital Library under DOI 10.1145/3731569.3764843.