CacheRoute is an innovative LLM scheduling scheme dedicated to enabling flexible KV cache reuse across LLM systems, improving task performance and system efficiency.
-
Updated
Jul 20, 2026 - Python
CacheRoute is an innovative LLM scheduling scheme dedicated to enabling flexible KV cache reuse across LLM systems, improving task performance and system efficiency.
Multimodal LLM inference gateway with KV-cache-aware routing and LMCache offload. OpenAI-compatible, benchmarked on GPUs.
KV-cache-aware LLM inference mesh on a single 6 GB GPU. 2× vLLM + LMCache shared KV + prefix-aware router on k3s, observed with Cilium/Hubble eBPF. Every number measured, every breakage documented.
Benchmarking LMCache under simulated RTT
Add a description, image, and links to the lmcache topic page so that developers can more easily learn about it.
To associate your repository with the lmcache topic, visit your repo's landing page and select "manage topics."