AIRust★ 206
novitalabs/pegaflow
PEGAFLOW KV CACHE FOR LLM
Pegaflow — high‑performance KV‑cache storage for LLM inference. Supports offloading to GPU, caching on SSD, and cross‑node exchange via RDMA. Compatible with vLLM and SGLang.
// KEY FEATURES
- Offloads KV cache to GPU memory, reducing model VRAM usage.
- Caches KV cache on NVMe SSD for fast access when GPU memory is exceeded.
- Provides KV cache transfer between cluster nodes via RDMA with low latency.
- Integrates with vLLM and SGLang as a plug‑in KV‑cache storage requiring no code changes.
#inference#kv-cache#llm#vllm
Open on GitHub →New repositories every 30 minutes
REDDYX AI scans GitHub 24/7 and ships the best AI/ML/Web3 projects to Telegram.
Join on Telegram