AIGo★ 202
defilantech/LLMKube
LLMKUBE: LLM IN KUBERNETES
LLMKube - Kubernetes operator for LLM inference on heterogeneous GPU fleet: NVIDIA CUDA, AMD Vulkan, Apple Silicon Metal. Supports llama.cpp, vLLm, TGI, mlx-server, multi‑GPU sharding, model cache and OpenAI‑compatible API.
// KEY FEATURES
- Runs LLM inference on NVIDIA, AMD and Apple GPU via Kubernetes operator.
- Supports runtimes llama.cpp, vLLm, TGI and mlx-server with selection based on load.
- Provides multi‑GPU sharding of models and automatic caching of GGUF files.
- Provides OpenAI‑compatible REST endpoints, compatible with chat.completions clients.
#ai#apple-silicon#autoscaling#edge-computing#gguf#gpu#homelab#inference#kubernetes#kubernetes-operator#llama-cpp#llm
Open on GitHub →New repositories every 30 minutes
REDDYX AI scans GitHub 24/7 and ships the best AI/ML/Web3 projects to Telegram.
Join on Telegram