R_REDDYX.XYZ
AIGo202

defilantech/LLMKube

LLMKUBE: LLM IN KUBERNETES

LLMKube - Kubernetes operator for LLM inference on heterogeneous GPU fleet: NVIDIA CUDA, AMD Vulkan, Apple Silicon Metal. Supports llama.cpp, vLLm, TGI, mlx-server, multi‑GPU sharding, model cache and OpenAI‑compatible API.

// KEY FEATURES

  • Runs LLM inference on NVIDIA, AMD and Apple GPU via Kubernetes operator.
  • Supports runtimes llama.cpp, vLLm, TGI and mlx-server with selection based on load.
  • Provides multi‑GPU sharding of models and automatic caching of GGUF files.
  • Provides OpenAI‑compatible REST endpoints, compatible with chat.completions clients.
#ai#apple-silicon#autoscaling#edge-computing#gguf#gpu#homelab#inference#kubernetes#kubernetes-operator#llama-cpp#llm
Open on GitHub →

New repositories every 30 minutes

REDDYX AI scans GitHub 24/7 and ships the best AI/ML/Web3 projects to Telegram.

Join on Telegram
← Full catalog·Full index

// SIMILAR REPOSITORIES

Claw Code: Fast Code192068n8n Process Automation187437DeepSeek Plugin Framework186631Agent AI Optimization186627AutoGPT: AI Agent184170Everything for Claude Code179304

← FULL CATALOG

defilantech/LLMKube — LLMKUBE: LLM IN KUBERNETES | REDDYX AI