AIPython★ 1,734
vllm-project/vllm-metal
VLLM-METAL: ACCELERATING LLM ON APPLE SILICON
vllm-metal is a community‑maintained plugin that brings Metal‑based acceleration to vLLM on Apple Silicon. It uses MLX to run LLMs on macOS with low latency and high throughput, no CUDA required.
// KEY FEATURES
- Supports hardware acceleration via Metal on Apple Silicon (M1/M2/M3).
- Integrates with vLLM as a plugin, replacing the CUDA backend with MLX.
- Provides low‑latency LLM inference on macOS without external drivers.
- Supports quantization and dynamic batching to increase throughput.
#apple-silicon#llm#macos#metal#mlx#vllm
Open on GitHub →New repositories every 30 minutes
REDDYX AI scans GitHub 24/7 and ships the best AI/ML/Web3 projects to Telegram.
Join on Telegram