AIPython★ 2,339
youssofal/MTPLX
MTPLX SUPERFAST LAUNCH
MTPLX lets you run Qwen 3.8 Flash Next and Qwen 3.8 27B on Mac at 125 tokens/sec in OpenCode. Thanks to MTP speculative decoding on Apple Silicon, results stay accurate at any temperature and server works with OpenAI and Anthropic APIs.
// KEY FEATURES
- Native MTP speculative decoding on Apple Silicon for inference at any temperature.
- Run Qwen 3.8 Flash Next and Qwen 3.8 27B on Mac with performance up to 125 tokens/sec.
- Full compatibility with OpenAI and Anthropic APIs — the server replaces cloud calls.
- Runs on Python, optimized for M-series chips, no extra drivers needed.
#anthropic-compatible#apple-silicon#claude-code#flash-next#inference-engine#llm-inference#local-ai#local-llm#macos#metal#mlx#mtp
Open on GitHub →New repositories every 30 minutes
REDDYX AI scans GitHub 24/7 and ships the best AI/ML/Web3 projects to Telegram.
Join on Telegram