AISwift★ 232
carloslfu/slotstream
SLOTSTREAM LAUNCHES QWEN3.8 ON MAC
Run Qwen3.8‑Flash‑Next (125B MoE, 104 GB at 4‑bit) on Mac by streaming experts from SSD on demand. Implemented with MLX + Swift, it offers an Ollama‑compatible API.
// KEY FEATURES
- Streams MoE experts from SSD, cutting RAM usage to a fraction of 104 GB.
- Runs on Apple Silicon via MLX and Swift, delivering high inference speed.
- Offers an API fully compatible with Ollama for easy integration into existing workflows.
- Supports running Qwen3.8‑Flash‑Next (125B MoE, 4‑bit) on Mac with minimal resources.
#apple-silicon#llm#llm-inference#local-llm#macos#mixture-of-experts#mlx#ollama#qwen#swift
Open on GitHub →New repositories every 30 minutes
REDDYX AI scans GitHub 24/7 and ships the best AI/ML/Web3 projects to Telegram.
Join on Telegram