R_REDDYX.XYZ
AISwift232

carloslfu/slotstream

SLOTSTREAM LAUNCHES QWEN3.8 ON MAC

Run Qwen3.8‑Flash‑Next (125B MoE, 104 GB at 4‑bit) on Mac by streaming experts from SSD on demand. Implemented with MLX + Swift, it offers an Ollama‑compatible API.

// KEY FEATURES

  • Streams MoE experts from SSD, cutting RAM usage to a fraction of 104 GB.
  • Runs on Apple Silicon via MLX and Swift, delivering high inference speed.
  • Offers an API fully compatible with Ollama for easy integration into existing workflows.
  • Supports running Qwen3.8‑Flash‑Next (125B MoE, 4‑bit) on Mac with minimal resources.
#apple-silicon#llm#llm-inference#local-llm#macos#mixture-of-experts#mlx#ollama#qwen#swift
Open on GitHub →

New repositories every 30 minutes

REDDYX AI scans GitHub 24/7 and ships the best AI/ML/Web3 projects to Telegram.

Join on Telegram
← Full catalog·Full index

// SIMILAR REPOSITORIES

Claw Code: Fast Code192068n8n Process Automation187437DeepSeek Plugin Framework186631Agent AI Optimization186627AutoGPT: AI Agent184170Everything for Claude Code179304

← FULL CATALOG

carloslfu/slotstream — SLOTSTREAM LAUNCHES QWEN3.8 ON MAC…