AIShell★ 276
peonist-ai/halogen-flash-server
FAST LAUNCH QWEN3.8 ON AMD
The project provides scripts for instant deployment of Qwen3.8-Flash-Next on AMD Strix Halo (gfx1151) GPU via HIP. All dependencies are built automatically, and inference runs with maximum throughput.
// KEY FEATURES
- Launches Qwen3.8-Flash-Next on AMD Strix Halo (gfx1151) GPU in seconds via HIP backend.
- Collects dependencies LLVM, ROCm, HIP, PyTorch with a single install.sh script.
- Optimizes inference cores for gfx1151, achieving peak TFLOPs and minimal latency.
- Supports FP16/BF16 precision and dynamic batching for maximum throughput.
#amd#gfx1151#gpu-inference#hip#inference-engine#llm#llm-serving#local-llm#long-context#mixture-of-experts#openai-api#qwen
Open on GitHub →New repositories every 30 minutes
REDDYX AI scans GitHub 24/7 and ships the best AI/ML/Web3 projects to Telegram.
Join on Telegram