AIOpenEdge ABL★ 328
avifenesh/memra-archive-20260901
MEMRA ARCHIVE CUDA INFERENCE
Rust + CUDA inference engine for RTX PRO 6000 Blackwell and RTX 5090. Serves safetensors and GGUF via OpenAI‑compatible API with per‑device tuned defaults and speculative decode byte‑identical to plain decode. Hosted at inference.tiyuvta.ai.
// KEY FEATURES
- Runs safetensors and GGUF models on RTX PRO 6000 Blackwell and RTX 5090 via CUDA kernels.
- Provides an OpenAI‑compatible API for text generation with latency under 10 ms per token.
- Automatically tunes kernels for each device (per‑device tuned defaults).
- Speculative decoding yields byte‑identical output to regular decoding.
#blackwell#cuda#gemma#gguf#gpu-kernels#inference-engine#llm#llm-inference#llm-serving#moe#nvfp4#openai-api
Open on GitHub →New repositories every 30 minutes
REDDYX AI scans GitHub 24/7 and ships the best AI/ML/Web3 projects to Telegram.
Join on Telegram