R_REDDYX.XYZ
AIOpenEdge ABL328

avifenesh/memra-archive-20260901

MEMRA ARCHIVE CUDA INFERENCE

Rust + CUDA inference engine for RTX PRO 6000 Blackwell and RTX 5090. Serves safetensors and GGUF via OpenAI‑compatible API with per‑device tuned defaults and speculative decode byte‑identical to plain decode. Hosted at inference.tiyuvta.ai.

// KEY FEATURES

  • Runs safetensors and GGUF models on RTX PRO 6000 Blackwell and RTX 5090 via CUDA kernels.
  • Provides an OpenAI‑compatible API for text generation with latency under 10 ms per token.
  • Automatically tunes kernels for each device (per‑device tuned defaults).
  • Speculative decoding yields byte‑identical output to regular decoding.
#blackwell#cuda#gemma#gguf#gpu-kernels#inference-engine#llm#llm-inference#llm-serving#moe#nvfp4#openai-api
Open on GitHub →

New repositories every 30 minutes

REDDYX AI scans GitHub 24/7 and ships the best AI/ML/Web3 projects to Telegram.

Join on Telegram
← Full catalog·Full index

// SIMILAR REPOSITORIES

Claw Code: Fast Code192068n8n Process Automation187437DeepSeek Plugin Framework186631Agent AI Optimization186627AutoGPT: AI Agent184170Everything for Claude Code179304

← FULL CATALOG

avifenesh/memra-archive-20260901 — MEMRA ARCHIVE CUDA…