memra-server
Rust + CUDA LLM inference engine for Blackwell (Tuned specifically on RTX PRO 6000, RTX 5090, B200): OpenAI-compatible (+converse and ant) serving, per-model X hardware exactness gates. NVFP4/mixed (fp8 hybrid, 4o6, etc - correctness, performance, hardware specific adapted) main quant support.
Activity
- Latest release
- 5d ago
- Total releases
- 83
- Cadence
- ~daily
- Last 12 months
- 83
Reach
- Downloads
- 1.2k
- Stars
- 1
Details
- License
- unknown
- First release
- Aug 04, 2026
Releases
1–50 of 83