local-ai
Memra
memra is a Rust + CUDA inference engine for serving GGUF models on NVIDIA GPUs. It is optimized for Blackwell (sm_120a), with a separately gated Hopper/H100
328 stars38 forksRustUpdated 8/30/2026100% free · open source
What it does
Memra is a Rust + CUDA inference engine that serves GGUF models on NVIDIA GPUs, optimized for Blackwell and Hopper/H100 architectures.
When to use it
- •You need to deploy GGUF models on NVIDIA GPUs for high-performance inference.
- •Your startup requires a customized inference engine for specific GPU architectures.
- •You want to leverage Rust and CUDA for building a high-speed AI application.
Ready-to-paste prompt
cargo run --release -- -m path/to/model.gguf -i path/to/input.data
Heads up: Ensure you have the CUDA toolkit installed and configured properly on your system, as memra relies on it for GPU acceleration.
Saves to your device
Use with Claude
New
Skip the builder — one click puts this in Claude, Cursor, Antigravity and more.
✅ Light setup
Installs with a command or two; your AI agent can do it for you.
Try it instantly — no install
Claude Code
mkdir -p ~/.claude/skills/memra && curl -fsSL https://workflowstacks.com/api/skills/memra/claude-skill -o ~/.claude/skills/memra/SKILL.mdOpen in another AI app
Opens the app with this repo with the prompt ready to go — no copy-paste needed.
Use with Claude
New
Skip the builder — one click puts this in Claude, Cursor, Antigravity and more.
✅ Light setup
Installs with a command or two; your AI agent can do it for you.
Try it instantly — no install
Claude Code
mkdir -p ~/.claude/skills/memra && curl -fsSL https://workflowstacks.com/api/skills/memra/claude-skill -o ~/.claude/skills/memra/SKILL.mdOpen in another AI app
Opens the app with this repo with the prompt ready to go — no copy-paste needed.
Details
Creator
avifenesh
Language
Rust
Category
local-ai
Published
7/5/2026
Are you the creator of this tool? Claim your listing → and earn 85% of every sale.
Related skills
More local-ai tools founders pair with this one.
local-ai★ 126,617
Run LLMs with llama.cpp
Get fast LLM inference with llama.cpp, for founders using C++.
local-ai★ 91,165
Stirling-PDF: Edit PDFs Anywhere
Edit PDFs on any device with Stirling-PDF, a tool for startup founders
local-ai★ 31,165
Run AI Locally with llmfit
Find compatible AI models for your hardware with llmfit, a tool.
local-ai★ 29,040
koreader: Multi-format eBook Reader
Get a versatile eBook reader for various devices.
local-ai★ 16,018
MNN: Fast On-Device AI
Get high-performance Edge AI with MNN, battle-tested by Alibaba, for startup founders
local-ai★ 10,300
Run AI Locally with runanywhere-sdks
Run AI models on-device with this production-ready toolkit, ideal for startup founders