local-ai

Memra

memra is a Rust + CUDA inference engine for serving GGUF models on NVIDIA GPUs. It is optimized for Blackwell (sm_120a), with a separately gated Hopper/H100
328 stars38 forksRustUpdated 8/30/2026100% free · open source
What it does

Memra is a Rust + CUDA inference engine that serves GGUF models on NVIDIA GPUs, optimized for Blackwell and Hopper/H100 architectures.

When to use it
  • You need to deploy GGUF models on NVIDIA GPUs for high-performance inference.
  • Your startup requires a customized inference engine for specific GPU architectures.
  • You want to leverage Rust and CUDA for building a high-speed AI application.
Ready-to-paste prompt
cargo run --release -- -m path/to/model.gguf -i path/to/input.data
Heads up: Ensure you have the CUDA toolkit installed and configured properly on your system, as memra relies on it for GPU acceleration.
Saves to your device
Use with Claude
New

Skip the builder — one click puts this in Claude, Cursor, Antigravity and more.

✅ Light setup

Installs with a command or two; your AI agent can do it for you.

Try it instantly — no install
Claude Code
mkdir -p ~/.claude/skills/memra && curl -fsSL https://workflowstacks.com/api/skills/memra/claude-skill -o ~/.claude/skills/memra/SKILL.md
Open in another AI app

Opens the app with this repo with the prompt ready to go — no copy-paste needed.

Connect the whole catalog (MCP)
claude mcp add --transport http workflowstacks https://workflowstacks.com/api/mcp

Adds a WorkflowStacks connector to Claude Code: search and load any skill here by chatting.

Quick Actions
Details
Creator
avifenesh
Language
Rust
Category
local-ai
Published
7/5/2026

Are you the creator of this tool? Claim your listing → and earn 85% of every sale.