ai-agent

openinfer: Fast LLM Inference

Get fast and compatible LLM inference with openinfer, a Rust-based engine, serving models from Qwen3 to Kimi-K2.
628 stars97 forksRustGuide quality 9/10Updated 8/6/2026100% free · open source
What it does

OpenInfer provides fast and compatible large language model (LLM) inference for startup founders, serving models from Qwen3 to Kimi-K2 using a Rust-based engine

When to use it
  • When you need to deploy LLM models without relying on PyTorch
  • When you require OpenAI-compatible model serving for your startup
  • When you prefer a Rust-based solution for performance and compatibility
Ready-to-paste prompt
curl -X POST -H 'Content-Type: application/json' -d '{"prompt": "Write a short story about a character who discovers a hidden world."}' http://localhost:8000/infer
Heads up: Make sure you have Rust and CUDA installed on your system, as OpenInfer relies on these dependencies to function correctly
Saves to your device
Use with Claude
New

Skip the builder — one click puts this in Claude, Cursor, Antigravity and more.

✅ Light setup

Installs with a command or two; your AI agent can do it for you.

Try it instantly — no install
Claude Code
mkdir -p ~/.claude/skills/openinfer && curl -fsSL https://workflowstacks.com/api/skills/openinfer/claude-skill -o ~/.claude/skills/openinfer/SKILL.md
Open in another AI app

Opens the app with this repo with the prompt ready to go — no copy-paste needed.

Connect the whole catalog (MCP)
claude mcp add --transport http workflowstacks https://workflowstacks.com/api/mcp

Adds a WorkflowStacks connector to Claude Code: search and load any skill here by chatting.

Quick Actions
Details
Creator
openinfer-project
Language
Rust
Category
ai-agent
Published
2/17/2026

Are you the creator of this tool? Claim your listing → and earn 85% of every sale.