ai-agent

xllm: Fast AI Inference

Get high-performance AI model inference with xllm, ideal for startup founders working with large language models
1,518 stars277 forksC++Guide quality 8/10Updated 8/13/2026100% free · open source
What it does

xllm provides high-performance inference for large language models, enabling fast and efficient processing of AI workloads

When to use it
  • You need to deploy large language models in a production environment with strict latency requirements
  • Your startup relies on real-time natural language processing for tasks like sentiment analysis or text classification
  • You want to optimize the performance of your AI models on diverse hardware accelerators
Ready-to-paste prompt
xllm_infer -c config.json -m your_model.xllm -i 'What is the capital of France?'
Heads up: Ensure you have the necessary dependencies installed, including CMake and a compatible C++ compiler, and that your model is compatible with the xllm format
Saves to your device
Use with Claude
New

Skip the builder — one click puts this in Claude, Cursor, Antigravity and more.

✅ Light setup

Installs with a command or two; your AI agent can do it for you.

Try it instantly — no install
Claude Code
mkdir -p ~/.claude/skills/xllm-ai-xllm && curl -fsSL https://workflowstacks.com/api/skills/xllm-ai-xllm/claude-skill -o ~/.claude/skills/xllm-ai-xllm/SKILL.md
Open in another AI app

Opens the app with this repo with the prompt ready to go — no copy-paste needed.

Connect the whole catalog (MCP)
claude mcp add --transport http workflowstacks https://workflowstacks.com/api/mcp

Adds a WorkflowStacks connector to Claude Code: search and load any skill here by chatting.

Quick Actions
Details
Creator
xLLM-AI
Language
C++
Category
ai-agent
Published
8/12/2025

Are you the creator of this tool? Claim your listing → and earn 85% of every sale.