ai-agent

Speculators: Fast LLM Inference

Get unified decoding algorithms for LLM inference with Speculators, a Python library.
754 stars194 forksPythonGuide quality 8/10Updated 8/21/2026100% free · open source
What it does

Speculators provides a unified library for building, evaluating, and storing speculative decoding algorithms for Large Language Model (LLM) inference

When to use it
  • When you need to optimize LLM inference for your startup's natural language processing tasks
  • When you want to compare and evaluate different decoding algorithms for your LLM models
  • When you need to store and manage multiple speculative decoding algorithms for different LLM models
Ready-to-paste prompt
python examples/example.py --model_name-or-path vllm/lm_base --decoder topk --topk 5
Heads up: Make sure you have the necessary dependencies installed, including PyTorch and transformers, as Speculators relies on these libraries to function correctly
Saves to your device
Use with Claude
New

Skip the builder — one click puts this in Claude, Cursor, Antigravity and more.

✅ Light setup

Installs with a command or two; your AI agent can do it for you.

Try it instantly — no install
Claude Code
mkdir -p ~/.claude/skills/speculators && curl -fsSL https://workflowstacks.com/api/skills/speculators/claude-skill -o ~/.claude/skills/speculators/SKILL.md
Open in another AI app

Opens the app with this repo with the prompt ready to go — no copy-paste needed.

Connect the whole catalog (MCP)
claude mcp add --transport http workflowstacks https://workflowstacks.com/api/mcp

Adds a WorkflowStacks connector to Claude Code: search and load any skill here by chatting.

Quick Actions
Details
Creator
vllm-project
Language
Python
Category
ai-agent
Published
4/8/2025

Are you the creator of this tool? Claim your listing → and earn 85% of every sale.