ai-agent

LLM Coding Benchmark

Simple benchmark to test the most popular open source and commercial LLMs with automated OpenCode
332 stars39 forksPythonUpdated 9/15/2026100% free · open source
What it does

Runs a ready‑made coding benchmark that evaluates LLMs (open‑source or commercial) on the OpenCode dataset and returns accuracy and runtime metrics.

When to use it
  • You want to compare several LLM providers (e.g., OpenAI, Anthropic, Llama‑2) on realistic programming tasks before choosing a model for your product.
  • You need an objective, repeatable score to justify a budget request for a higher‑cost LLM.
  • You are building an AI‑coding assistant and want a quick sanity‑check of how new model releases perform on code generation.
Ready-to-paste prompt
llm-coding-benchmark --model openai/gpt-4o --model anthropic/claude-3-opus --tasks data/open_code/tasks.jsonl --output benchmark_results.json
Heads up: The tool requires a valid API key for every commercial model you list; if the key is missing or the model name is mistyped the run aborts with a 401 error before any benchmark is executed.
Saves to your device
Use with Claude
New

Skip the builder — one click puts this in Claude, Cursor, Antigravity and more.

⚡ Runs out of the box

A prompt/skill package — nothing to install beyond adding it to your AI tool.

Try it instantly — no install
Claude Code
mkdir -p ~/.claude/skills/llm-coding-benchmark && curl -fsSL https://workflowstacks.com/api/skills/llm-coding-benchmark/claude-skill -o ~/.claude/skills/llm-coding-benchmark/SKILL.md
Open in another AI app

Opens the app with this repo with the prompt ready to go — no copy-paste needed.

Connect the whole catalog (MCP)
claude mcp add --transport http workflowstacks https://workflowstacks.com/api/mcp

Adds a WorkflowStacks connector to Claude Code: search and load any skill here by chatting.

How LLM Coding Benchmark works
Codeflow
Free to inspect

LLM Coding Benchmark is mostly documents (386 doc files, 174 small scripts) — something you read, not something you run. There is nothing to install. Last commit this month, no license file.

Size
Mostly documents
386 documents · 174 small scripts · days to read — use, don't read
Setup
No code to run
It's prompts / instructions. Copy them into your AI tool.
Runs on
Inside your AI tool
Nothing to install on your machine
Python 80%Ruby 9%HTML 7%JavaScript 2%Go 1%
Where to start reading
  1. 1
    README.md
    Start here — what it does and how to install it
  2. 2
    CLAUDE.md
    The instructions the AI actually follows
What's in each folder
prompts/Prompts, skills & agent definitions6 files
docs/Documentation31 files
results-v3/Folder346 files
results-v2/Folder309 files
results-v4/Folder192 files
benchmark-v4/Folder190 files
results/Folder178 files
benchmark-v3/Folder148 files
READMEHas testsDocumentedNo licenseUpdated this month
Quick Actions
Details
Creator
akitaonrails
Language
Python
Category
ai-agent
Published
4/5/2026

Are you the creator of this tool? and earn 85% of every sale.

Show it off in your README:

Featured on WorkflowStacks
[![Featured on WorkflowStacks](https://workflowstacks.com/api/badge/llm-coding-benchmark.svg)](https://workflowstacks.com/skills/llm-coding-benchmark?utm_source=github&utm_medium=badge)
🔥 Hot this week, in your inbox
Skills like this, every Monday.

The five fastest-growing open-source AI skills, ranked by GitHub star growth. One email a week, unsubscribe anytime.