ai-agent

llm_benchmark: AI Performance

Measure AI model performance with llm_benchmark, used by founders, for informed decisions.
1,586 stars23 forksGuide quality 8/10Updated 9/7/2026100% free · open source
What it does

llm_benchmark measures the performance of AI models, providing founders with actionable insights to inform their decisions.

When to use it
  • When evaluating the accuracy of different AI models for a startup's application
  • To compare the performance of various AI models on a specific dataset
  • Before deploying an AI model to production, to ensure it meets performance requirements
Ready-to-paste prompt
python scripts/run_benchmark.py --model_name bert-base-uncased --dataset_name sst-2
Heads up: The repository uses Python 3.8 or later, and some models may require specific dependencies or GPU acceleration, so ensure your environment meets these requirements before running the benchmark
Saves to your device
Use with Claude
New

Skip the builder — one click puts this in Claude, Cursor, Antigravity and more.

🛠️ Technical setup

Expect 20–40 minutes in a terminal — or let your AI agent drive it.

Try it instantly — no install
Claude Code
mkdir -p ~/.claude/skills/llm-benchmark && curl -fsSL https://workflowstacks.com/api/skills/llm-benchmark/claude-skill -o ~/.claude/skills/llm-benchmark/SKILL.md
Open in another AI app

Opens the app with this repo with the prompt ready to go — no copy-paste needed.

Connect the whole catalog (MCP)
claude mcp add --transport http workflowstacks https://workflowstacks.com/api/mcp

Adds a WorkflowStacks connector to Claude Code: search and load any skill here by chatting.

How llm_benchmark: AI Performance works
Codeflow
Free to inspect

Llm_benchmark: AI Performance is a medium project (~3.9k lines across 4 code files). Expect some technical setup — comfortable with a terminal, or ask a developer. Last commit this month, no license file, no tests found.

Size
Medium codebase
~3.9k lines · 4 code files · 32 min skim
Setup
Some technical setup
Comfortable with a terminal? 20–40 min. Otherwise ask a dev.
Runs on
See README
No API keys detected
Where to start reading
  1. 1
    README.md
    Start here — what it does and how to install it
What's in each folder
docs/Documentation116 files
logic/Folder19 files
code/Folder5 files
vision/Folder4 files
code_v2/Folder1 files
READMENo tests foundDocumentedNo CINo licenseUpdated this month
Quick Actions
Details
Creator
llm2014
Category
ai-agent
Published
2/7/2025

Are you the creator of this tool? Claim your listing → and earn 85% of every sale.