ClawProBench
ClawProBench runs repeatable, deterministic benchmarks on LLM agents inside the OpenClaw runtime so you can measure performance with confidence.
- •You need to compare multiple LLM‑agent implementations before picking one for production.
- •Your product relies on OpenClaw‑compatible agents and you want reliable, repeatable grading across runs.
- •You want to embed a benchmark suite into CI/CD to catch regressions in agent behavior.
clawprobench run --bench examples/bench.yaml --agent ./agents/customer‑support-agent --repeat 5 --output results/cs_report.json
Skip the builder — one click puts this in Claude, Cursor, Antigravity and more.
Expect 20–40 minutes in a terminal — or let your AI agent drive it.
mkdir -p ~/.claude/skills/clawprobench && curl -fsSL https://workflowstacks.com/api/skills/clawprobench/claude-skill -o ~/.claude/skills/clawprobench/SKILL.mdOpens the app with this repo with the prompt ready to go — no copy-paste needed.
ClawProBench is a very large Rust project (~1.2M lines across 2782 code files, plus 773 test files). It is a full software project: use it through its install path rather than reading it end to end. Last commit a month ago, Apache-2.0 license, has a test suite.
- 1README.mdStart here — what it does and how to install it
- 2run.pyWhere the program starts running
- 3requirements.txtDependencies and the commands it exposes
- 4ironclaw/src/app.rsInside ironclaw/src/ — the main logic begins here
Skip the builder — one click puts this in Claude, Cursor, Antigravity and more.
Expect 20–40 minutes in a terminal — or let your AI agent drive it.
mkdir -p ~/.claude/skills/clawprobench && curl -fsSL https://workflowstacks.com/api/skills/clawprobench/claude-skill -o ~/.claude/skills/clawprobench/SKILL.mdOpens the app with this repo with the prompt ready to go — no copy-paste needed.
Are you the creator of this tool? Claim your listing → and earn 85% of every sale.
Related skills
More ai-agent tools founders pair with this one.