analytics

Improve LLMs with deepeval

Get reliable evaluation metrics for your language models with deepeval, used by founders
17,589 stars1,798 forksPythonGuide quality 9/10Updated 8/13/2026100% free · open source
What it does

Deepeval provides reliable evaluation metrics for language models, allowing founders to assess their model's performance accurately.

When to use it
  • When you need to compare the performance of different language models
  • When you want to evaluate your language model's performance on a specific task or dataset
  • When you need to identify areas for improvement in your language model
Ready-to-paste prompt
python -m deepeval.examples.eval_example --model-name bert-base-uncased --task-name sentiment-analysis --dataset-name imdb
Heads up: Deepeval requires the `transformers` library, which can be installed using `pip install transformers`, and also requires a GPU with at least 8 GB of VRAM for large-scale evaluations
Saves to your device
Use with Claude
New

Skip the builder — one click puts this in Claude, Cursor, Antigravity and more.

⚡ Runs out of the box

A prompt/skill package — nothing to install beyond adding it to your AI tool.

Try it instantly — no install
Claude Code
mkdir -p ~/.claude/skills/deepeval && curl -fsSL https://workflowstacks.com/api/skills/deepeval/claude-skill -o ~/.claude/skills/deepeval/SKILL.md
Open in another AI app

Opens the app with this repo with the prompt ready to go — no copy-paste needed.

Connect the whole catalog (MCP)
claude mcp add --transport http workflowstacks https://workflowstacks.com/api/mcp

Adds a WorkflowStacks connector to Claude Code: search and load any skill here by chatting.

How Improve LLMs with deepeval works
Codeflow
Free to inspect

Improve LLMs with deepeval is a very large Python project (~149k lines across 1109 code files, plus 474 test files). You install it into your AI tool with one command; there is nothing to run yourself. Last commit this month, Apache-2.0 license, has a test suite.

Size
Very large codebase
~149k lines · 1109 code files · days to read — use, don't read
Setup
Install as a skill / plugin
Add it to Claude Code (or your AI tool) with one command — 6 skills inside. Nothing to run yourself.
Runs on
Inside your AI tool
Helper scripts use Python · needs API keys
Python 81%TypeScript 19%
Where to start reading
  1. 1
    README.md
    Start here — what it does and how to install it
  2. 2
    .env.example
    The API keys and settings you must provide
  3. 3
    deepeval/__init__.py
    Inside deepeval/ — the main logic begins here
  4. 4
    examples/community/chatbot_evaluation/dataset.json
    A worked example — copy this to get going
What's in each folder
deepeval/Core code — the actual logic782 files
skills/Prompts, skills & agent definitions28 files
docs/Documentation638 files
examples/Examples you can copy19 files
typescript/Folder459 files
voice_simulations_ws/Folder11 files
scripts/Helper scripts4 files
.claude-plugin/Plugin manifest — what gets installed2 files
READMEHas testsDocumentedCI checksExamples includedApache-2.0 licenseUpdated this month
Quick Actions
Details
Creator
confident-ai
Language
Python
Category
analytics
Published
8/10/2023

Are you the creator of this tool? Claim your listing → and earn 85% of every sale.

Ready-to-run pack · one-time $29

Competitor Watch

Get emailed the moment a rival changes their price or pitch.

Includes the tested workflow and a written setup playbook. Needs: No extra API — just watches the pages you list.

See what you get