Best by category · September 2026

Best Evals AI skills

6 open-source evals in the catalog, ranked by GitHub stars and refreshed daily. The Hot section uses stars gained in the last 7 days, so newer repos can appear there before they reach the top list.

⭐ Top Evals by stars

Total GitHub stars
  1. 1
    Evals · ★ 9,107 · BoundaryML

    Baml is an AI framework that helps founders engineer better prompts for their AI models, making it easier to get structured and accurate output from AI tools.

  2. 2
    Evals · ★ 6,192 · Trusted-AI

    A Python library that lets you test and improve your ML model’s security by generating and defending against adversarial attacks.

  3. 3
    Evals · ★ 2,152 · center-for-threat-informed-defense

    Provides ready‑to‑use, MITRE ATT&CK‑aligned adversary emulation plans so you can safely test your security controls against real‑world tactics.

  4. 4
    Evals · ★ 685 · eunomia-bpf

    Agentsight provides system-level insights into AI agents using zero instrument tracing with eBPF, allowing founders to monitor and optimize their AI systems without modifying the code.

  5. 5
    Evals · ★ 623 · framerslab

    Agentos is a TypeScript AI agent framework that enables the creation of autonomous AI agents for multi-agent orchestration with cognitive memory and runtime tool forging capabilities.

  6. 6
    Evals · ★ 345 · verifywise-ai

    VerifyWise provides comprehensive AI governance and risk management, supporting 20+ AI frameworks and regulations, including EU AI Act, ISO 42001, and NIST AI RMF.