Best Evals AI skills
6 open-source evals in the catalog, ranked by GitHub stars and refreshed daily. The Hot section uses stars gained in the last 7 days, so newer repos can appear there before they reach the top list.
⭐ Top Evals by stars
Total GitHub stars- 1Evals · ★ 9,107 · BoundaryML
Baml is an AI framework that helps founders engineer better prompts for their AI models, making it easier to get structured and accurate output from AI tools.
- 2Evals · ★ 6,192 · Trusted-AI
A Python library that lets you test and improve your ML model’s security by generating and defending against adversarial attacks.
- 3Evals · ★ 2,152 · center-for-threat-informed-defense
Provides ready‑to‑use, MITRE ATT&CK‑aligned adversary emulation plans so you can safely test your security controls against real‑world tactics.
- 4Evals · ★ 685 · eunomia-bpf
Agentsight provides system-level insights into AI agents using zero instrument tracing with eBPF, allowing founders to monitor and optimize their AI systems without modifying the code.
- 5Evals · ★ 623 · framerslab
Agentos is a TypeScript AI agent framework that enables the creation of autonomous AI agents for multi-agent orchestration with cognitive memory and runtime tool forging capabilities.
- 6Evals · ★ 345 · verifywise-ai
VerifyWise provides comprehensive AI governance and risk management, supporting 20+ AI frameworks and regulations, including EU AI Act, ISO 42001, and NIST AI RMF.