SkillEvaluator — tool profile and history
Multi-tier framework for evaluating AI agent skills with quality gates, semantic overlap detection, synthetic evaluation dataset generation, and live agent evaluation that measures how skills affect agent behavior.
- Language
- Python
- License
- Apache-2.0
- Created
- Last activity
- Topics
- agent-evaluation
- agent-security
- agent-skills
- agentic-ai
- benchmark
- claude-code
- codex
- evaluate
- evaluation
- security-scanner
- skill-eval
- skill-evals
- skill-evaluation
- skill-evaluator
- skills
Install git clone https://github.com/NVIDIA/SkillEvaluator ~/.claude/skills/SkillEvaluator
Current measurement
363GitHub stars+70 (+23.9 %)
measured growthMomentum uses available readings from the last 7 days: measured with at least two comparable readings, estimated otherwise.
Reading history
| Date | Value | Metric |
|---|---|---|
| 363 | GitHub stars | |
| 351 | GitHub stars | |
| 348 | GitHub stars | |
| 332 | GitHub stars | |
| 326 | GitHub stars | |
| 305 | GitHub stars | |
| 293 | GitHub stars | |
| 266 | GitHub stars | |
| 245 | GitHub stars | |
| 209 | GitHub stars | |
| 168 | GitHub stars | |
| 94 | GitHub stars |
Classifications
- Domains
- Type
- Agent skills
- Use cases