SkillEvaluator — tool profile and history

Multi-tier framework for evaluating AI agent skills with quality gates, semantic overlap detection, synthetic evaluation dataset generation, and live agent evaluation that measures how skills affect agent behavior.

Active

Canonical source ↗

Language
Python
License
Apache-2.0
Created
Last activity
Topics
  • agent-evaluation
  • agent-security
  • agent-skills
  • agentic-ai
  • benchmark
  • claude-code
  • codex
  • evaluate
  • evaluation
  • security-scanner
  • skill-eval
  • skill-evals
  • skill-evaluation
  • skill-evaluator
  • skills

Install git clone https://github.com/NVIDIA/SkillEvaluator ~/.claude/skills/SkillEvaluator

Current measurement

363GitHub stars+70 (+23.9 %)
measured growth

Momentum uses available readings from the last 7 days: measured with at least two comparable readings, estimated otherwise.

Reading history

90 days · 90 maximum readings
DateValueMetric
363GitHub stars
351GitHub stars
348GitHub stars
332GitHub stars
326GitHub stars
305GitHub stars
293GitHub stars
266GitHub stars
245GitHub stars
209GitHub stars
168GitHub stars
94GitHub stars

Classifications

Domains
Type
Agent skills
Use cases