coder_eval — tool profile and history

Test that your Claude Code skills, MCP servers, and CLIs actually work when an agent uses them — sandboxed YAML suites, activation checks, A/B experiments, CI gates.

Active

Canonical source ↗

Language
Python
License
Apache-2.0
Created
Last activity
Topics
  • agent-evaluation
  • agent-skills
  • agent-testing
  • anthropic
  • claude
  • claude-agent-sdk
  • claude-code
  • claude-code-plugins-marketplace
  • claude-code-skills
  • claude-skills
  • codex
  • coding-agents
  • evaluation-framework
  • gemini
  • github-actions
  • long-horizon-agents
  • regression-testing
  • skill-testing
  • swe-bench
  • terminal-bench

Current measurement

119GitHub stars+3 (+2.6 %)
measured growth

Momentum uses available readings from the last 7 days: measured with at least two comparable readings, estimated otherwise.

Reading history

90 days · 90 maximum readings
DateValueMetric
119GitHub stars
119GitHub stars
118GitHub stars
116GitHub stars
116GitHub stars
116GitHub stars
116GitHub stars
116GitHub stars
116GitHub stars
116GitHub stars
116GitHub stars
116GitHub stars
116GitHub stars
115GitHub stars
115GitHub stars
113GitHub stars
113GitHub stars
113GitHub stars
111GitHub stars
110GitHub stars
109GitHub stars
109GitHub stars

Classifications

Domains
Type
MCP servers
Use cases

These entries declare the same topics. No similarity is inferred: only the shared topics are stated.