coder_eval — tool profile and history
Test that your Claude Code skills, MCP servers, and CLIs actually work when an agent uses them — sandboxed YAML suites, activation checks, A/B experiments, CI gates.
- Language
- Python
- License
- Apache-2.0
- Created
- Last activity
- Topics
- agent-evaluation
- agent-skills
- agent-testing
- anthropic
- claude
- claude-agent-sdk
- claude-code
- claude-code-plugins-marketplace
- claude-code-skills
- claude-skills
- codex
- coding-agents
- evaluation-framework
- gemini
- github-actions
- long-horizon-agents
- regression-testing
- skill-testing
- swe-bench
- terminal-bench
Current measurement
119GitHub stars+3 (+2.6 %)
measured growthMomentum uses available readings from the last 7 days: measured with at least two comparable readings, estimated otherwise.
Reading history
| Date | Value | Metric |
|---|---|---|
| 119 | GitHub stars | |
| 119 | GitHub stars | |
| 118 | GitHub stars | |
| 116 | GitHub stars | |
| 116 | GitHub stars | |
| 116 | GitHub stars | |
| 116 | GitHub stars | |
| 116 | GitHub stars | |
| 116 | GitHub stars | |
| 116 | GitHub stars | |
| 116 | GitHub stars | |
| 116 | GitHub stars | |
| 116 | GitHub stars | |
| 115 | GitHub stars | |
| 115 | GitHub stars | |
| 113 | GitHub stars | |
| 113 | GitHub stars | |
| 113 | GitHub stars | |
| 111 | GitHub stars | |
| 110 | GitHub stars | |
| 109 | GitHub stars | |
| 109 | GitHub stars |
Classifications
- Domains
- Type
- MCP servers
- Use cases
Neighbouring tools
These entries declare the same topics. No similarity is inferred: only the shared topics are stated.