LLM observability: tools for developers
Monitoring, traces, evaluation and quality tracking for LLM applications. An entry can appear under several use cases when its topics or description provide several explicit signals.
Use case
Activity
Sort by
122 entries in this view.
Tool ranking
Ranked by measured growth, normalized across sources. Raw values keep their source-specific window, shown under each tool; a change requires at least two snapshots from the last 7 days.
- 1Active
AgentStack
AgentProvide clear documentation for AgentStack’s MCP protocol, plugins, and ecosystem API with usage examples and tool references.
githubmeasured growthOpen source ↗
0GitHub starsstable - 2Active
AgentX-Python
OtherAgentX python SDK. Build multi-agent AI workforce. Run evaluation. Trace your agent. Full Observability.
githubmeasured growthOpen source ↗
68GitHub starsstable - 3Active
Production reliability, benchmarking, and evaluation layer for MCP agent runtimes — protocol negotiation, A2A interop, circuit breakers, connection pooling, adaptive streaming, and autoscaling, with measured throughput/latency and a…
githubmeasured growthOpen source ↗
3GitHub starsstable - 4Active
opus-fable-playbook
SkillMake Claude Opus 4.8 behave like Claude Fable 5 — doctrine output style, drift-catching hooks, and an eval loop against golden Fable transcripts. Claude Code plugin.
githubmeasured growthOpen source ↗
Install
/plugin marketplace add rennf93/opus-fable-playbook34GitHub stars+1 (+3.0 %) - 5Active
adherencia-reglas
SkillMide si tu agente de IA cumple las reglas que le escribiste. Lee el historial local de Claude Code y devuelve un porcentaje por regla. Sin instalar nada, sin red, biblioteca estándar.
githubmeasured growthOpen source ↗
Install
git clone https://github.com/jleonceo/adherencia-reglas ~/.claude/skills/adherencia-reglas1GitHub starsstable - 6Active
claudestat
SkillReal-time execution trace and cost intelligence for Claude Code
githubmeasured growthOpen source ↗
Install
/plugin marketplace add DeibyGS/claudestat34GitHub starsstable - 7Active
seo-skill-bench
SkillOpen benchmark for Claude Code SEO skills — real headless execution against fixture sites with planted-defect answer keys. Deterministic scoring, pre-registered rubric.
githubmeasured growthOpen source ↗
Install
git clone https://github.com/aleclindz/seo-skill-bench ~/.claude/skills/seo-skill-bench51GitHub stars-5 (-8.9 %) - 8Active
skill-receipts
SkillClaude Code skills where every entry ships receipts — accuracy-gated benchmarks against baseline and placebo, rejects published
githubmeasured growthOpen source ↗
Install
git clone https://github.com/sjh9714/skill-receipts ~/.claude/skills/skill-receipts2GitHub starsstable - 9Active
green-agent
AgentA2A green-agent orchestrator for evaluating agents on the AppWorld benchmark, built on the AgentBeats SDK
githubmeasured growthOpen source ↗
0GitHub starsstable - 10Active
SkillForge
SkillA skill creator that proves its skills work. Evidence-driven skill creation for Claude Code and Codex: baseline-tested generation, per-skill regression evals, ecosystem doctor, cross-runtime compile, and an opt-in proactive advisor.
githubmeasured growthOpen source ↗
Install
git clone https://github.com/tripleyak/SkillForge ~/.claude/skills/SkillForge885GitHub starsstable - 11Active
bakeoff
SkillTurn one decision into a judged tournament of solutions, then pick the best — a Claude Code skill that generates candidates, auto-derives the rubric, judges independently, and returns a defensible winner.
githubmeasured growthOpen source ↗
Install
/plugin marketplace add CoriChui/bakeoff10GitHub starsstable - 12Active
nora
MCPOpen-source, self-hosted control plane for OpenClaw and Hermes AI-agent fleets on Docker/Kubernetes — REST, CLI, and MCP.
githubmeasured growthOpen source ↗
50GitHub stars+1 (+2.0 %) - 13Active
yao-meta-skill
SkillYAO = Yielding AI Outcomes. A rigorous engineering, evaluation, governance, and portability system for reusable agent skills.
githubmeasured growthOpen source ↗
Install
git clone https://github.com/yaojingang/yao-meta-skill ~/.claude/skills/yao-meta-skill2 596GitHub stars+37 (+1.4 %) - 14Active
Skill instalable para Claude Code que mide si tu agente cumple las reglas establecidas. Lee el historial de sesiones de tu disco y devuelve un porcentaje por regla de cumplimiento. Sin instalar nada, sin red, biblioteca estándar.
githubmeasured growthOpen source ↗
Install
/plugin marketplace add jleonceo/skill-adherencia-reglas1GitHub starsstable - 15Dormant
forge-skills
SkillAn assembly line for AI software development. 35 skills, 11 agent personas, 29 commands. From raw idea to shipped code.
githubmeasured growthOpen source ↗
Install
/plugin marketplace add aneja5/forge-skills3GitHub starsstable - 16Active
skill-graveyard
SkillAudit which Claude Code skills you actually use — surface dead installs and hallucinated invocations from your session logs.
githubmeasured growthOpen source ↗
Install
/plugin marketplace add sfrangulov/skill-graveyard10GitHub starsstable - 17Dormant
A Python proof-of-concept for tracing multi-turn Agent-to-Agent (A2A) conversations as a single unified MLflow trace for LLM observability and evaluation.
githubmeasured growthOpen source ↗
0GitHub starsstable - 18Active
ai-dev-stack
SkillProduction-grade AI coding rules for Cursor and Claude Code. 15 rules + 9 doc templates + skills + agents + MCP setup. Drop into any project.
githubmeasured growthOpen source ↗
Install
git clone https://github.com/aiagentwithdhruv/ai-dev-stack ~/.claude/skills/ai-dev-stack10GitHub starsstable - 19Active
DriftSentinel
AgentAgent 降智检测与自愈公评网络 — an immune system for the AI agent society
githubmeasured growthOpen source ↗
1GitHub starsstable - 20Active
claude-code-karma
SkillDashboard for monitoring claude code sessions.
githubmeasured growthOpen source ↗
Install
/plugin marketplace add JayantDevkar/claude-code-karma324GitHub stars+3 (+0.93 %) - 21Active
galdor
OtherA Go-native framework for LLM agents, with OpenTelemetry observability built in.
githubmeasured growthOpen source ↗
10GitHub starsstable - 22Active
agent-kernel
AgentThe Operating System for Scalable Enterprise AI Agents - Run, orchestrate, and deploy Compliant Enterprise AI Agents at scale across frameworks, without lock-in, rewrites or fragile glue code. Native support for MCP, A2A. Interface with all…
githubmeasured growthOpen source ↗
166GitHub stars+25 (+17.7 %) - 23Dormant
kiboserve
MCPOpen-source Python framework to deploy AI agents via HTTP, A2A, and MCP with built-in observability
githubmeasured growthOpen source ↗
5GitHub starsstable - 24Active
boundary-bench
AgentDeterministic crash tests for agent control planes: paired worlds, hard stops, budgets, scopes, approvals, and hash-linked receipts.
githubmeasured growthOpen source ↗
1GitHub starsstable - 25Active
dsh-plugins
AgentGeneric DeepSeek Harness (dsh) plugins: A2A protocol server, session storage mirror, and Langfuse observability.
githubmeasured growthOpen source ↗
1GitHub starsstable