LLM observability: tools for developers
Monitoring, traces, evaluation and quality tracking for LLM applications. An entry can appear under several use cases when its topics or description provide several explicit signals.
Use case
Activity
Sort by
123 entries in this view.
Tool ranking
Ranked by creation date, newest first; undated entries come last. Raw values keep their source-specific window, shown under each tool; a change requires at least two snapshots from the last 7 days.
- 1Active
CodeFlow
Agent面向长程研发任务的 Java Agent Harness,支持 A2A 跨语言协作、可恢复执行、上下文工程与 Eval 驱动开发。
githubmeasured growthOpen source ↗
0GitHub starsstable - 2Active
deslop-GPT
SkillDeletion-first Agent Skill for removing test bloat, verification theater, and speculative fallbacks while preserving behavior.
githubmeasured growthOpen source ↗
Install
git clone https://github.com/MrZoyo/deslop-GPT ~/.claude/skills/deslop-GPT55GitHub stars+27 (+96.4 %) - 3Active
dsh-plugins
AgentGeneric DeepSeek Harness (dsh) plugins: A2A protocol server, session storage mirror, and Langfuse observability.
githubmeasured growthOpen source ↗
1GitHub starsstable - 4Active
Production reliability, benchmarking, and evaluation layer for MCP agent runtimes — protocol negotiation, A2A interop, circuit breakers, connection pooling, adaptive streaming, and autoscaling, with measured throughput/latency and a…
githubmeasured growthOpen source ↗
3GitHub starsstable - 5Active
Multi-agent quality gate skill for Claude Code that researches, reviews, tests and challenges AI-generated work before the final answer.
githubmeasured growthOpen source ↗
Install
git clone https://github.com/ma-nucho-pro/supervisor-skill-claude ~/.claude/skills/supervisor-skill-claude2GitHub starsstable - 6Active
Mini Program Engineering Skill Suite is an Agent Skill suite for evidence‑first mini‑program development. It helps agents bring WeChat or other mini‑program projects from vague intent to reliable engineering work: project intake, product…
githubestimated momentumOpen source ↗
Install
git clone https://github.com/NocodeMrLi/mini-program-engineering-skill-suite ~/.claude/skills/mini-program-engineering-skill-suite39GitHub stars— - 7Active
Deterministic, local-only audit reports for Claude Code AI agent sessions. Rust + SQLite. Zero network calls. MCP server for agents.
githubmeasured growthOpen source ↗
22GitHub stars+2 (+10.0 %) - 8Active
Amazon Bedrock AgentCore enterprise platform accelerator with AWS CDK, Terraform organization guardrails, MCP/A2A agents, security, memory, and observability.
githubmeasured growthOpen source ↗
5GitHub stars+1 (+25.0 %) - 9Active
SkillCorpus
SkillOpen-source infrastructure that turns scattered SKILL.md files into curated, retrieval-ready agent-skill corpora—with retrieval and evaluation tooling included.
githubmeasured growthOpen source ↗
Install
git clone https://github.com/EverMind-AI/SkillCorpus ~/.claude/skills/SkillCorpus424GitHub stars+170 (+66.9 %) - 10Active
evoagent-os
AgentLocal-first control plane for durable, governed agent teams: DAG orchestration, approvals, memory, signed skills, trace contracts, and realtime.
githubmeasured growthOpen source ↗
0GitHub starsstable - 11Active
agent-stack
SkillProduction patterns for agent orchestrators, harnesses, evals, MCP/A2A interoperability, memory, provider routing, and LLM usage metering.
githubmeasured growthOpen source ↗
Install
git clone https://github.com/ssheleg/agent-stack ~/.claude/skills/agent-stack2GitHub stars+1 (+100.0 %) - 12Active
craft-skills
SkillResearch-backed, eval-driven skills for AI agents
githubmeasured growthOpen source ↗
Install
git clone https://github.com/ZSeven-W/craft-skills ~/.claude/skills/craft-skills142GitHub stars+18 (+14.5 %) - 13Active
simple-output-styles
SkillMake Claude write clearly, for everyone.
githubmeasured growthOpen source ↗
Install
/plugin marketplace add stefanobaghino/simple-output-styles16GitHub starsstable - 14Active
Skill instalable para Claude Code que mide si tu agente cumple las reglas establecidas. Lee el historial de sesiones de tu disco y devuelve un porcentaje por regla de cumplimiento. Sin instalar nada, sin red, biblioteca estándar.
githubmeasured growthOpen source ↗
Install
/plugin marketplace add jleonceo/skill-adherencia-reglas1GitHub starsstable - 15Active
adherencia-reglas
SkillMide si tu agente de IA cumple las reglas que le escribiste. Lee el historial local de Claude Code y devuelve un porcentaje por regla. Sin instalar nada, sin red, biblioteca estándar.
githubmeasured growthOpen source ↗
Install
git clone https://github.com/jleonceo/adherencia-reglas ~/.claude/skills/adherencia-reglas1GitHub starsstable - 16Active
Agent Skill that asks: should there be an agentic system at all? Then designs, builds, audits, verifies, and ports systems—from workflows to governed teams.
githubmeasured growthOpen source ↗
Install
git clone https://github.com/ingcontartese-netizen/multi-agent-system-architect-v2 ~/.claude/skills/multi-agent-system-architect-v21GitHub starsstable - 17Active
VeriRun
OtherEvidence-first infrastructure for reproducible, isolated executable evaluation and online rewards.
githubmeasured growthOpen source ↗
140GitHub stars+24 (+20.7 %) - 18Active
skill-receipts
SkillClaude Code skills where every entry ships receipts — accuracy-gated benchmarks against baseline and placebo, rejects published
githubmeasured growthOpen source ↗
Install
git clone https://github.com/sjh9714/skill-receipts ~/.claude/skills/skill-receipts2GitHub starsstable - 19Active
boundary-bench
AgentDeterministic crash tests for agent control planes: paired worlds, hard stops, budgets, scopes, approvals, and hash-linked receipts.
githubmeasured growthOpen source ↗
1GitHub starsstable - 20Active
adlc-team-skills
Skill🐙 ADLC Team Skills — Agentic SDLC for Engineering Teams
githubmeasured growthOpen source ↗
Install
git clone https://github.com/tikalk/adlc-team-skills ~/.claude/skills/adlc-team-skills132GitHub stars+1 (+0.76 %) - 21Active
nudge
MCPA typed, replayable, budget-aware programming language for LLM agents. Compiles to Python & TypeScript.
githubmeasured growthOpen source ↗
2GitHub starsstable - 22Active
a2a-otel-kit
MCPVendor-neutral OpenTelemetry tracing for A2A agents and MCP services, with W3C context propagation and privacy-safe telemetry.
githubmeasured growthOpen source ↗
1GitHub starsstable - 23Active
Architecture-first Python scaffold for an auditable multi-agent corporate credit desk using A2A, MCP, deterministic credit policies, model routing, and OpenTelemetry.
githubmeasured growthOpen source ↗
0GitHub starsstable
Learning resources
Ranked by creation date, newest first; undated entries come last. These resources remain available separately and do not take part in the main tool ranking.
- 1ActiveOpen source ↗
Agentic_AI_Engineer
AgentMy complete journey to becoming an Agentic AI Engineer through structured learning, projects, experiments, and production-ready implementations of modern AI systems.
githubResourcemeasured growth18GitHub starsstable - 2ActiveOpen source ↗
trigger_tree
SkillDocumentation-discovery telemetry for Claude Code — heat/cold maps, health grade, evidence-backed router fixes. 100% local, zero tokens. /tt
githubInstall
Resourcemeasured growth/plugin marketplace add Hedde/trigger_tree14GitHub stars+1 (+7.7 %)