LLM observability: tools for developers
Monitoring, traces, evaluation and quality tracking for LLM applications. An entry can appear under several use cases when its topics or description provide several explicit signals.
Use case
Activity
Sort by
117 entries in this view.
Tool ranking
Ranked by measured growth, normalized across sources. Raw values keep their source-specific window, shown under each tool; a change requires at least two snapshots from the last 7 days.
- 1Active
agent-stack
SkillProduction patterns for agent orchestrators, harnesses, evals, MCP/A2A interoperability, memory, provider routing, and LLM usage metering.
githubmeasured growthOpen source ↗
Install
git clone https://github.com/ssheleg/agent-stack ~/.claude/skills/agent-stack2GitHub stars+1 (+100.0 %) - 2Active
untell
OtherAI-detector auditing toolkit: measures how often a detector flags genuine human writing, how stable its verdict is across seeds, and whether that verdict survives meaning-preserving edits. Detector-in-the-loop measurement harness. Claude…
githubmeasured growthOpen source ↗
18GitHub starsstable - 3Active
Log Claude Code sessions to Opik, the open-source LLM observability and evaluation platform, built by Comet. Tracing, evaluation, and skills for observable AI applications.
githubmeasured growthOpen source ↗
22GitHub stars+1 (+4.8 %) - 4Active
Deterministic, local-only audit reports for Claude Code AI agent sessions. Rust + SQLite. Zero network calls. MCP server for agents.
githubmeasured growthOpen source ↗
22GitHub stars+2 (+10.0 %) - 5Active
deslop-GPT
SkillDeletion-first Agent Skill for removing test bloat, verification theater, and speculative fallbacks while preserving behavior.
githubmeasured growthOpen source ↗
Install
git clone https://github.com/MrZoyo/deslop-GPT ~/.claude/skills/deslop-GPT87GitHub stars+59 (+210.7 %) - 6Dormant
agenttap
AgentReal-time debugging proxy for Agent2Agent (A2A) multi-agent systems
githubmeasured growthOpen source ↗
3GitHub starsstable - 7Dormant
agentanvil
AgentContract-based testing for LLM agents: hybrid evaluation, multi-agent + A2A, deterministic replay.
githubmeasured growthOpen source ↗
0GitHub starsstable - 8Active
Sentry instrumentation skill for system-behavior tracking
githubmeasured growthOpen source ↗
Install
git clone https://github.com/tortastudios/sentry-instrumentation ~/.claude/skills/sentry-instrumentation24GitHub starsstable - 9Dormant
adaptive-harness
SkillA self-improving harness router for Claude Code.
githubmeasured growthOpen source ↗
Install
/plugin marketplace add SeongwoongCho/adaptive-harness8GitHub starsstable - 10Active
axiom
SkillAxiom is a curated marketplace of shared plugins for Claude Code and Codex.
githubmeasured growthOpen source ↗
Install
git clone https://github.com/netopsengineer/axiom ~/.claude/skills/axiom5GitHub starsstable - 11Dormant
astragraph
AgentPolicy-enforced observability and fail-closed guardrails for MCP/A2A multi-agent systems.
githubmeasured growthOpen source ↗
26GitHub starsstable - 12Active
CodeFlow
Agent面向长程研发任务的 Java Agent Harness,支持 A2A 跨语言协作、可恢复执行、上下文工程与 Eval 驱动开发。
githubmeasured growthOpen source ↗
0GitHub starsstable - 13Dormant
🔍 AI observability skill for Claude Code. Debug LangChain/LangGraph agents by fetching execution traces from LangSmith Studio directly in your terminal.
githubmeasured growthOpen source ↗
Install
git clone https://github.com/OthmanAdi/langsmith-fetch-skill ~/.claude/skills/langsmith-fetch-skill27GitHub starsstable - 14Active
8 MCP SMB products — standalone AI servers for SMBs: Guardrails, FinOps, Observability, Router, Trust Score, Memory, ThinkSecure, A2A Lite. 37 tools, TNC credits billing.
githubmeasured growthOpen source ↗
1GitHub starsstable - 15Dormant
agentops
AgentAgentOps: Multi-agent infrastructure remediation platform. A2A protocol, agent coordination, HITL approval, auto-rollback.
githubmeasured growthOpen source ↗
0GitHub starsstable - 16Active
AgentStack
AgentProvide clear documentation for AgentStack’s MCP protocol, plugins, and ecosystem API with usage examples and tool references.
githubmeasured growthOpen source ↗
0GitHub starsstable - 17Active
Production reliability, benchmarking, and evaluation layer for MCP agent runtimes — protocol negotiation, A2A interop, circuit breakers, connection pooling, adaptive streaming, and autoscaling, with measured throughput/latency and a…
githubmeasured growthOpen source ↗
3GitHub starsstable - 18Active
opus-fable-playbook
SkillMake Claude Opus 4.8 behave like Claude Fable 5 — doctrine output style, drift-catching hooks, and an eval loop against golden Fable transcripts. Claude Code plugin.
githubmeasured growthOpen source ↗
Install
/plugin marketplace add rennf93/opus-fable-playbook34GitHub stars+1 (+3.0 %) - 19Active
adherencia-reglas
SkillMide si tu agente de IA cumple las reglas que le escribiste. Lee el historial local de Claude Code y devuelve un porcentaje por regla. Sin instalar nada, sin red, biblioteca estándar.
githubmeasured growthOpen source ↗
Install
git clone https://github.com/jleonceo/adherencia-reglas ~/.claude/skills/adherencia-reglas1GitHub starsstable - 20Active
claudestat
SkillReal-time execution trace and cost intelligence for Claude Code
githubmeasured growthOpen source ↗
Install
/plugin marketplace add DeibyGS/claudestat34GitHub starsstable - 21Active
skill-receipts
SkillClaude Code skills where every entry ships receipts — accuracy-gated benchmarks against baseline and placebo, rejects published
githubmeasured growthOpen source ↗
Install
git clone https://github.com/sjh9714/skill-receipts ~/.claude/skills/skill-receipts2GitHub starsstable - 22Active
green-agent
AgentA2A green-agent orchestrator for evaluating agents on the AppWorld benchmark, built on the AgentBeats SDK
githubmeasured growthOpen source ↗
0GitHub starsstable - 23Active
bakeoff
SkillTurn one decision into a judged tournament of solutions, then pick the best — a Claude Code skill that generates candidates, auto-derives the rubric, judges independently, and returns a defensible winner.
githubmeasured growthOpen source ↗
Install
/plugin marketplace add CoriChui/bakeoff10GitHub starsstable - 24Active
nora
MCPOpen-source, self-hosted control plane for OpenClaw and Hermes AI-agent fleets on Docker/Kubernetes — REST, CLI, and MCP.
githubmeasured growthOpen source ↗
50GitHub stars+1 (+2.0 %) - 25Active
Skill instalable para Claude Code que mide si tu agente cumple las reglas establecidas. Lee el historial de sesiones de tu disco y devuelve un porcentaje por regla de cumplimiento. Sin instalar nada, sin red, biblioteca estándar.
githubmeasured growthOpen source ↗
Install
/plugin marketplace add jleonceo/skill-adherencia-reglas1GitHub starsstable