LLM observability: tools for developers

Monitoring, traces, evaluation and quality tracking for LLM applications. An entry can appear under several use cases when its topics or description provide several explicit signals.

Use case

Activity

Sort by

117 entries in this view.

Tool ranking

Ranked by measured growth, normalized across sources. Raw values keep their source-specific window, shown under each tool; a change requires at least two snapshots from the last 7 days.

  1. 1

    Skill
    Active

    Production patterns for agent orchestrators, harnesses, evals, MCP/A2A interoperability, memory, provider routing, and LLM usage metering.

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/ssheleg/agent-stack ~/.claude/skills/agent-stack

    2GitHub stars+1 (+100.0 %)
  2. 2

    Other
    Active

    AI-detector auditing toolkit: measures how often a detector flags genuine human writing, how stable its verdict is across seeds, and whether that verdict survives meaning-preserving edits. Detector-in-the-loop measurement harness. Claude…

    githubmeasured growthOpen source ↗

    18GitHub starsstable
  3. 3
    Active

    Log Claude Code sessions to Opik, the open-source LLM observability and evaluation platform, built by Comet. Tracing, evaluation, and skills for observable AI applications.

    githubmeasured growthOpen source ↗

    22GitHub stars+1 (+4.8 %)
  4. 4
    Active

    Deterministic, local-only audit reports for Claude Code AI agent sessions. Rust + SQLite. Zero network calls. MCP server for agents.

    githubmeasured growthOpen source ↗

    22GitHub stars+2 (+10.0 %)
  5. 5

    Skill
    Active

    Deletion-first Agent Skill for removing test bloat, verification theater, and speculative fallbacks while preserving behavior.

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/MrZoyo/deslop-GPT ~/.claude/skills/deslop-GPT

    87GitHub stars+59 (+210.7 %)
  6. 6

    Agent
    Dormant

    Real-time debugging proxy for Agent2Agent (A2A) multi-agent systems

    githubmeasured growthOpen source ↗

    3GitHub starsstable
  7. 7

    Agent
    Dormant

    Contract-based testing for LLM agents: hybrid evaluation, multi-agent + A2A, deterministic replay.

    githubmeasured growthOpen source ↗

    0GitHub starsstable
  8. 8
    Active

    Sentry instrumentation skill for system-behavior tracking

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/tortastudios/sentry-instrumentation ~/.claude/skills/sentry-instrumentation

    24GitHub starsstable
  9. 9
    Dormant

    A self-improving harness router for Claude Code.

    githubmeasured growthOpen source ↗

    Install /plugin marketplace add SeongwoongCho/adaptive-harness

    8GitHub starsstable
  10. 10

    Skill
    Active

    Axiom is a curated marketplace of shared plugins for Claude Code and Codex.

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/netopsengineer/axiom ~/.claude/skills/axiom

    5GitHub starsstable
  11. 11

    Agent
    Dormant

    Policy-enforced observability and fail-closed guardrails for MCP/A2A multi-agent systems.

    githubmeasured growthOpen source ↗

    26GitHub starsstable
  12. 12

    Agent
    Active

    面向长程研发任务的 Java Agent Harness,支持 A2A 跨语言协作、可恢复执行、上下文工程与 Eval 驱动开发。

    githubmeasured growthOpen source ↗

    0GitHub starsstable
  13. 13
    Dormant

    🔍 AI observability skill for Claude Code. Debug LangChain/LangGraph agents by fetching execution traces from LangSmith Studio directly in your terminal.

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/OthmanAdi/langsmith-fetch-skill ~/.claude/skills/langsmith-fetch-skill

    27GitHub starsstable
  14. 14
    Active

    8 MCP SMB products — standalone AI servers for SMBs: Guardrails, FinOps, Observability, Router, Trust Score, Memory, ThinkSecure, A2A Lite. 37 tools, TNC credits billing.

    githubmeasured growthOpen source ↗

    1GitHub starsstable
  15. 15

    Agent
    Dormant

    AgentOps: Multi-agent infrastructure remediation platform. A2A protocol, agent coordination, HITL approval, auto-rollback.

    githubmeasured growthOpen source ↗

    0GitHub starsstable
  16. 16

    Agent
    Active

    Provide clear documentation for AgentStack’s MCP protocol, plugins, and ecosystem API with usage examples and tool references.

    githubmeasured growthOpen source ↗

    0GitHub starsstable
  17. 17
    Active

    Production reliability, benchmarking, and evaluation layer for MCP agent runtimes — protocol negotiation, A2A interop, circuit breakers, connection pooling, adaptive streaming, and autoscaling, with measured throughput/latency and a…

    githubmeasured growthOpen source ↗

    3GitHub starsstable
  18. 18
    Active

    Make Claude Opus 4.8 behave like Claude Fable 5 — doctrine output style, drift-catching hooks, and an eval loop against golden Fable transcripts. Claude Code plugin.

    githubmeasured growthOpen source ↗

    Install /plugin marketplace add rennf93/opus-fable-playbook

    34GitHub stars+1 (+3.0 %)
  19. 19
    Active

    Mide si tu agente de IA cumple las reglas que le escribiste. Lee el historial local de Claude Code y devuelve un porcentaje por regla. Sin instalar nada, sin red, biblioteca estándar.

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/jleonceo/adherencia-reglas ~/.claude/skills/adherencia-reglas

    1GitHub starsstable
  20. 20

    Skill
    Active

    Real-time execution trace and cost intelligence for Claude Code

    githubmeasured growthOpen source ↗

    Install /plugin marketplace add DeibyGS/claudestat

    34GitHub starsstable
  21. 21
    Active

    Claude Code skills where every entry ships receipts — accuracy-gated benchmarks against baseline and placebo, rejects published

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/sjh9714/skill-receipts ~/.claude/skills/skill-receipts

    2GitHub starsstable
  22. 22

    Agent
    Active

    A2A green-agent orchestrator for evaluating agents on the AppWorld benchmark, built on the AgentBeats SDK

    githubmeasured growthOpen source ↗

    0GitHub starsstable
  23. 23

    Skill
    Active

    Turn one decision into a judged tournament of solutions, then pick the best — a Claude Code skill that generates candidates, auto-derives the rubric, judges independently, and returns a defensible winner.

    githubmeasured growthOpen source ↗

    Install /plugin marketplace add CoriChui/bakeoff

    10GitHub starsstable
  24. 24

    MCP
    Active

    Open-source, self-hosted control plane for OpenClaw and Hermes AI-agent fleets on Docker/Kubernetes — REST, CLI, and MCP.

    githubmeasured growthOpen source ↗

    50GitHub stars+1 (+2.0 %)
  25. 25
    Active

    Skill instalable para Claude Code que mide si tu agente cumple las reglas establecidas. Lee el historial de sesiones de tu disco y devuelve un porcentaje por regla de cumplimiento. Sin instalar nada, sin red, biblioteca estándar.

    githubmeasured growthOpen source ↗

    Install /plugin marketplace add jleonceo/skill-adherencia-reglas

    1GitHub starsstable