LLM observability: tools for developers

Monitoring, traces, evaluation and quality tracking for LLM applications. An entry can appear under several use cases when its topics or description provide several explicit signals.

Use case

Activity

Sort by

122 entries in this view.

Tool ranking

Ranked by normalized popularity across sources. Raw values keep their source-specific window, shown under each tool; a change requires at least two snapshots from the last 7 days.

  1. 1
    Active

    Optimize any AI agent’s skills, tools/MCP, and prompts against your own evals.

    githubestimated momentumOpen source ↗

    47GitHub stars
  2. 2
    Active

    Mini Program Engineering Skill Suite is an Agent Skill suite for evidence‑first mini‑program development. It helps agents bring WeChat or other mini‑program projects from vague intent to reliable engineering work: project intake, product…

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/NocodeMrLi/mini-program-engineering-skill-suite ~/.claude/skills/mini-program-engineering-skill-suite

    43GitHub stars+4 (+10.3 %)
  3. 3

    Skill
    Active

    Real-time execution trace and cost intelligence for Claude Code

    githubmeasured growthOpen source ↗

    Install /plugin marketplace add DeibyGS/claudestat

    34GitHub starsstable
  4. 4
    Active

    Make Claude Opus 4.8 behave like Claude Fable 5 — doctrine output style, drift-catching hooks, and an eval loop against golden Fable transcripts. Claude Code plugin.

    githubmeasured growthOpen source ↗

    Install /plugin marketplace add rennf93/opus-fable-playbook

    34GitHub stars+1 (+3.0 %)
  5. 5
    Dormant

    🔍 AI observability skill for Claude Code. Debug LangChain/LangGraph agents by fetching execution traces from LangSmith Studio directly in your terminal.

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/OthmanAdi/langsmith-fetch-skill ~/.claude/skills/langsmith-fetch-skill

    27GitHub starsstable
  6. 6

    Agent
    Dormant

    Policy-enforced observability and fail-closed guardrails for MCP/A2A multi-agent systems.

    githubmeasured growthOpen source ↗

    26GitHub starsstable
  7. 7
    Active

    Sentry instrumentation skill for system-behavior tracking

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/tortastudios/sentry-instrumentation ~/.claude/skills/sentry-instrumentation

    24GitHub starsstable
  8. 8
    Active

    Deterministic, local-only audit reports for Claude Code AI agent sessions. Rust + SQLite. Zero network calls. MCP server for agents.

    githubmeasured growthOpen source ↗

    22GitHub stars+2 (+10.0 %)
  9. 9
    Active

    Log Claude Code sessions to Opik, the open-source LLM observability and evaluation platform, built by Comet. Tracing, evaluation, and skills for observable AI applications.

    githubmeasured growthOpen source ↗

    22GitHub stars+1 (+4.8 %)
  10. 10

    Other
    Active

    AI-detector auditing toolkit: measures how often a detector flags genuine human writing, how stable its verdict is across seeds, and whether that verdict survives meaning-preserving edits. Detector-in-the-loop measurement harness. Claude…

    githubmeasured growthOpen source ↗

    18GitHub starsstable
  11. 11

    Skill
    Active

    Measure prompt and skill improvements with blind A/B comparison.

    githubmeasured growthOpen source ↗

    Install /plugin marketplace add shinpr/rashomon

    18GitHub starsstable
  12. 12
    Active

    Two Claude Code skills for adversarial code review (single-model PAR and multi-model MMAR with cross-critique to catch hallucinations), plus a fixture-based eval suite.

    githubmeasured growthOpen source ↗

    Install /plugin marketplace add prime-radiant-inc/parallel-adversarial-review

    17GitHub starsstable
  13. 13
    Dormant

    MCP as a Judge: a behavioral MCP that strengthens AI coding assistants via explicit LLM evaluations

    mcpmeasured growthOpen source ↗

    Install claude mcp add mcp-as-a-judge -- uvx mcp-as-a-judge

    17GitHub starsstable
  14. 14
    Active

    Make Claude write clearly, for everyone.

    githubmeasured growthOpen source ↗

    Install /plugin marketplace add stefanobaghino/simple-output-styles

    16GitHub starsstable
  15. 15

    Other
    Active

    A command-line tool to scaffold and manage enterprise-ready AI Agents powered by the A2A (Agent-to-Agent) protocol

    githubmeasured growthOpen source ↗

    14GitHub stars+1 (+7.7 %)
  16. 16
    Active

    Claude Code plugin marketplace — agentic-engineering (spec-driven shape→decide→execute→measure→eval with adversarial review) + github-keeper (audit/elevate READMEs and make a repo open-source-ready).

    githubmeasured growthOpen source ↗

    Install /plugin marketplace add GiustoPiedimonte/agentic-engineering-marketplace

    13GitHub starsstable
  17. 17

    Other
    Dormant

    A Claude Code skill that adds a rubric-based eval layer to any agent project. Framework-agnostic — generates rubric, test cases, judge prompt, and harness. Returns a weighted score plus a judge-leniency signal.

    githubmeasured growthOpen source ↗

    13GitHub starsstable
  18. 18
    Active

    Query Spanlens LLM observability from Cursor, Claude Desktop, or Continue via MCP.

    mcpmeasured growthOpen source ↗

    Install claude mcp add mcp-server -- npx @spanlens/mcp-server

    12GitHub starsstable
  19. 19

    Skill
    Active

    Production-grade AI coding rules for Cursor and Claude Code. 15 rules + 9 doc templates + skills + agents + MCP setup. Drop into any project.

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/aiagentwithdhruv/ai-dev-stack ~/.claude/skills/ai-dev-stack

    10GitHub stars+1 (+11.1 %)
  20. 20
    Active

    Audit which Claude Code skills you actually use — surface dead installs and hallucinated invocations from your session logs.

    githubmeasured growthOpen source ↗

    Install /plugin marketplace add sfrangulov/skill-graveyard

    10GitHub stars+1 (+11.1 %)
  21. 21

    Skill
    Active

    Turn one decision into a judged tournament of solutions, then pick the best — a Claude Code skill that generates candidates, auto-derives the rubric, judges independently, and returns a defensible winner.

    githubmeasured growthOpen source ↗

    Install /plugin marketplace add CoriChui/bakeoff

    10GitHub starsstable
  22. 22

    Other
    Active

    A Go-native framework for LLM agents, with OpenTelemetry observability built in.

    githubmeasured growthOpen source ↗

    10GitHub starsstable

Learning resources

Ranked by normalized popularity across sources. These resources remain available separately and do not take part in the main tool ranking.

  1. 1
    ActiveOpen source ↗

    50+ curated LLM observability tools PLUS 26 Agent Skills (several with runnable, unit-tested scripts) to build, evaluate, debug, secure & monitor reliable LLM apps. Tracing, evals, guardrails, LLMOps.

    github

    Install git clone https://github.com/ContextJet-ai/awesome-llm-observability ~/.claude/skills/awesome-llm-observability

    Resourcemeasured growth
    33GitHub stars+3 (+10.0 %)
  2. 2
    ActiveOpen source ↗

    My complete journey to becoming an Agentic AI Engineer through structured learning, projects, experiments, and production-ready implementations of modern AI systems.

    githubResourcemeasured growth
    18GitHub starsstable
  3. 3
    ActiveOpen source ↗

    Documentation-discovery telemetry for Claude Code — heat/cold maps, health grade, evidence-backed router fixes. 100% local, zero tokens. /tt

    github

    Install /plugin marketplace add Hedde/trigger_tree

    Resourcemeasured growth
    14GitHub stars+1 (+7.7 %)