LLM observability: tools for developers

Monitoring, traces, evaluation and quality tracking for LLM applications. An entry can appear under several use cases when its topics or description provide several explicit signals.

Use case

Activity

Sort by

117 entries in this view.

Tool ranking

Ranked by measured growth, normalized across sources. Raw values keep their source-specific window, shown under each tool; a change requires at least two snapshots from the last 7 days.

  1. 1
    Active

    Self-hosted AI SRE for Kubernetes — zero-instrumentation eBPF observability plus a copilot that fixes issues through guardrailed, self-verifying actions. BYO-LLM, air-gapped capable.

    githubmeasured growthOpen source ↗

    130GitHub stars-11 (-7.8 %)
  2. 2
    Active

    Open-source Agent Skills for Claude Code and Codex: ship-DAG orchestration, A-F code review, AI evals, CI gates, design, copy, SEO/AEO/GEO, Instagram growth, app shipping, creator rights, and consumer recovery.

    githubmeasured growthOpen source ↗

    129GitHub stars-10 (-7.2 %)
  3. 3
    Active

    A Model Context Protocol (MCP) server for Langfuse, enabling AI agents to query Langfuse trace data for enhanced debugging and observability

    githubmeasured growthOpen source ↗

    105GitHub starsstable
  4. 4

    Skill
    Dormant

    Don't make LLMs honest. Make every factual claim auditable. — An LLM Claim Auditing Layer with T1-T7 truth gradients. 98.1% business effectiveness on LiarBench v0.2.

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/lc198707/anti-lie ~/.claude/skills/anti-lie

    89GitHub starsstable
  5. 5
    Active

    AgentX python SDK. Build multi-agent AI workforce. Run evaluation. Trace your agent. Full Observability.

    githubmeasured growthOpen source ↗

    68GitHub starsstable
  6. 6
    Active

    Open benchmark for Claude Code SEO skills — real headless execution against fixture sites with planted-defect answer keys. Deterministic scoring, pre-registered rubric.

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/aleclindz/seo-skill-bench ~/.claude/skills/seo-skill-bench

    50GitHub stars-6 (-10.7 %)
  7. 7

    Skill
    Active

    Agent skills for Arize — datasets, experiments, and traces via the ax CLI

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/Arize-ai/arize-skills ~/.claude/skills/arize-skills

    47GitHub starsstable
  8. 8
    Dormant

    Pressure-test research claims with falsifiable evidence plans, adversarial checks, frozen verifiers, and proof ledgers.

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/tonyblu331/research-proof ~/.claude/skills/research-proof

    45GitHub starsstable
  9. 9
    Dormant

    🔍 AI observability skill for Claude Code. Debug LangChain/LangGraph agents by fetching execution traces from LangSmith Studio directly in your terminal.

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/OthmanAdi/langsmith-fetch-skill ~/.claude/skills/langsmith-fetch-skill

    27GitHub starsstable
  10. 10

    Agent
    Dormant

    Policy-enforced observability and fail-closed guardrails for MCP/A2A multi-agent systems.

    githubmeasured growthOpen source ↗

    26GitHub starsstable
  11. 11
    Active

    Sentry instrumentation skill for system-behavior tracking

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/tortastudios/sentry-instrumentation ~/.claude/skills/sentry-instrumentation

    24GitHub starsstable
  12. 12
    Active

    Log Claude Code sessions to Opik, the open-source LLM observability and evaluation platform, built by Comet. Tracing, evaluation, and skills for observable AI applications.

    githubmeasured growthOpen source ↗

    21GitHub starsstable
  13. 13

    Skill
    Active

    Measure prompt and skill improvements with blind A/B comparison.

    githubmeasured growthOpen source ↗

    Install /plugin marketplace add shinpr/rashomon

    18GitHub starsstable
  14. 14
    Dormant

    MCP as a Judge: a behavioral MCP that strengthens AI coding assistants via explicit LLM evaluations

    mcpmeasured growthOpen source ↗

    Install claude mcp add mcp-as-a-judge -- uvx mcp-as-a-judge

    17GitHub starsstable
  15. 15
    Active

    Two Claude Code skills for adversarial code review (single-model PAR and multi-model MMAR with cross-critique to catch hallucinations), plus a fixture-based eval suite.

    githubmeasured growthOpen source ↗

    Install /plugin marketplace add prime-radiant-inc/parallel-adversarial-review

    17GitHub starsstable
  16. 16
    Active

    Make Claude write clearly, for everyone.

    githubmeasured growthOpen source ↗

    Install /plugin marketplace add stefanobaghino/simple-output-styles

    16GitHub starsstable
  17. 17
    Active

    Claude Code plugin marketplace — agentic-engineering (spec-driven shape→decide→execute→measure→eval with adversarial review) + github-keeper (audit/elevate READMEs and make a repo open-source-ready).

    githubmeasured growthOpen source ↗

    Install /plugin marketplace add GiustoPiedimonte/agentic-engineering-marketplace

    13GitHub starsstable
  18. 18

    Other
    Active

    A command-line tool to scaffold and manage enterprise-ready AI Agents powered by the A2A (Agent-to-Agent) protocol

    githubmeasured growthOpen source ↗

    13GitHub starsstable
  19. 19
    Active

    Query Spanlens LLM observability from Cursor, Claude Desktop, or Continue via MCP.

    mcpmeasured growthOpen source ↗

    Install claude mcp add mcp-server -- npx @spanlens/mcp-server

    12GitHub starsstable
  20. 20
    Active

    CLI for benchmarks & evals of AI coding agents — on tasks you already understand, using your Claude / Codex / Gemini individual subscriptions or API keys.

    githubmeasured growthOpen source ↗

    12GitHub starsstable
  21. 21

    Skill
    Active

    Turn one decision into a judged tournament of solutions, then pick the best — a Claude Code skill that generates candidates, auto-derives the rubric, judges independently, and returns a defensible winner.

    githubmeasured growthOpen source ↗

    Install /plugin marketplace add CoriChui/bakeoff

    10GitHub starsstable
  22. 22

    Other
    Active

    A Go-native framework for LLM agents, with OpenTelemetry observability built in.

    githubmeasured growthOpen source ↗

    10GitHub starsstable
  23. 23
    Dormant

    Local open-source dev tool to debug, secure, and evaluate LLM agents. Provides static analysis, dynamic security checks, and runtime monitoring - integrates with Cursor and Claude Code.

    githubmeasured growthOpen source ↗

    Install /plugin marketplace add cylestio/agent-inspector

    9GitHub starsstable
  24. 24

    Skill
    Active

    Anchor — the production-grade AGENTS.md template for AI coding agents. 51 sections of battle-tested rules. Keep your agents grounded.

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/peva3/anchor ~/.claude/skills/anchor

    8GitHub starsstable

Learning resources

Ranked by measured growth, normalized across sources. These resources remain available separately and do not take part in the main tool ranking.

  1. 1
    ActiveOpen source ↗

    My complete journey to becoming an Agentic AI Engineer through structured learning, projects, experiments, and production-ready implementations of modern AI systems.

    githubResourcemeasured growth
    18GitHub starsstable