LLM observability: tools for developers

Monitoring, traces, evaluation and quality tracking for LLM applications. An entry can appear under several use cases when its topics or description provide several explicit signals.

Use case

Activity

Sort by

74 entries in this view.

Tool ranking

Ranked by measured growth, normalized across sources. Raw values keep their source-specific window, shown under each tool; a change requires at least two snapshots from the last 7 days.

  1. 1
    Active

    Agent Skill evaluation harness for paired variants, trace artifacts, and runner adapters

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/adewale/skill-eval-harness ~/.claude/skills/skill-eval-harness

    73GitHub stars+4 (+5.8 %)
  2. 2
    Active

    🪢 Langfuse documentation -- Langfuse is the open source LLM Engineering Platform. Observability, evals, prompt management, playground and metrics to debug and improve LLM apps

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/langfuse/langfuse-docs ~/.claude/skills/langfuse-docs

    239GitHub stars+3 (+1.3 %)
  3. 3

    MCP
    Active

    Open-source, self-hosted control plane for OpenClaw and Hermes AI-agent fleets on Docker/Kubernetes — REST, CLI, and MCP.

    githubmeasured growthOpen source ↗

    50GitHub stars+2 (+4.2 %)
  4. 4
    Active

    Deterministic, local-only audit reports for Claude Code AI agent sessions. Rust + SQLite. Zero network calls. MCP server for agents.

    githubmeasured growthOpen source ↗

    22GitHub stars+2 (+10.0 %)
  5. 5
    Active

    Dashboard for monitoring claude code sessions.

    githubmeasured growthOpen source ↗

    Install /plugin marketplace add JayantDevkar/claude-code-karma

    323GitHub stars+2 (+0.62 %)
  6. 6

    Other
    Active

    Open-source observability tool that uses AI agents to self-heal your software

    githubmeasured growthOpen source ↗

    1 404GitHub stars+2 (+0.14 %)
  7. 7

    Other
    Active

    AI SRE AgenticOps for Kubernetes and cloud infrastructure.

    githubmeasured growthOpen source ↗

    782GitHub stars+1 (+0.13 %)
  8. 8

    Skill
    Active

    Production patterns for agent orchestrators, harnesses, evals, MCP/A2A interoperability, memory, provider routing, and LLM usage metering.

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/ssheleg/agent-stack ~/.claude/skills/agent-stack

    2GitHub stars+1 (+100.0 %)
  9. 9
    Active

    🐙 ADLC Team Skills — Agentic SDLC for Engineering Teams

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/tikalk/adlc-team-skills ~/.claude/skills/adlc-team-skills

    132GitHub stars+1 (+0.76 %)
  10. 10
    Active

    Self improving agents through iterations

    githubmeasured growthOpen source ↗

    105GitHub stars+1 (+0.96 %)
  11. 11

    Skill
    Active

    local-first analytics for AI agent skills

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/crafter-station/skill-kit ~/.claude/skills/skill-kit

    77GitHub stars+1 (+1.3 %)
  12. 12

    Skill
    Active

    Real-time execution trace and cost intelligence for Claude Code

    githubmeasured growthOpen source ↗

    Install /plugin marketplace add DeibyGS/claudestat

    34GitHub stars+1 (+3.0 %)
  13. 13

    Other
    Active

    A command-line tool to scaffold and manage enterprise-ready AI Agents powered by the A2A (Agent-to-Agent) protocol

    githubmeasured growthOpen source ↗

    14GitHub stars+1 (+7.7 %)
  14. 14
    Active

    Amazon Bedrock AgentCore enterprise platform accelerator with AWS CDK, Terraform organization guardrails, MCP/A2A agents, security, memory, and observability.

    githubmeasured growthOpen source ↗

    5GitHub stars+1 (+25.0 %)
  15. 15
    Active

    Optimize any AI agent’s skills, tools/MCP, and prompts against your own evals.

    githubmeasured growthOpen source ↗

    47GitHub stars+1 (+2.2 %)
  16. 16
    Active

    Production reliability, benchmarking, and evaluation layer for MCP agent runtimes — protocol negotiation, A2A interop, circuit breakers, connection pooling, adaptive streaming, and autoscaling, with measured throughput/latency and a…

    githubmeasured growthOpen source ↗

    3GitHub starsstable
  17. 17
    Active

    Architecture-first Python scaffold for an auditable multi-agent corporate credit desk using A2A, MCP, deterministic credit policies, model routing, and OpenTelemetry.

    githubmeasured growthOpen source ↗

    0GitHub starsstable
  18. 18

    Agent
    Active

    Generic DeepSeek Harness (dsh) plugins: A2A protocol server, session storage mirror, and Langfuse observability.

    githubmeasured growthOpen source ↗

    1GitHub starsstable
  19. 19

    MCP
    Active

    A typed, replayable, budget-aware programming language for LLM agents. Compiles to Python & TypeScript.

    githubmeasured growthOpen source ↗

    2GitHub starsstable
  20. 20
    Active

    Multi-agent quality gate skill for Claude Code that researches, reviews, tests and challenges AI-generated work before the final answer.

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/ma-nucho-pro/supervisor-skill-claude ~/.claude/skills/supervisor-skill-claude

    2GitHub starsstable
  21. 21
    Active

    Deterministic crash tests for agent control planes: paired worlds, hard stops, budgets, scopes, approvals, and hash-linked receipts.

    githubmeasured growthOpen source ↗

    1GitHub starsstable
  22. 22
    Active

    Local-first Agentic AI Infrastructure Platform with LLM Gateway, Agent DAG Runtime, MCP Tool Hub, A2A Agent Mesh, Local RAG, Tool Sandbox and Observability.

    githubmeasured growthOpen source ↗

    1GitHub starsstable
  23. 23

    Agent
    Active

    面向长程研发任务的 Java Agent Harness,支持 A2A 跨语言协作、可恢复执行、上下文工程与 Eval 驱动开发。

    githubmeasured growthOpen source ↗

    0GitHub starsstable

Learning resources

Ranked by measured growth, normalized across sources. These resources remain available separately and do not take part in the main tool ranking.

  1. 1
    ActiveOpen source ↗

    50+ curated LLM observability tools PLUS 26 Agent Skills (several with runnable, unit-tested scripts) to build, evaluate, debug, secure & monitor reliable LLM apps. Tracing, evals, guardrails, LLMOps.

    github

    Install git clone https://github.com/ContextJet-ai/awesome-llm-observability ~/.claude/skills/awesome-llm-observability

    Resourcemeasured growth
    32GitHub stars+2 (+6.7 %)
  2. 2
    ActiveOpen source ↗

    Documentation-discovery telemetry for Claude Code — heat/cold maps, health grade, evidence-backed router fixes. 100% local, zero tokens. /tt

    github

    Install /plugin marketplace add Hedde/trigger_tree

    Resourcemeasured growth
    14GitHub stars+1 (+7.7 %)