LLM observability: tools for developers

Monitoring, traces, evaluation and quality tracking for LLM applications. An entry can appear under several use cases when its topics or description provide several explicit signals.

Use case

Activity

Sort by

123 entries in this view.

Tool ranking

Ranked by measured growth, normalized across sources. Raw values keep their source-specific window, shown under each tool; a change requires at least two snapshots from the last 7 days.

  1. 1
    Active

    Amazon Bedrock AgentCore enterprise platform accelerator with AWS CDK, Terraform organization guardrails, MCP/A2A agents, security, memory, and observability.

    githubmeasured growthOpen source ↗

    5GitHub stars+1 (+25.0 %)
  2. 2

    Skill
    Active

    Real-time execution trace and cost intelligence for Claude Code

    githubmeasured growthOpen source ↗

    Install /plugin marketplace add DeibyGS/claudestat

    34GitHub stars+3 (+9.7 %)
  3. 3
    Active

    🫖 Status page with uptime monitoring & API monitoring as code 🫖

    githubmeasured growthOpen source ↗

    9 036GitHub stars+25 (+0.28 %)
  4. 4

    Other
    Active

    A cron that remembers what it did—a CLI job scheduler with run history, captured output, and a live terminal dashboard.

    githubmeasured growthOpen source ↗

    216GitHub stars+7 (+3.3 %)
  5. 5

    Other
    Active

    AI-detector auditing toolkit: measures how often a detector flags genuine human writing, how stable its verdict is across seeds, and whether that verdict survives meaning-preserving edits. Detector-in-the-loop measurement harness. Claude…

    githubmeasured growthOpen source ↗

    19GitHub stars+2 (+11.8 %)
  6. 6
    Active

    Agent Skill evaluation harness for paired variants, trace artifacts, and runner adapters

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/adewale/skill-eval-harness ~/.claude/skills/skill-eval-harness

    72GitHub stars+4 (+5.9 %)
  7. 7

    Other
    Active

    eBPF-based Networking, Security, and Observability

    githubmeasured growthOpen source ↗

    25 028GitHub stars+24 (+0.10 %)
  8. 8

    Agent
    Active

    DataBuff is an AI-native APM built on Opentelemetry,with multi-agent troubleshooting out of the box.

    githubmeasured growthOpen source ↗

    615GitHub stars+10 (+1.7 %)
  9. 9
    Active

    Deterministic, local-only audit reports for Claude Code AI agent sessions. Rust + SQLite. Zero network calls. MCP server for agents.

    githubmeasured growthOpen source ↗

    22GitHub stars+2 (+10.0 %)
  10. 10
    Active

    Skills, prompts, and instructions for building AI agents on top of Dynatrace production context

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/Dynatrace/dynatrace-for-ai ~/.claude/skills/dynatrace-for-ai

    131GitHub stars+5 (+4.0 %)
  11. 11

    Skill
    Active

    OpenJudge: A Unified Framework for Holistic Evaluation and Quality Rewards

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/agentscope-ai/OpenJudge ~/.claude/skills/OpenJudge

    807GitHub stars+10 (+1.3 %)
  12. 12
    Active

    Vendor-neutral OpenTelemetry skills for AI coding agents, grounded in upstream sources

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/ollygarden/opentelemetry-agent-skills ~/.claude/skills/opentelemetry-agent-skills

    96GitHub stars+4 (+4.3 %)
  13. 13

    Skill
    Active

    Production-grade AI coding rules for Cursor and Claude Code. 15 rules + 9 doc templates + skills + agents + MCP setup. Drop into any project.

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/aiagentwithdhruv/ai-dev-stack ~/.claude/skills/ai-dev-stack

    10GitHub stars+1 (+11.1 %)
  14. 14

    Other
    Active

    Open-source observability tool that uses AI agents to self-heal your software

    githubmeasured growthOpen source ↗

    1 404GitHub stars+10 (+0.72 %)
  15. 15

    Other
    Dormant

    A Claude Code skill that adds a rubric-based eval layer to any agent project. Framework-agnostic — generates rubric, test cases, judge prompt, and harness. Returns a weighted score plus a judge-leniency signal.

    githubmeasured growthOpen source ↗

    13GitHub stars+1 (+8.3 %)
  16. 16

    Skill
    Active

    A skill creator that proves its skills work. Evidence-driven skill creation for Claude Code and Codex: baseline-tested generation, per-skill regression evals, ecosystem doctor, cross-runtime compile, and an opt-in proactive advisor.

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/tripleyak/SkillForge ~/.claude/skills/SkillForge

    884GitHub stars+6 (+0.68 %)
  17. 17

    Skill
    Active

    The container platform tailored for Kubernetes multi-cloud, datacenter, and edge management ⎈ 🖥 ☁️

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/kubesphere/kubesphere ~/.claude/skills/kubesphere

    17 035GitHub stars+6 (+0.04 %)
  18. 18
    Active

    Make Claude Opus 4.8 behave like Claude Fable 5 — doctrine output style, drift-catching hooks, and an eval loop against golden Fable transcripts. Claude Code plugin.

    githubmeasured growthOpen source ↗

    Install /plugin marketplace add rennf93/opus-fable-playbook

    33GitHub stars+1 (+3.1 %)
  19. 19

    Agent
    Active

    eBPF-powered network observability for Kubernetes. Indexes L4/L7 traffic with full K8s context, decrypts TLS without keys. Queryable by AI agents via MCP and humans via dashboard.

    githubmeasured growthOpen source ↗

    12 063GitHub stars+5 (+0.04 %)
  20. 20
    Active

    Optimize any AI agent’s skills, tools/MCP, and prompts against your own evals.

    githubmeasured growthOpen source ↗

    47GitHub stars+1 (+2.2 %)
  21. 21

    MCP
    Active

    Open-source, self-hosted control plane for OpenClaw and Hermes AI-agent fleets on Docker/Kubernetes — REST, CLI, and MCP.

    githubmeasured growthOpen source ↗

    49GitHub stars+1 (+2.1 %)
  22. 22
    Active

    🪢 Langfuse documentation -- Langfuse is the open source LLM Engineering Platform. Observability, evals, prompt management, playground and metrics to debug and improve LLM apps

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/langfuse/langfuse-docs ~/.claude/skills/langfuse-docs

    237GitHub stars+2 (+0.85 %)
  23. 23

    MCP
    Active

    AURA is a production-tested SRE agent platform you can deploy in minutes. AURA handles the guardrails, APIs, state management, streaming, and failure handling required to put AI to work safely on production infrastructure.

    githubmeasured growthOpen source ↗

    258GitHub stars+2 (+0.78 %)
  24. 24
    Active

    Dashboard for monitoring claude code sessions.

    githubmeasured growthOpen source ↗

    Install /plugin marketplace add JayantDevkar/claude-code-karma

    321GitHub stars+2 (+0.63 %)

Learning resources

Ranked by measured growth, normalized across sources. These resources remain available separately and do not take part in the main tool ranking.

  1. 1
    ActiveOpen source ↗

    50+ curated LLM observability tools PLUS 26 Agent Skills (several with runnable, unit-tested scripts) to build, evaluate, debug, secure & monitor reliable LLM apps. Tracing, evals, guardrails, LLMOps.

    github

    Install git clone https://github.com/ContextJet-ai/awesome-llm-observability ~/.claude/skills/awesome-llm-observability

    Resourcemeasured growth
    31GitHub stars+1 (+3.3 %)