LLM observability: tools for developers

Monitoring, traces, evaluation and quality tracking for LLM applications. An entry can appear under several use cases when its topics or description provide several explicit signals.

Use case

Activity

Sort by

122 entries in this view.

Tool ranking

Ranked by normalized popularity across sources. Raw values keep their source-specific window, shown under each tool; a change requires at least two snapshots from the last 7 days.

  1. 1
    Active

    Dashboard for monitoring claude code sessions.

    githubmeasured growthOpen source ↗

    Install /plugin marketplace add JayantDevkar/claude-code-karma

    323GitHub stars+2 (+0.62 %)
  2. 2

    MCP
    Active

    AURA is a production-tested SRE agent platform you can deploy in minutes. AURA handles the guardrails, APIs, state management, streaming, and failure handling required to put AI to work safely on production infrastructure.

    githubmeasured growthOpen source ↗

    254GitHub stars-3 (-1.2 %)
  3. 3
    Active

    🪢 Langfuse documentation -- Langfuse is the open source LLM Engineering Platform. Observability, evals, prompt management, playground and metrics to debug and improve LLM apps

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/langfuse/langfuse-docs ~/.claude/skills/langfuse-docs

    239GitHub stars+3 (+1.3 %)
  4. 4
    Active

    🟪 Open-source runtime that ships any LangGraph or Google ADK agent as a production-ready FastAPI service. Bundled , AG-UI copilotkit API, chat UI, 15+ guardrails, MCP, OpenTelemetry, OIDC. One pip install. Self-hosted, no vendor lock-in.

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/Idun-Group/idun-agent-platform ~/.claude/skills/idun-agent-platform

    198GitHub stars-1 (-0.50 %)
  5. 5

    Agent
    Active

    The Operating System for Scalable Enterprise AI Agents - Run, orchestrate, and deploy Compliant Enterprise AI Agents at scale across frameworks, without lock-in, rewrites or fragile glue code. Native support for MCP, A2A. Interface with all…

    githubmeasured growthOpen source ↗

    156GitHub stars+19 (+13.9 %)
  6. 6

    Skill
    Active

    Research-backed, eval-driven skills for AI agents

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/ZSeven-W/craft-skills ~/.claude/skills/craft-skills

    142GitHub stars+18 (+14.5 %)
  7. 7

    Other
    Active

    Evidence-first infrastructure for reproducible, isolated executable evaluation and online rewards.

    githubmeasured growthOpen source ↗

    140GitHub stars+24 (+20.7 %)
  8. 8
    Active

    🐙 ADLC Team Skills — Agentic SDLC for Engineering Teams

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/tikalk/adlc-team-skills ~/.claude/skills/adlc-team-skills

    132GitHub stars+1 (+0.76 %)
  9. 9
    Active

    Skills, prompts, and instructions for building AI agents on top of Dynatrace production context

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/Dynatrace/dynatrace-for-ai ~/.claude/skills/dynatrace-for-ai

    132GitHub stars+4 (+3.1 %)
  10. 10
    Active

    Self-hosted AI SRE for Kubernetes — zero-instrumentation eBPF observability plus a copilot that fixes issues through guardrailed, self-verifying actions. BYO-LLM, air-gapped capable.

    githubmeasured growthOpen source ↗

    131GitHub stars-4 (-3.0 %)
  11. 11
    Active

    74 open-source Agent Skills for Claude Code and Codex: AI SEO, AEO and GEO, code review with an A-F ship grade, CI gates, AI evals, design systems, conversion copy, Instagram growth, iOS and Android app shipping, creator rights, and…

    githubmeasured growthOpen source ↗

    129GitHub stars-4 (-3.0 %)
  12. 12
    Active

    A Model Context Protocol (MCP) server for Langfuse, enabling AI agents to query Langfuse trace data for enhanced debugging and observability

    githubmeasured growthOpen source ↗

    105GitHub starsstable
  13. 13
    Active

    Self improving agents through iterations

    githubmeasured growthOpen source ↗

    105GitHub stars+1 (+0.96 %)
  14. 14
    Active

    Vendor-neutral OpenTelemetry skills for AI coding agents, grounded in upstream sources

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/ollygarden/opentelemetry-agent-skills ~/.claude/skills/opentelemetry-agent-skills

    97GitHub stars+4 (+4.3 %)
  15. 15

    Skill
    Dormant

    Don't make LLMs honest. Make every factual claim auditable. — An LLM Claim Auditing Layer with T1-T7 truth gradients. 98.1% business effectiveness on LiarBench v0.2.

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/lc198707/anti-lie ~/.claude/skills/anti-lie

    89GitHub starsstable
  16. 16
    Active

    MCP server for Langfuse LLM observability — trace and observation analysis.

    mcpmeasured growthOpen source ↗

    Install claude mcp add langfuse -- npx langfuse-observability-mcp-server

    78GitHub starsstable
  17. 17

    Skill
    Active

    local-first analytics for AI agent skills

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/crafter-station/skill-kit ~/.claude/skills/skill-kit

    77GitHub stars+1 (+1.3 %)
  18. 18
    Active

    Agent Skill evaluation harness for paired variants, trace artifacts, and runner adapters

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/adewale/skill-eval-harness ~/.claude/skills/skill-eval-harness

    73GitHub stars+4 (+5.8 %)
  19. 19
    Active

    AgentX python SDK. Build multi-agent AI workforce. Run evaluation. Trace your agent. Full Observability.

    githubmeasured growthOpen source ↗

    68GitHub starsstable
  20. 20

    Skill
    Active

    Deletion-first Agent Skill for removing test bloat, verification theater, and speculative fallbacks while preserving behavior.

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/MrZoyo/deslop-GPT ~/.claude/skills/deslop-GPT

    55GitHub stars+27 (+96.4 %)
  21. 21
    Active

    Open benchmark for Claude Code SEO skills — real headless execution against fixture sites with planted-defect answer keys. Deterministic scoring, pre-registered rubric.

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/aleclindz/seo-skill-bench ~/.claude/skills/seo-skill-bench

    51GitHub stars-5 (-8.9 %)
  22. 22

    MCP
    Active

    Open-source, self-hosted control plane for OpenClaw and Hermes AI-agent fleets on Docker/Kubernetes — REST, CLI, and MCP.

    githubmeasured growthOpen source ↗

    50GitHub stars+2 (+4.2 %)
  23. 23

    Skill
    Active

    Agent skills for Arize — datasets, experiments, and traces via the ax CLI

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/Arize-ai/arize-skills ~/.claude/skills/arize-skills

    47GitHub starsstable
  24. 24
    Active

    Optimize any AI agent’s skills, tools/MCP, and prompts against your own evals.

    githubmeasured growthOpen source ↗

    47GitHub stars+1 (+2.2 %)

Learning resources

Ranked by normalized popularity across sources. These resources remain available separately and do not take part in the main tool ranking.

  1. 1
    ActiveOpen source ↗

    Curated + scored map of official MCP servers and agents for DevOps, Cloud, SRE, and Platform Engineering — every entry rated on production access, approval gates, and audit evidence.

    githubResourcemeasured growth
    77GitHub starsstable