LLM observability: tools for developers

Monitoring, traces, evaluation and quality tracking for LLM applications. An entry can appear under several use cases when its topics or description provide several explicit signals.

Use case

Activity

Sort by

110 entries in this view.

Tool ranking

Ranked by normalized popularity across sources. Raw values keep their source-specific window, shown under each tool; a change requires at least two snapshots from the last 7 days.

  1. 1

    Other
    Active

    Test your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare performance of GPT, Claude, Gemini, DeepSeek, and more. Simple declarative configs with command line and CI/CD integration. Used by OpenAI…

    githubmeasured growthOpen source ↗

    24 768GitHub stars+142 (+0.58 %)
  2. 2

    Other
    Active

    SigNoz is an open-source, OpenTelemetry-native observability platform for your team and their AI agents. Get logs, metrics, and traces in one tool with features like APM, distributed tracing, log management, infra monitoring, etc. Combined…

    githubmeasured growthOpen source ↗

    32 003GitHub stars+50 (+0.16 %)
  3. 3

    Other
    Active

    The fastest path to AI-powered full stack observability, even for lean teams.

    githubmeasured growthOpen source ↗

    80 412GitHub stars+85 (+0.11 %)
  4. 4

    Skill
    Active

    The container platform tailored for Kubernetes multi-cloud, datacenter, and edge management ⎈ 🖥 ☁️

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/kubesphere/kubesphere ~/.claude/skills/kubesphere

    17 037GitHub stars+8 (+0.05 %)
  5. 5

    Other
    Active

    the LLM vulnerability scanner

    githubmeasured growthOpen source ↗

    9 079GitHub stars+31 (+0.34 %)
  6. 6
    Active

    YAO = Yielding AI Outcomes. A rigorous engineering, evaluation, governance, and portability system for reusable agent skills.

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/yaojingang/yao-meta-skill ~/.claude/skills/yao-meta-skill

    2 577GitHub stars+35 (+1.4 %)
  7. 7
    Active

    Build tested agent skills and govern their lifecycle through a user-defined marketplace: evidence, discovery, updates, rollback, quarantine, and 17-platform distribution.

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/FrancyJGLisboa/agent-skill-creator ~/.claude/skills/agent-skill-creator

    2 367GitHub stars+28 (+1.2 %)
  8. 8

    Skill
    Active

    The Fable Workflow: how Claude Fable 5 worked, distilled into skills any model can run, with the eval that keeps it honest. Think / act / prove.

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/Sahir619/fable-method ~/.claude/skills/fable-method

    2 272GitHub stars+16 (+0.71 %)
  9. 9
    Active

    A Go framework for building production agent systems with graph workflows, tools, memory, A2A, AG-UI, MCP, evaluation, and observability.

    githubmeasured growthOpen source ↗

    1 759GitHub stars+18 (+1.0 %)
  10. 10

    Skill
    Active

    A skill creator that proves its skills work. Evidence-driven skill creation for Claude Code and Codex: baseline-tested generation, per-skill regression evals, ecosystem doctor, cross-runtime compile, and an opt-in proactive advisor.

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/tripleyak/SkillForge ~/.claude/skills/SkillForge

    884GitHub stars+1 (+0.11 %)
  11. 11

    Skill
    Active

    OpenJudge: A Unified Framework for Holistic Evaluation and Quality Rewards

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/agentscope-ai/OpenJudge ~/.claude/skills/OpenJudge

    809GitHub stars+6 (+0.75 %)
  12. 12

    Other
    Active

    AI SRE AgenticOps for Kubernetes and cloud infrastructure.

    githubmeasured growthOpen source ↗

    782GitHub stars+1 (+0.13 %)
  13. 13
    Active

    A test runner for agentskills.io-style AI agent skills

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/darkrishabh/agent-skills-eval ~/.claude/skills/agent-skills-eval

    719GitHub stars+13 (+1.8 %)
  14. 14

    Skill
    Active

    Open-source infrastructure that turns scattered SKILL.md files into curated, retrieval-ready agent-skill corpora—with retrieval and evaluation tooling included.

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/EverMind-AI/SkillCorpus ~/.claude/skills/SkillCorpus

    424GitHub stars+170 (+66.9 %)
  15. 15
    Active

    Multi-tier framework for evaluating AI agent skills with quality gates, semantic overlap detection, synthetic evaluation dataset generation, and live agent evaluation that measures how skills affect agent behavior.

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/NVIDIA/SkillEvaluator ~/.claude/skills/SkillEvaluator

    389GitHub stars+57 (+17.2 %)
  16. 16
    Active

    Dashboard for monitoring claude code sessions.

    githubmeasured growthOpen source ↗

    Install /plugin marketplace add JayantDevkar/claude-code-karma

    323GitHub stars+2 (+0.62 %)
  17. 17

    MCP
    Active

    AURA is a production-tested SRE agent platform you can deploy in minutes. AURA handles the guardrails, APIs, state management, streaming, and failure handling required to put AI to work safely on production infrastructure.

    githubmeasured growthOpen source ↗

    254GitHub stars-3 (-1.2 %)
  18. 18
    Active

    🪢 Langfuse documentation -- Langfuse is the open source LLM Engineering Platform. Observability, evals, prompt management, playground and metrics to debug and improve LLM apps

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/langfuse/langfuse-docs ~/.claude/skills/langfuse-docs

    239GitHub stars+3 (+1.3 %)
  19. 19
    Active

    🟪 Open-source runtime that ships any LangGraph or Google ADK agent as a production-ready FastAPI service. Bundled , AG-UI copilotkit API, chat UI, 15+ guardrails, MCP, OpenTelemetry, OIDC. One pip install. Self-hosted, no vendor lock-in.

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/Idun-Group/idun-agent-platform ~/.claude/skills/idun-agent-platform

    198GitHub stars-1 (-0.50 %)
  20. 20

    Agent
    Active

    The Operating System for Scalable Enterprise AI Agents - Run, orchestrate, and deploy Compliant Enterprise AI Agents at scale across frameworks, without lock-in, rewrites or fragile glue code. Native support for MCP, A2A. Interface with all…

    githubmeasured growthOpen source ↗

    156GitHub stars+19 (+13.9 %)
  21. 21

    Skill
    Active

    Research-backed, eval-driven skills for AI agents

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/ZSeven-W/craft-skills ~/.claude/skills/craft-skills

    142GitHub stars+18 (+14.5 %)
  22. 22
    Active

    🐙 ADLC Team Skills — Agentic SDLC for Engineering Teams

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/tikalk/adlc-team-skills ~/.claude/skills/adlc-team-skills

    132GitHub stars+1 (+0.76 %)
  23. 23
    Active

    Skills, prompts, and instructions for building AI agents on top of Dynatrace production context

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/Dynatrace/dynatrace-for-ai ~/.claude/skills/dynatrace-for-ai

    132GitHub stars+4 (+3.1 %)
  24. 24
    Active

    74 open-source Agent Skills for Claude Code and Codex: AI SEO, AEO and GEO, code review with an A-F ship grade, CI gates, AI evals, design systems, conversion copy, Instagram growth, iOS and Android app shipping, creator rights, and…

    githubmeasured growthOpen source ↗

    129GitHub stars-4 (-3.0 %)
  25. 25
    Active

    A Model Context Protocol (MCP) server for Langfuse, enabling AI agents to query Langfuse trace data for enhanced debugging and observability

    githubmeasured growthOpen source ↗

    105GitHub starsstable