LLM observability: tools for developers

Monitoring, traces, evaluation and quality tracking for LLM applications. An entry can appear under several use cases when its topics or description provide several explicit signals.

Use case

Activity

Sort by

75 entries in this view.

Tool ranking

Ranked by measured growth, normalized across sources. Raw values keep their source-specific window, shown under each tool; a change requires at least two snapshots from the last 7 days.

  1. 1

    Skill
    Active

    Open-source infrastructure that turns scattered SKILL.md files into curated, retrieval-ready agent-skill corpora—with retrieval and evaluation tooling included.

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/EverMind-AI/SkillCorpus ~/.claude/skills/SkillCorpus

    424GitHub stars+170 (+66.9 %)
  2. 2

    Other
    Active

    Evidence-first infrastructure for reproducible, isolated executable evaluation and online rewards.

    githubmeasured growthOpen source ↗

    140GitHub stars+24 (+20.7 %)
  3. 3
    Active

    Multi-tier framework for evaluating AI agent skills with quality gates, semantic overlap detection, synthetic evaluation dataset generation, and live agent evaluation that measures how skills affect agent behavior.

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/NVIDIA/SkillEvaluator ~/.claude/skills/SkillEvaluator

    389GitHub stars+57 (+17.2 %)
  4. 4

    Skill
    Active

    Research-backed, eval-driven skills for AI agents

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/ZSeven-W/craft-skills ~/.claude/skills/craft-skills

    142GitHub stars+18 (+14.5 %)
  5. 5

    Agent
    Active

    The Operating System for Scalable Enterprise AI Agents - Run, orchestrate, and deploy Compliant Enterprise AI Agents at scale across frameworks, without lock-in, rewrites or fragile glue code. Native support for MCP, A2A. Interface with all…

    githubmeasured growthOpen source ↗

    156GitHub stars+19 (+13.9 %)
  6. 6
    Active

    Agent Skill evaluation harness for paired variants, trace artifacts, and runner adapters

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/adewale/skill-eval-harness ~/.claude/skills/skill-eval-harness

    73GitHub stars+4 (+5.8 %)
  7. 7

    Agent
    Active

    DataBuff is an AI-native APM built on Opentelemetry,with multi-agent troubleshooting out of the box.

    githubmeasured growthOpen source ↗

    642GitHub stars+32 (+5.2 %)
  8. 8
    Active

    Vendor-neutral OpenTelemetry skills for AI coding agents, grounded in upstream sources

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/ollygarden/opentelemetry-agent-skills ~/.claude/skills/opentelemetry-agent-skills

    97GitHub stars+4 (+4.3 %)
  9. 9
    Active

    Skills, prompts, and instructions for building AI agents on top of Dynatrace production context

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/Dynatrace/dynatrace-for-ai ~/.claude/skills/dynatrace-for-ai

    132GitHub stars+4 (+3.1 %)
  10. 10
    Active

    A test runner for agentskills.io-style AI agent skills

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/darkrishabh/agent-skills-eval ~/.claude/skills/agent-skills-eval

    719GitHub stars+13 (+1.8 %)
  11. 11
    Active

    YAO = Yielding AI Outcomes. A rigorous engineering, evaluation, governance, and portability system for reusable agent skills.

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/yaojingang/yao-meta-skill ~/.claude/skills/yao-meta-skill

    2 577GitHub stars+35 (+1.4 %)
  12. 12

    Skill
    Active

    local-first analytics for AI agent skills

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/crafter-station/skill-kit ~/.claude/skills/skill-kit

    77GitHub stars+1 (+1.3 %)
  13. 13
    Active

    🪢 Langfuse documentation -- Langfuse is the open source LLM Engineering Platform. Observability, evals, prompt management, playground and metrics to debug and improve LLM apps

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/langfuse/langfuse-docs ~/.claude/skills/langfuse-docs

    239GitHub stars+3 (+1.3 %)
  14. 14
    Active

    Build tested agent skills and govern their lifecycle through a user-defined marketplace: evidence, discovery, updates, rollback, quarantine, and 17-platform distribution.

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/FrancyJGLisboa/agent-skill-creator ~/.claude/skills/agent-skill-creator

    2 367GitHub stars+28 (+1.2 %)
  15. 15
    Active

    A Go framework for building production agent systems with graph workflows, tools, memory, A2A, AG-UI, MCP, evaluation, and observability.

    githubmeasured growthOpen source ↗

    1 759GitHub stars+18 (+1.0 %)
  16. 16
    Active

    Self improving agents through iterations

    githubmeasured growthOpen source ↗

    105GitHub stars+1 (+0.96 %)
  17. 17
    Active

    🐙 ADLC Team Skills — Agentic SDLC for Engineering Teams

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/tikalk/adlc-team-skills ~/.claude/skills/adlc-team-skills

    132GitHub stars+1 (+0.76 %)
  18. 18
    Active

    Dashboard for monitoring claude code sessions.

    githubmeasured growthOpen source ↗

    Install /plugin marketplace add JayantDevkar/claude-code-karma

    323GitHub stars+2 (+0.62 %)
  19. 19

    Other
    Active

    Test your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare performance of GPT, Claude, Gemini, DeepSeek, and more. Simple declarative configs with command line and CI/CD integration. Used by OpenAI…

    githubmeasured growthOpen source ↗

    24 768GitHub stars+142 (+0.58 %)
  20. 20

    Other
    Active

    Mastra is the modern TypeScript framework for AI-powered applications and agents.

    githubmeasured growthOpen source ↗

    27 658GitHub stars+128 (+0.46 %)
  21. 21

    Other
    Active

    Open source observability platform for logs, metrics, traces, frontend monitoring, pipelines and LLM observability. A sophisticated, simple and highly performant alternative to Datadog, Splunk, and Elasticsearch with 140x lower storage…

    githubmeasured growthOpen source ↗

    21 615GitHub stars+96 (+0.45 %)
  22. 22
    Active

    🫖 Status page with uptime monitoring & API monitoring as code 🫖

    githubmeasured growthOpen source ↗

    9 056GitHub stars+31 (+0.34 %)
  23. 23

    Other
    Active

    the LLM vulnerability scanner

    githubmeasured growthOpen source ↗

    9 079GitHub stars+31 (+0.34 %)
  24. 24

    Agent
    Active

    The open source AI engineering platform for agents, LLMs, and ML models. MLflow enables teams of all sizes to debug, evaluate, monitor, and optimize production-quality AI applications while controlling costs and managing access to models…

    githubmeasured growthOpen source ↗

    27 783GitHub stars+82 (+0.30 %)
  25. 25

    Library
    Active

    Prefect is a workflow orchestration framework for building resilient data pipelines in Python.

    githubmeasured growthOpen source ↗

    23 766GitHub stars+66 (+0.28 %)