LLM observability: tools for developers

Monitoring, traces, evaluation and quality tracking for LLM applications. An entry can appear under several use cases when its topics or description provide several explicit signals.

Use case

Activity

Sort by

118 entries in this view.

Tool ranking

Ranked by measured growth, normalized across sources. Raw values keep their source-specific window, shown under each tool; a change requires at least two snapshots from the last 7 days.

  1. 1

    Skill
    Active

    Open-source infrastructure that turns scattered SKILL.md files into curated, retrieval-ready agent-skill corpora—with retrieval and evaluation tooling included.

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/EverMind-AI/SkillCorpus ~/.claude/skills/SkillCorpus

    360GitHub stars+200 (+125.0 %)
  2. 2

    Skill
    Active

    Research-backed, eval-driven skills for AI agents

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/ZSeven-W/craft-skills ~/.claude/skills/craft-skills

    138GitHub stars+56 (+68.3 %)
  3. 3
    Active

    Multi-tier framework for evaluating AI agent skills with quality gates, semantic overlap detection, synthetic evaluation dataset generation, and live agent evaluation that measures how skills affect agent behavior.

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/NVIDIA/SkillEvaluator ~/.claude/skills/SkillEvaluator

    363GitHub stars+70 (+23.9 %)
  4. 4

    Skill
    Active

    Deletion-first Agent Skill for removing test bloat, verification theater, and speculative fallbacks while preserving behavior.

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/MrZoyo/deslop-GPT ~/.claude/skills/deslop-GPT

    38GitHub stars+10 (+35.7 %)
  5. 5
    Active

    YAO = Yielding AI Outcomes. A rigorous engineering, evaluation, governance, and portability system for reusable agent skills.

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/yaojingang/yao-meta-skill ~/.claude/skills/yao-meta-skill

    2 565GitHub stars+69 (+2.8 %)
  6. 6

    Other
    Active

    Test your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare performance of GPT, Claude, Gemini, DeepSeek, and more. Simple declarative configs with command line and CI/CD integration. Used by OpenAI…

    githubmeasured growthOpen source ↗

    24 683GitHub stars+139 (+0.57 %)
  7. 7

    Other
    Active

    Mastra is the modern TypeScript framework for AI-powered applications and agents.

    githubmeasured growthOpen source ↗

    27 580GitHub stars+139 (+0.51 %)
  8. 8

    Other
    Active

    Open source observability platform for logs, metrics, traces, frontend monitoring, pipelines and LLM observability. A sophisticated, simple and highly performant alternative to Datadog, Splunk, and Elasticsearch with 140x lower storage…

    githubmeasured growthOpen source ↗

    21 575GitHub stars+128 (+0.60 %)
  9. 9

    Agent
    Active

    The Operating System for Scalable Enterprise AI Agents - Run, orchestrate, and deploy Compliant Enterprise AI Agents at scale across frameworks, without lock-in, rewrites or fragile glue code. Native support for MCP, A2A. Interface with all…

    githubmeasured growthOpen source ↗

    146GitHub stars+16 (+12.3 %)
  10. 10

    MCP
    Active

    A typed, replayable, budget-aware programming language for LLM agents. Compiles to Python & TypeScript.

    githubmeasured growthOpen source ↗

    2GitHub stars+1 (+100.0 %)
  11. 11

    Skill
    Active

    Production patterns for agent orchestrators, harnesses, evals, MCP/A2A interoperability, memory, provider routing, and LLM usage metering.

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/ssheleg/agent-stack ~/.claude/skills/agent-stack

    2GitHub stars+1 (+100.0 %)
  12. 12

    Agent
    Active

    A local A2A event and continuity service for agent applications, with structured history, provenance-safe transcripts, literal search, and natural-language queries.

    githubmeasured growthOpen source ↗

    1GitHub stars+1
  13. 13
    Active

    Audit which Claude Code skills you actually use — surface dead installs and hallucinated invocations from your session logs.

    githubmeasured growthOpen source ↗

    Install /plugin marketplace add sfrangulov/skill-graveyard

    10GitHub stars+3 (+42.9 %)
  14. 14

    Agent
    Active

    The open source AI engineering platform for agents, LLMs, and ML models. MLflow enables teams of all sizes to debug, evaluate, monitor, and optimize production-quality AI applications while controlling costs and managing access to models…

    githubmeasured growthOpen source ↗

    27 741GitHub stars+78 (+0.28 %)
  15. 15

    Other
    Active

    the LLM vulnerability scanner

    githubmeasured growthOpen source ↗

    9 079GitHub stars+56 (+0.62 %)
  16. 16

    Other
    Active

    The fastest path to AI-powered full stack observability, even for lean teams.

    githubmeasured growthOpen source ↗

    80 361GitHub stars+77 (+0.10 %)
  17. 17

    Library
    Active

    Prefect is a workflow orchestration framework for building resilient data pipelines in Python.

    githubmeasured growthOpen source ↗

    23 727GitHub stars+56 (+0.24 %)
  18. 18
    Active

    Build tested agent skills and govern their lifecycle through a user-defined marketplace: evidence, discovery, updates, rollback, quarantine, and 17-platform distribution.

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/FrancyJGLisboa/agent-skill-creator ~/.claude/skills/agent-skill-creator

    2 358GitHub stars+29 (+1.2 %)
  19. 19
    Active

    A Go framework for building production agent systems with graph workflows, tools, memory, A2A, AG-UI, MCP, evaluation, and observability.

    githubmeasured growthOpen source ↗

    1 752GitHub stars+26 (+1.5 %)
  20. 20

    Other
    Active

    A high-performance observability data pipeline.

    githubmeasured growthOpen source ↗

    22 489GitHub stars+41 (+0.18 %)
  21. 21

    Other
    Active

    SigNoz is an open-source, OpenTelemetry-native observability platform for your team and their AI agents. Get logs, metrics, and traces in one tool with features like APM, distributed tracing, log management, infra monitoring, etc. Combined…

    githubmeasured growthOpen source ↗

    31 973GitHub stars+42 (+0.13 %)
  22. 22

    Skill
    Active

    The Fable Workflow: how Claude Fable 5 worked, distilled into skills any model can run, with the eval that keeps it honest. Think / act / prove.

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/Sahir619/fable-method ~/.claude/skills/fable-method

    2 269GitHub stars+24 (+1.1 %)
  23. 23
    Active

    A test runner for agentskills.io-style AI agent skills

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/darkrishabh/agent-skills-eval ~/.claude/skills/agent-skills-eval

    713GitHub stars+14 (+2.0 %)
  24. 24

    MCP
    Dormant

    Open-source Python framework to deploy AI agents via HTTP, A2A, and MCP with built-in observability

    githubmeasured growthOpen source ↗

    5GitHub stars+1 (+25.0 %)

Learning resources

Ranked by measured growth, normalized across sources. These resources remain available separately and do not take part in the main tool ranking.

  1. 1
    ActiveOpen source ↗

    Documentation-discovery telemetry for Claude Code — heat/cold maps, health grade, evidence-backed router fixes. 100% local, zero tokens. /tt

    github

    Install /plugin marketplace add Hedde/trigger_tree

    Resourcemeasured growth
    14GitHub stars+2 (+16.7 %)