LLM observability: tools for developers

Monitoring, traces, evaluation and quality tracking for LLM applications. An entry can appear under several use cases when its topics or description provide several explicit signals.

Use case

Activity

Sort by

123 entries in this view.

Tool ranking

Ranked by known GitHub stars, highest first; archived repositories come after maintained repositories. Raw values keep their source-specific window, shown under each tool; a change requires at least two snapshots from the last 7 days.

  1. 1

    Other
    Active

    The fastest path to AI-powered full stack observability, even for lean teams.

    githubmeasured growthOpen source ↗

    80 402GitHub stars+91 (+0.11 %)
  2. 2

    Other
    Active

    SigNoz is an open-source, OpenTelemetry-native observability platform for your team and their AI agents. Get logs, metrics, and traces in one tool with features like APM, distributed tracing, log management, infra monitoring, etc. Combined…

    githubmeasured growthOpen source ↗

    31 992GitHub stars+62 (+0.19 %)
  3. 3

    Agent
    Active

    The open source AI engineering platform for agents, LLMs, and ML models. MLflow enables teams of all sizes to debug, evaluate, monitor, and optimize production-quality AI applications while controlling costs and managing access to models…

    githubmeasured growthOpen source ↗

    27 768GitHub stars+80 (+0.29 %)
  4. 4

    Other
    Active

    Mastra is the modern TypeScript framework for AI-powered applications and agents.

    githubmeasured growthOpen source ↗

    27 624GitHub stars+121 (+0.44 %)
  5. 5

    Other
    Active

    eBPF-based Networking, Security, and Observability

    githubmeasured growthOpen source ↗

    25 041GitHub stars+27 (+0.11 %)
  6. 6

    Other
    Active

    Test your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare performance of GPT, Claude, Gemini, DeepSeek, and more. Simple declarative configs with command line and CI/CD integration. Used by OpenAI…

    githubmeasured growthOpen source ↗

    24 738GitHub stars+131 (+0.53 %)
  7. 7

    Library
    Active

    Prefect is a workflow orchestration framework for building resilient data pipelines in Python.

    githubmeasured growthOpen source ↗

    23 759GitHub stars+66 (+0.28 %)
  8. 8

    Other
    Active

    A high-performance observability data pipeline.

    githubmeasured growthOpen source ↗

    22 507GitHub stars+45 (+0.20 %)
  9. 9

    Other
    Active

    Open source observability platform for logs, metrics, traces, frontend monitoring, pipelines and LLM observability. A sophisticated, simple and highly performant alternative to Datadog, Splunk, and Elasticsearch with 140x lower storage…

    githubmeasured growthOpen source ↗

    21 608GitHub stars+115 (+0.54 %)
  10. 10

    Skill
    Active

    The container platform tailored for Kubernetes multi-cloud, datacenter, and edge management ⎈ 🖥 ☁️

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/kubesphere/kubesphere ~/.claude/skills/kubesphere

    17 037GitHub stars+9 (+0.05 %)
  11. 11

    Agent
    Active

    eBPF-powered network observability for Kubernetes. Indexes L4/L7 traffic with full K8s context, decrypts TLS without keys. Queryable by AI agents via MCP and humans via dashboard.

    githubmeasured growthOpen source ↗

    12 066GitHub stars+7 (+0.06 %)
  12. 12

    Other
    Active

    the LLM vulnerability scanner

    githubmeasured growthOpen source ↗

    9 079GitHub stars+43 (+0.48 %)
  13. 13
    Active

    🫖 Status page with uptime monitoring & API monitoring as code 🫖

    githubmeasured growthOpen source ↗

    9 052GitHub stars+28 (+0.31 %)
  14. 14
    Active

    YAO = Yielding AI Outcomes. A rigorous engineering, evaluation, governance, and portability system for reusable agent skills.

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/yaojingang/yao-meta-skill ~/.claude/skills/yao-meta-skill

    2 577GitHub stars+54 (+2.1 %)
  15. 15
    Active

    Build tested agent skills and govern their lifecycle through a user-defined marketplace: evidence, discovery, updates, rollback, quarantine, and 17-platform distribution.

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/FrancyJGLisboa/agent-skill-creator ~/.claude/skills/agent-skill-creator

    2 367GitHub stars+30 (+1.3 %)
  16. 16

    Skill
    Active

    The Fable Workflow: how Claude Fable 5 worked, distilled into skills any model can run, with the eval that keeps it honest. Think / act / prove.

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/Sahir619/fable-method ~/.claude/skills/fable-method

    2 272GitHub stars+19 (+0.84 %)
  17. 17
    Active

    A Go framework for building production agent systems with graph workflows, tools, memory, A2A, AG-UI, MCP, evaluation, and observability.

    githubmeasured growthOpen source ↗

    1 759GitHub stars+24 (+1.4 %)
  18. 18

    Other
    Active

    Open-source observability tool that uses AI agents to self-heal your software

    githubmeasured growthOpen source ↗

    1 404GitHub stars+2 (+0.14 %)
  19. 19

    Skill
    Active

    A skill creator that proves its skills work. Evidence-driven skill creation for Claude Code and Codex: baseline-tested generation, per-skill regression evals, ecosystem doctor, cross-runtime compile, and an opt-in proactive advisor.

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/tripleyak/SkillForge ~/.claude/skills/SkillForge

    884GitHub stars+3 (+0.34 %)
  20. 20

    Skill
    Active

    OpenJudge: A Unified Framework for Holistic Evaluation and Quality Rewards

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/agentscope-ai/OpenJudge ~/.claude/skills/OpenJudge

    809GitHub stars+10 (+1.3 %)
  21. 21

    Other
    Active

    AI SRE AgenticOps for Kubernetes and cloud infrastructure.

    githubmeasured growthOpen source ↗

    782GitHub stars+1 (+0.13 %)
  22. 22
    Active

    A test runner for agentskills.io-style AI agent skills

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/darkrishabh/agent-skills-eval ~/.claude/skills/agent-skills-eval

    719GitHub stars+15 (+2.1 %)
  23. 23

    Agent
    Active

    DataBuff is an AI-native APM built on Opentelemetry,with multi-agent troubleshooting out of the box.

    githubmeasured growthOpen source ↗

    634GitHub stars+26 (+4.3 %)
  24. 24

    Skill
    Active

    Open-source infrastructure that turns scattered SKILL.md files into curated, retrieval-ready agent-skill corpora—with retrieval and evaluation tooling included.

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/EverMind-AI/SkillCorpus ~/.claude/skills/SkillCorpus

    424GitHub stars+204 (+92.7 %)
  25. 25
    Active

    Multi-tier framework for evaluating AI agent skills with quality gates, semantic overlap detection, synthetic evaluation dataset generation, and live agent evaluation that measures how skills affect agent behavior.

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/NVIDIA/SkillEvaluator ~/.claude/skills/SkillEvaluator

    389GitHub stars+63 (+19.3 %)