LLM observability: tools for developers

Monitoring, traces, evaluation and quality tracking for LLM applications. An entry can appear under several use cases when its topics or description provide several explicit signals.

Use case

Activity

Sort by

122 entries in this view.

Tool ranking

Ranked by measured growth, normalized across sources. Raw values keep their source-specific window, shown under each tool; a change requires at least two snapshots from the last 7 days.

  1. 1

    Other
    Dormant

    Official Python SDK for GT8004 — AI agent observability with MCP, A2A, x402 payment tracking. FastAPI, Flask, FastMCP middleware included.

    githubmeasured growthOpen source ↗

    1GitHub starsstable
  2. 2

    Agent
    Dormant

    Policy-enforced observability and fail-closed guardrails for MCP/A2A multi-agent systems.

    githubmeasured growthOpen source ↗

    26GitHub starsstable
  3. 3
    Active

    Claude Code skills where every entry ships receipts — accuracy-gated benchmarks against baseline and placebo, rejects published

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/sjh9714/skill-receipts ~/.claude/skills/skill-receipts

    2GitHub starsstable
  4. 4
    Active

    Sentry instrumentation skill for system-behavior tracking

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/tortastudios/sentry-instrumentation ~/.claude/skills/sentry-instrumentation

    24GitHub starsstable
  5. 5

    Other
    Dormant

    Multi-agent customer support system with Google ADK & Gemini 2.5 Flash Lite. Kaggle capstone demonstrating 11+ concepts. Automates 80%+ queries, <10s response time.

    githubmeasured growthOpen source ↗

    2GitHub starsstable
  6. 6

    MCP
    Active

    A typed, replayable, budget-aware programming language for LLM agents. Compiles to Python & TypeScript.

    githubmeasured growthOpen source ↗

    2GitHub starsstable
  7. 7

    Other
    Active

    AI-detector auditing toolkit: measures how often a detector flags genuine human writing, how stable its verdict is across seeds, and whether that verdict survives meaning-preserving edits. Detector-in-the-loop measurement harness. Claude…

    githubmeasured growthOpen source ↗

    18GitHub starsstable
  8. 8
    Dormant

    OpenTelemetry semantic conventions and instrumentation for agent provenance, derivation lineage, and acceptance criteria evaluation. Fills the Microsoft AI stack observability gap.

    githubmeasured growthOpen source ↗

    0GitHub starsstable
  9. 9

    Agent
    Dormant

    A Multi-Agent System (MAS) evaluation framework using PydanticAI that generates and evaluates scientific paper reviews through a three-tiered assessment approach: traditional metrics, LLM-as-a-Judge, and graph-based complexity analysis.

    githubmeasured growthOpen source ↗

    2GitHub starsstable
  10. 10
    Active

    Multi-agent quality gate skill for Claude Code that researches, reviews, tests and challenges AI-generated work before the final answer.

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/ma-nucho-pro/supervisor-skill-claude ~/.claude/skills/supervisor-skill-claude

    2GitHub starsstable
  11. 11

    Skill
    Active

    Measure prompt and skill improvements with blind A/B comparison.

    githubmeasured growthOpen source ↗

    Install /plugin marketplace add shinpr/rashomon

    18GitHub starsstable
  12. 12
    Dormant

    MCP as a Judge: a behavioral MCP that strengthens AI coding assistants via explicit LLM evaluations

    mcpmeasured growthOpen source ↗

    Install claude mcp add mcp-as-a-judge -- uvx mcp-as-a-judge

    17GitHub starsstable
  13. 13

    Agent
    Active

    Multi-agent SRE on-call investigator that auto-triages Slack/Discord infrastructure alerts via AWS Bedrock AgentCore, fanning out to specialized agents (CloudWatch, EKS, Slack/Discord scanners) for parallel investigation.

    githubmeasured growthOpen source ↗

    3GitHub starsstable
  14. 14

    Skill
    Active

    Agent skills for Arize — datasets, experiments, and traces via the ax CLI

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/Arize-ai/arize-skills ~/.claude/skills/arize-skills

    47GitHub starsstable
  15. 15

    Skill
    Active

    A skill creator that proves its skills work. Evidence-driven skill creation for Claude Code and Codex: baseline-tested generation, per-skill regression evals, ecosystem doctor, cross-runtime compile, and an opt-in proactive advisor.

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/tripleyak/SkillForge ~/.claude/skills/SkillForge

    884GitHub stars-1 (-0.11 %)
  16. 16
    Active

    Self-hosted AI SRE for Kubernetes — zero-instrumentation eBPF observability plus a copilot that fixes issues through guardrailed, self-verifying actions. BYO-LLM, air-gapped capable.

    githubmeasured growthOpen source ↗

    131GitHub stars-4 (-3.0 %)
  17. 17
    Active

    74 open-source Agent Skills for Claude Code and Codex: AI SEO, AEO and GEO, code review with an A-F ship grade, CI gates, AI evals, design systems, conversion copy, Instagram growth, iOS and Android app shipping, creator rights, and…

    githubmeasured growthOpen source ↗

    129GitHub stars-4 (-3.0 %)
  18. 18
    Active

    Open benchmark for Claude Code SEO skills — real headless execution against fixture sites with planted-defect answer keys. Deterministic scoring, pre-registered rubric.

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/aleclindz/seo-skill-bench ~/.claude/skills/seo-skill-bench

    51GitHub stars-5 (-8.9 %)

Learning resources

Ranked by measured growth, normalized across sources. These resources remain available separately and do not take part in the main tool ranking.

  1. 1
    ActiveOpen source ↗

    Curated + scored map of official MCP servers and agents for DevOps, Cloud, SRE, and Platform Engineering — every entry rated on production access, approval gates, and audit evidence.

    githubResourcemeasured growth
    77GitHub starsstable
  2. 2
    ActiveOpen source ↗

    Handbook técnico aberto sobre engenharia de IA em produção: ML tradicional, LLMs, RAG, agentes, segurança, observabilidade, FinOps e deployment.

    githubResourcemeasured growth
    1GitHub starsstable
  3. 3
    ActiveOpen source ↗

    My complete journey to becoming an Agentic AI Engineer through structured learning, projects, experiments, and production-ready implementations of modern AI systems.

    githubResourcemeasured growth
    18GitHub starsstable
  4. 4

    tunelab

    Skill
    ActiveOpen source ↗

    Claude Code plugin for LLM fine-tuning, distillation, and evaluation — decide whether you need fine-tuning at all, distill your LLM logs into small local models (MLX/LoRA), evaluate with held-out discipline, and learn the why at every step.

    github

    Install /plugin marketplace add rchaz/tunelab

    Resourcemeasured growth
    6GitHub starsstable