LLM observability: tools for developers

Monitoring, traces, evaluation and quality tracking for LLM applications. An entry can appear under several use cases when its topics or description provide several explicit signals.

Use case

Activity

Sort by

123 entries in this view.

Tool ranking

Ranked by creation date, newest first; undated entries come last. Raw values keep their source-specific window, shown under each tool; a change requires at least two snapshots from the last 7 days.

  1. 1
    Active

    Architecture-first Python scaffold for an auditable multi-agent corporate credit desk using A2A, MCP, deterministic credit policies, model routing, and OpenTelemetry.

    githubmeasured growthOpen source ↗

    0GitHub starsstable
  2. 2
    Active

    Governed local-LLM observability: model policy, prompt scanner, capture proxy, 21 tools.

    mcpmeasured growthOpen source ↗

    Install claude mcp add ai-guardian -- uvx ai-guardian-aiops

    0GitHub starsstable
  3. 3

    Other
    Active

    AI SRE AgenticOps for Kubernetes and cloud infrastructure.

    githubmeasured growthOpen source ↗

    782GitHub stars+1 (+0.13 %)
  4. 4

    Other
    Active

    🔑 HomeStream · 家园·流 — 零成本自托管多Agent协作框架,通往AI世界的那把钥匙 | Zero-cost self-hosted multi-agent framework — The key to AI world

    githubmeasured growthOpen source ↗

    0GitHub starsstable
  5. 5

    Skill
    Active

    Turn one decision into a judged tournament of solutions, then pick the best — a Claude Code skill that generates candidates, auto-derives the rubric, judges independently, and returns a defensible winner.

    githubmeasured growthOpen source ↗

    Install /plugin marketplace add CoriChui/bakeoff

    10GitHub starsstable
  6. 6

    Skill
    Active

    The Fable Workflow: how Claude Fable 5 worked, distilled into skills any model can run, with the eval that keeps it honest. Think / act / prove.

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/Sahir619/fable-method ~/.claude/skills/fable-method

    2 273GitHub stars+11 (+0.49 %)
  7. 7
    Active

    Self-hosted AI SRE for Kubernetes — zero-instrumentation eBPF observability plus a copilot that fixes issues through guardrailed, self-verifying actions. BYO-LLM, air-gapped capable.

    githubmeasured growthOpen source ↗

    131GitHub stars-4 (-3.0 %)
  8. 8
    Active

    Open benchmark for Claude Code SEO skills — real headless execution against fixture sites with planted-defect answer keys. Deterministic scoring, pre-registered rubric.

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/aleclindz/seo-skill-bench ~/.claude/skills/seo-skill-bench

    51GitHub stars-5 (-8.9 %)
  9. 9
    Active

    Make Claude Opus 4.8 behave like Claude Fable 5 — doctrine output style, drift-catching hooks, and an eval loop against golden Fable transcripts. Claude Code plugin.

    githubmeasured growthOpen source ↗

    Install /plugin marketplace add rennf93/opus-fable-playbook

    34GitHub stars+1 (+3.0 %)
  10. 10
    Active

    Claude Code plugin marketplace — agentic-engineering (spec-driven shape→decide→execute→measure→eval with adversarial review) + github-keeper (audit/elevate READMEs and make a repo open-source-ready).

    githubmeasured growthOpen source ↗

    Install /plugin marketplace add GiustoPiedimonte/agentic-engineering-marketplace

    13GitHub starsstable
  11. 11

    Other
    Active

    AI-detector auditing toolkit: measures how often a detector flags genuine human writing, how stable its verdict is across seeds, and whether that verdict survives meaning-preserving edits. Detector-in-the-loop measurement harness. Claude…

    githubmeasured growthOpen source ↗

    18GitHub starsstable
  12. 12
    Active

    Multi-tier framework for evaluating AI agent skills with quality gates, semantic overlap detection, synthetic evaluation dataset generation, and live agent evaluation that measures how skills affect agent behavior.

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/NVIDIA/SkillEvaluator ~/.claude/skills/SkillEvaluator

    399GitHub stars+51 (+14.7 %)
  13. 13
    Active

    Agent 降智检测与自愈公评网络 — an immune system for the AI agent society

    githubmeasured growthOpen source ↗

    1GitHub starsstable
  14. 14

    Agent
    Active

    DataBuff is an AI-native APM built on Opentelemetry,with multi-agent troubleshooting out of the box.

    githubmeasured growthOpen source ↗

    649GitHub stars+39 (+6.4 %)
  15. 15

    Agent
    Active

    A local A2A event and continuity service for agent applications, with structured history, provenance-safe transcripts, literal search, and natural-language queries.

    githubmeasured growthOpen source ↗

    1GitHub starsstable
  16. 16
    Active

    Terminal-first validation harness for reviewing PRDs, RFCs, strategy docs, and experiment plans.

    githubmeasured growthOpen source ↗

    1GitHub starsstable
  17. 17
    Active

    Optimize any AI agent’s skills, tools/MCP, and prompts against your own evals.

    githubestimated momentumOpen source ↗

    47GitHub stars
  18. 18

    Agent
    Active

    A lightweight Python library that decouples agentic runtime from applications it builds

    githubmeasured growthOpen source ↗

    1GitHub starsstable
  19. 19
    Active

    Agent Skill evaluation harness for paired variants, trace artifacts, and runner adapters

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/adewale/skill-eval-harness ~/.claude/skills/skill-eval-harness

    73GitHub stars+4 (+5.8 %)
  20. 20

    Agent
    Active

    The agent framework where the model never holds the trigger — every consequential action clears your policy first, waits for a human when it matters, and lands on a record you can verify. Build on it, or put it around the agent you already…

    githubmeasured growthOpen source ↗

    1GitHub starsstable
  21. 21

    Skill
    Active

    Axiom is a curated marketplace of shared plugins for Claude Code and Codex.

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/netopsengineer/axiom ~/.claude/skills/axiom

    5GitHub starsstable
  22. 22
    Active

    Local-first Agentic AI Infrastructure Platform with LLM Gateway, Agent DAG Runtime, MCP Tool Hub, A2A Agent Mesh, Local RAG, Tool Sandbox and Observability.

    githubmeasured growthOpen source ↗

    1GitHub starsstable

Learning resources

Ranked by creation date, newest first; undated entries come last. These resources remain available separately and do not take part in the main tool ranking.

  1. 1
    ActiveOpen source ↗

    50+ curated LLM observability tools PLUS 26 Agent Skills (several with runnable, unit-tested scripts) to build, evaluate, debug, secure & monitor reliable LLM apps. Tracing, evals, guardrails, LLMOps.

    github

    Install git clone https://github.com/ContextJet-ai/awesome-llm-observability ~/.claude/skills/awesome-llm-observability

    Resourcemeasured growth
    33GitHub stars+3 (+10.0 %)
  2. 2
    ActiveOpen source ↗

    Curated + scored map of official MCP servers and agents for DevOps, Cloud, SRE, and Platform Engineering — every entry rated on production access, approval gates, and audit evidence.

    githubResourcemeasured growth
    77GitHub starsstable
  3. 3

    tunelab

    Skill
    ActiveOpen source ↗

    Claude Code plugin for LLM fine-tuning, distillation, and evaluation — decide whether you need fine-tuning at all, distill your LLM logs into small local models (MLX/LoRA), evaluate with held-out discipline, and learn the why at every step.

    github

    Install /plugin marketplace add rchaz/tunelab

    Resourcemeasured growth
    6GitHub starsstable