LLM observability: tools for developers

Monitoring, traces, evaluation and quality tracking for LLM applications. An entry can appear under several use cases when its topics or description provide several explicit signals.

Use case

Activity

Sort by

75 entries in this view.

Tool ranking

Ranked by creation date, newest first; undated entries come last. Raw values keep their source-specific window, shown under each tool; a change requires at least two snapshots from the last 7 days.

  1. 1

    Agent
    Active

    面向长程研发任务的 Java Agent Harness,支持 A2A 跨语言协作、可恢复执行、上下文工程与 Eval 驱动开发。

    githubmeasured growthOpen source ↗

    0GitHub starsstable
  2. 2

    Skill
    Active

    Deletion-first Agent Skill for removing test bloat, verification theater, and speculative fallbacks while preserving behavior.

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/MrZoyo/deslop-GPT ~/.claude/skills/deslop-GPT

    55GitHub stars+27 (+96.4 %)
  3. 3

    Agent
    Active

    Generic DeepSeek Harness (dsh) plugins: A2A protocol server, session storage mirror, and Langfuse observability.

    githubmeasured growthOpen source ↗

    1GitHub starsstable
  4. 4
    Active

    Production reliability, benchmarking, and evaluation layer for MCP agent runtimes — protocol negotiation, A2A interop, circuit breakers, connection pooling, adaptive streaming, and autoscaling, with measured throughput/latency and a…

    githubmeasured growthOpen source ↗

    3GitHub starsstable
  5. 5
    Active

    Multi-agent quality gate skill for Claude Code that researches, reviews, tests and challenges AI-generated work before the final answer.

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/ma-nucho-pro/supervisor-skill-claude ~/.claude/skills/supervisor-skill-claude

    2GitHub starsstable
  6. 6
    Active

    Mini Program Engineering Skill Suite is an Agent Skill suite for evidence‑first mini‑program development. It helps agents bring WeChat or other mini‑program projects from vague intent to reliable engineering work: project intake, product…

    githubestimated momentumOpen source ↗

    Install git clone https://github.com/NocodeMrLi/mini-program-engineering-skill-suite ~/.claude/skills/mini-program-engineering-skill-suite

    39GitHub stars
  7. 7
    Active

    Deterministic, local-only audit reports for Claude Code AI agent sessions. Rust + SQLite. Zero network calls. MCP server for agents.

    githubmeasured growthOpen source ↗

    22GitHub stars+2 (+10.0 %)
  8. 8
    Active

    Amazon Bedrock AgentCore enterprise platform accelerator with AWS CDK, Terraform organization guardrails, MCP/A2A agents, security, memory, and observability.

    githubmeasured growthOpen source ↗

    5GitHub stars+1 (+25.0 %)
  9. 9

    Skill
    Active

    Open-source infrastructure that turns scattered SKILL.md files into curated, retrieval-ready agent-skill corpora—with retrieval and evaluation tooling included.

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/EverMind-AI/SkillCorpus ~/.claude/skills/SkillCorpus

    424GitHub stars+170 (+66.9 %)
  10. 10

    Agent
    Active

    Local-first control plane for durable, governed agent teams: DAG orchestration, approvals, memory, signed skills, trace contracts, and realtime.

    githubmeasured growthOpen source ↗

    0GitHub starsstable
  11. 11

    Skill
    Active

    Production patterns for agent orchestrators, harnesses, evals, MCP/A2A interoperability, memory, provider routing, and LLM usage metering.

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/ssheleg/agent-stack ~/.claude/skills/agent-stack

    2GitHub stars+1 (+100.0 %)
  12. 12

    Skill
    Active

    Research-backed, eval-driven skills for AI agents

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/ZSeven-W/craft-skills ~/.claude/skills/craft-skills

    142GitHub stars+18 (+14.5 %)
  13. 13
    Active

    Make Claude write clearly, for everyone.

    githubmeasured growthOpen source ↗

    Install /plugin marketplace add stefanobaghino/simple-output-styles

    16GitHub starsstable
  14. 14

    Other
    Active

    Evidence-first infrastructure for reproducible, isolated executable evaluation and online rewards.

    githubmeasured growthOpen source ↗

    140GitHub stars+24 (+20.7 %)
  15. 15
    Active

    Deterministic crash tests for agent control planes: paired worlds, hard stops, budgets, scopes, approvals, and hash-linked receipts.

    githubmeasured growthOpen source ↗

    1GitHub starsstable
  16. 16
    Active

    🐙 ADLC Team Skills — Agentic SDLC for Engineering Teams

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/tikalk/adlc-team-skills ~/.claude/skills/adlc-team-skills

    132GitHub stars+1 (+0.76 %)
  17. 17

    MCP
    Active

    A typed, replayable, budget-aware programming language for LLM agents. Compiles to Python & TypeScript.

    githubmeasured growthOpen source ↗

    2GitHub starsstable
  18. 18
    Active

    Vendor-neutral OpenTelemetry tracing for A2A agents and MCP services, with W3C context propagation and privacy-safe telemetry.

    githubmeasured growthOpen source ↗

    1GitHub starsstable
  19. 19
    Active

    Architecture-first Python scaffold for an auditable multi-agent corporate credit desk using A2A, MCP, deterministic credit policies, model routing, and OpenTelemetry.

    githubmeasured growthOpen source ↗

    0GitHub starsstable
  20. 20
    Active

    Governed local-LLM observability: model policy, prompt scanner, capture proxy, 21 tools.

    mcpmeasured growthOpen source ↗

    Install claude mcp add ai-guardian -- uvx ai-guardian-aiops

    0GitHub starsstable
  21. 21

    Other
    Active

    AI SRE AgenticOps for Kubernetes and cloud infrastructure.

    githubmeasured growthOpen source ↗

    782GitHub stars+1 (+0.13 %)
  22. 22
    Active

    Open benchmark for Claude Code SEO skills — real headless execution against fixture sites with planted-defect answer keys. Deterministic scoring, pre-registered rubric.

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/aleclindz/seo-skill-bench ~/.claude/skills/seo-skill-bench

    51GitHub stars-5 (-8.9 %)

Learning resources

Ranked by creation date, newest first; undated entries come last. These resources remain available separately and do not take part in the main tool ranking.

  1. 1
    ActiveOpen source ↗

    My complete journey to becoming an Agentic AI Engineer through structured learning, projects, experiments, and production-ready implementations of modern AI systems.

    githubResourcemeasured growth
    18GitHub starsstable
  2. 2
    ActiveOpen source ↗

    Documentation-discovery telemetry for Claude Code — heat/cold maps, health grade, evidence-backed router fixes. 100% local, zero tokens. /tt

    github

    Install /plugin marketplace add Hedde/trigger_tree

    Resourcemeasured growth
    14GitHub stars+1 (+7.7 %)
  3. 3
    ActiveOpen source ↗

    50+ curated LLM observability tools PLUS 26 Agent Skills (several with runnable, unit-tested scripts) to build, evaluate, debug, secure & monitor reliable LLM apps. Tracing, evals, guardrails, LLMOps.

    github

    Install git clone https://github.com/ContextJet-ai/awesome-llm-observability ~/.claude/skills/awesome-llm-observability

    Resourcemeasured growth
    32GitHub stars+2 (+6.7 %)