LLM observability: tools for developers

Monitoring, traces, evaluation and quality tracking for LLM applications. An entry can appear under several use cases when its topics or description provide several explicit signals.

Use case

Activity

Sort by

117 entries in this view.

Tool ranking

Ranked by measured growth, normalized across sources. Raw values keep their source-specific window, shown under each tool; a change requires at least two snapshots from the last 7 days.

  1. 1

    Skill
    Active

    Anchor — the production-grade AGENTS.md template for AI coding agents. 51 sections of battle-tested rules. Keep your agents grounded.

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/peva3/anchor ~/.claude/skills/anchor

    8GitHub starsstable
  2. 2
    Dormant

    A self-improving harness router for Claude Code.

    githubmeasured growthOpen source ↗

    Install /plugin marketplace add SeongwoongCho/adaptive-harness

    8GitHub starsstable
  3. 3

    Agent
    Active

    MemroOS / memroos: memory OS and governance layer for AI agents, agent workflows, dispatch, proof, and context continuity.

    githubmeasured growthOpen source ↗

    7GitHub starsstable
  4. 4

    MCP
    Dormant

    Open-source Python framework to deploy AI agents via HTTP, A2A, and MCP with built-in observability

    githubmeasured growthOpen source ↗

    5GitHub starsstable
  5. 5

    Skill
    Active

    Axiom is a curated marketplace of shared plugins for Claude Code and Codex.

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/netopsengineer/axiom ~/.claude/skills/axiom

    5GitHub starsstable
  6. 6
    Active

    Production reliability, benchmarking, and evaluation layer for MCP agent runtimes — protocol negotiation, A2A interop, circuit breakers, connection pooling, adaptive streaming, and autoscaling, with measured throughput/latency and a…

    githubmeasured growthOpen source ↗

    3GitHub starsstable
  7. 7

    Skill
    Dormant

    An assembly line for AI software development. 35 skills, 11 agent personas, 29 commands. From raw idea to shipped code.

    githubmeasured growthOpen source ↗

    Install /plugin marketplace add aneja5/forge-skills

    3GitHub starsstable
  8. 8

    Agent
    Active

    Multi-agent SRE on-call investigator that auto-triages Slack/Discord infrastructure alerts via AWS Bedrock AgentCore, fanning out to specialized agents (CloudWatch, EKS, Slack/Discord scanners) for parallel investigation.

    githubmeasured growthOpen source ↗

    3GitHub starsstable
  9. 9

    Agent
    Dormant

    Real-time debugging proxy for Agent2Agent (A2A) multi-agent systems

    githubmeasured growthOpen source ↗

    3GitHub starsstable
  10. 10

    MCP
    Active

    A typed, replayable, budget-aware programming language for LLM agents. Compiles to Python & TypeScript.

    githubmeasured growthOpen source ↗

    2GitHub starsstable
  11. 11
    Active

    Claude Code skills where every entry ships receipts — accuracy-gated benchmarks against baseline and placebo, rejects published

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/sjh9714/skill-receipts ~/.claude/skills/skill-receipts

    2GitHub starsstable
  12. 12
    Active

    Multi-agent quality gate skill for Claude Code that researches, reviews, tests and challenges AI-generated work before the final answer.

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/ma-nucho-pro/supervisor-skill-claude ~/.claude/skills/supervisor-skill-claude

    2GitHub starsstable
  13. 13

    Other
    Dormant

    Multi-agent customer support system with Google ADK & Gemini 2.5 Flash Lite. Kaggle capstone demonstrating 11+ concepts. Automates 80%+ queries, <10s response time.

    githubmeasured growthOpen source ↗

    2GitHub starsstable
  14. 14

    Agent
    Dormant

    A Multi-Agent System (MAS) evaluation framework using PydanticAI that generates and evaluates scientific paper reviews through a three-tiered assessment approach: traditional metrics, LLM-as-a-Judge, and graph-based complexity analysis.

    githubmeasured growthOpen source ↗

    2GitHub starsstable
  15. 15

    Agent
    Active

    Generic DeepSeek Harness (dsh) plugins: A2A protocol server, session storage mirror, and Langfuse observability.

    githubmeasured growthOpen source ↗

    1GitHub starsstable
  16. 16

    Agent
    Active

    A lightweight Python library that decouples agentic runtime from applications it builds

    githubmeasured growthOpen source ↗

    1GitHub starsstable
  17. 17
    Active

    Deterministic crash tests for agent control planes: paired worlds, hard stops, budgets, scopes, approvals, and hash-linked receipts.

    githubmeasured growthOpen source ↗

    1GitHub starsstable
  18. 18
    Active

    Agent 降智检测与自愈公评网络 — an immune system for the AI agent society

    githubmeasured growthOpen source ↗

    1GitHub starsstable
  19. 19
    Active

    Local-first Agentic AI Infrastructure Platform with LLM Gateway, Agent DAG Runtime, MCP Tool Hub, A2A Agent Mesh, Local RAG, Tool Sandbox and Observability.

    githubmeasured growthOpen source ↗

    1GitHub starsstable
  20. 20

    Agent
    Dormant

    The reliability and interoperability layer for AI agents. A2A-native. MCP-bridged.

    githubmeasured growthOpen source ↗

    1GitHub starsstable
  21. 21
    Active

    Skill instalable para Claude Code que mide si tu agente cumple las reglas establecidas. Lee el historial de sesiones de tu disco y devuelve un porcentaje por regla de cumplimiento. Sin instalar nada, sin red, biblioteca estándar.

    githubmeasured growthOpen source ↗

    Install /plugin marketplace add jleonceo/skill-adherencia-reglas

    1GitHub starsstable
  22. 22
    Active

    Mide si tu agente de IA cumple las reglas que le escribiste. Lee el historial local de Claude Code y devuelve un porcentaje por regla. Sin instalar nada, sin red, biblioteca estándar.

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/jleonceo/adherencia-reglas ~/.claude/skills/adherencia-reglas

    1GitHub starsstable
  23. 23
    Active

    8 MCP SMB products — standalone AI servers for SMBs: Guardrails, FinOps, Observability, Router, Trust Score, Memory, ThinkSecure, A2A Lite. 37 tools, TNC credits billing.

    githubmeasured growthOpen source ↗

    1GitHub starsstable
  24. 24

    Skill
    Active

    Glanceable Claude Code and Codex state in your terminal tabs: white=idle, blue=working, orange=waiting. Multi-terminal (iTerm2, WezTerm, AI Power Term), a live status bar with Anthropic usage limits, /sfl and /nil window save-and-restore,…

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/wasulajr/headsup ~/.claude/skills/headsup

    1GitHub starsstable

Learning resources

Ranked by measured growth, normalized across sources. These resources remain available separately and do not take part in the main tool ranking.

  1. 1

    tunelab

    Skill
    ActiveOpen source ↗

    Claude Code plugin for LLM fine-tuning, distillation, and evaluation — decide whether you need fine-tuning at all, distill your LLM logs into small local models (MLX/LoRA), evaluate with held-out discipline, and learn the why at every step.

    github

    Install /plugin marketplace add rchaz/tunelab

    Resourcemeasured growth
    6GitHub starsstable