LLM observability: tools for developers

Monitoring, traces, evaluation and quality tracking for LLM applications. An entry can appear under several use cases when its topics or description provide several explicit signals.

Use case

Activity

Sort by

122 entries in this view.

Tool ranking

Ranked by measured growth, normalized across sources. Raw values keep their source-specific window, shown under each tool; a change requires at least two snapshots from the last 7 days.

  1. 1

    Agent
    Active

    Provide clear documentation for AgentStack’s MCP protocol, plugins, and ecosystem API with usage examples and tool references.

    githubmeasured growthOpen source ↗

    0GitHub starsstable
  2. 2
    Active

    AgentX python SDK. Build multi-agent AI workforce. Run evaluation. Trace your agent. Full Observability.

    githubmeasured growthOpen source ↗

    68GitHub starsstable
  3. 3
    Active

    Production reliability, benchmarking, and evaluation layer for MCP agent runtimes — protocol negotiation, A2A interop, circuit breakers, connection pooling, adaptive streaming, and autoscaling, with measured throughput/latency and a…

    githubmeasured growthOpen source ↗

    3GitHub starsstable
  4. 4
    Active

    Make Claude Opus 4.8 behave like Claude Fable 5 — doctrine output style, drift-catching hooks, and an eval loop against golden Fable transcripts. Claude Code plugin.

    githubmeasured growthOpen source ↗

    Install /plugin marketplace add rennf93/opus-fable-playbook

    34GitHub stars+1 (+3.0 %)
  5. 5
    Active

    Mide si tu agente de IA cumple las reglas que le escribiste. Lee el historial local de Claude Code y devuelve un porcentaje por regla. Sin instalar nada, sin red, biblioteca estándar.

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/jleonceo/adherencia-reglas ~/.claude/skills/adherencia-reglas

    1GitHub starsstable
  6. 6

    Skill
    Active

    Real-time execution trace and cost intelligence for Claude Code

    githubmeasured growthOpen source ↗

    Install /plugin marketplace add DeibyGS/claudestat

    34GitHub starsstable
  7. 7
    Active

    Open benchmark for Claude Code SEO skills — real headless execution against fixture sites with planted-defect answer keys. Deterministic scoring, pre-registered rubric.

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/aleclindz/seo-skill-bench ~/.claude/skills/seo-skill-bench

    51GitHub stars-5 (-8.9 %)
  8. 8
    Active

    Claude Code skills where every entry ships receipts — accuracy-gated benchmarks against baseline and placebo, rejects published

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/sjh9714/skill-receipts ~/.claude/skills/skill-receipts

    2GitHub starsstable
  9. 9

    Agent
    Active

    A2A green-agent orchestrator for evaluating agents on the AppWorld benchmark, built on the AgentBeats SDK

    githubmeasured growthOpen source ↗

    0GitHub starsstable
  10. 10

    Skill
    Active

    A skill creator that proves its skills work. Evidence-driven skill creation for Claude Code and Codex: baseline-tested generation, per-skill regression evals, ecosystem doctor, cross-runtime compile, and an opt-in proactive advisor.

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/tripleyak/SkillForge ~/.claude/skills/SkillForge

    885GitHub starsstable
  11. 11

    Skill
    Active

    Turn one decision into a judged tournament of solutions, then pick the best — a Claude Code skill that generates candidates, auto-derives the rubric, judges independently, and returns a defensible winner.

    githubmeasured growthOpen source ↗

    Install /plugin marketplace add CoriChui/bakeoff

    10GitHub starsstable
  12. 12

    MCP
    Active

    Open-source, self-hosted control plane for OpenClaw and Hermes AI-agent fleets on Docker/Kubernetes — REST, CLI, and MCP.

    githubmeasured growthOpen source ↗

    50GitHub stars+1 (+2.0 %)
  13. 13
    Active

    YAO = Yielding AI Outcomes. A rigorous engineering, evaluation, governance, and portability system for reusable agent skills.

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/yaojingang/yao-meta-skill ~/.claude/skills/yao-meta-skill

    2 596GitHub stars+37 (+1.4 %)
  14. 14
    Active

    Skill instalable para Claude Code que mide si tu agente cumple las reglas establecidas. Lee el historial de sesiones de tu disco y devuelve un porcentaje por regla de cumplimiento. Sin instalar nada, sin red, biblioteca estándar.

    githubmeasured growthOpen source ↗

    Install /plugin marketplace add jleonceo/skill-adherencia-reglas

    1GitHub starsstable
  15. 15

    Skill
    Dormant

    An assembly line for AI software development. 35 skills, 11 agent personas, 29 commands. From raw idea to shipped code.

    githubmeasured growthOpen source ↗

    Install /plugin marketplace add aneja5/forge-skills

    3GitHub starsstable
  16. 16
    Active

    Audit which Claude Code skills you actually use — surface dead installs and hallucinated invocations from your session logs.

    githubmeasured growthOpen source ↗

    Install /plugin marketplace add sfrangulov/skill-graveyard

    10GitHub starsstable
  17. 17
    Dormant

    A Python proof-of-concept for tracing multi-turn Agent-to-Agent (A2A) conversations as a single unified MLflow trace for LLM observability and evaluation.

    githubmeasured growthOpen source ↗

    0GitHub starsstable
  18. 18

    Skill
    Active

    Production-grade AI coding rules for Cursor and Claude Code. 15 rules + 9 doc templates + skills + agents + MCP setup. Drop into any project.

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/aiagentwithdhruv/ai-dev-stack ~/.claude/skills/ai-dev-stack

    10GitHub starsstable
  19. 19
    Active

    Agent 降智检测与自愈公评网络 — an immune system for the AI agent society

    githubmeasured growthOpen source ↗

    1GitHub starsstable
  20. 20
    Active

    Dashboard for monitoring claude code sessions.

    githubmeasured growthOpen source ↗

    Install /plugin marketplace add JayantDevkar/claude-code-karma

    324GitHub stars+3 (+0.93 %)
  21. 21

    Other
    Active

    A Go-native framework for LLM agents, with OpenTelemetry observability built in.

    githubmeasured growthOpen source ↗

    10GitHub starsstable
  22. 22

    Agent
    Active

    The Operating System for Scalable Enterprise AI Agents - Run, orchestrate, and deploy Compliant Enterprise AI Agents at scale across frameworks, without lock-in, rewrites or fragile glue code. Native support for MCP, A2A. Interface with all…

    githubmeasured growthOpen source ↗

    166GitHub stars+25 (+17.7 %)
  23. 23

    MCP
    Dormant

    Open-source Python framework to deploy AI agents via HTTP, A2A, and MCP with built-in observability

    githubmeasured growthOpen source ↗

    5GitHub starsstable
  24. 24
    Active

    Deterministic crash tests for agent control planes: paired worlds, hard stops, budgets, scopes, approvals, and hash-linked receipts.

    githubmeasured growthOpen source ↗

    1GitHub starsstable
  25. 25

    Agent
    Active

    Generic DeepSeek Harness (dsh) plugins: A2A protocol server, session storage mirror, and Langfuse observability.

    githubmeasured growthOpen source ↗

    1GitHub starsstable