LLM observability: tools for developers

Monitoring, traces, evaluation and quality tracking for LLM applications. An entry can appear under several use cases when its topics or description provide several explicit signals.

Use case

Activity

Sort by

117 entries in this view.

Tool ranking

Ranked by measured growth, normalized across sources. Raw values keep their source-specific window, shown under each tool; a change requires at least two snapshots from the last 7 days.

  1. 1
    Active

    Production reliability, benchmarking, and evaluation layer for MCP agent runtimes — protocol negotiation, A2A interop, circuit breakers, connection pooling, adaptive streaming, and autoscaling, with measured throughput/latency and a…

    githubmeasured growthOpen source ↗

    3GitHub starsstable
  2. 2

    Other
    Active

    A command-line tool to scaffold and manage enterprise-ready AI Agents powered by the A2A (Agent-to-Agent) protocol

    githubmeasured growthOpen source ↗

    14GitHub stars+1 (+7.7 %)
  3. 3

    Agent
    Dormant

    Real-time debugging proxy for Agent2Agent (A2A) multi-agent systems

    githubmeasured growthOpen source ↗

    3GitHub starsstable
  4. 4

    Skill
    Active

    Axiom is a curated marketplace of shared plugins for Claude Code and Codex.

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/netopsengineer/axiom ~/.claude/skills/axiom

    5GitHub starsstable
  5. 5
    Active

    Claude Code plugin marketplace — agentic-engineering (spec-driven shape→decide→execute→measure→eval with adversarial review) + github-keeper (audit/elevate READMEs and make a repo open-source-ready).

    githubmeasured growthOpen source ↗

    Install /plugin marketplace add GiustoPiedimonte/agentic-engineering-marketplace

    13GitHub starsstable
  6. 6

    MCP
    Dormant

    Open-source Python framework to deploy AI agents via HTTP, A2A, and MCP with built-in observability

    githubmeasured growthOpen source ↗

    5GitHub starsstable
  7. 7

    Other
    Dormant

    A Claude Code skill that adds a rubric-based eval layer to any agent project. Framework-agnostic — generates rubric, test cases, judge prompt, and harness. Returns a weighted score plus a judge-leniency signal.

    githubmeasured growthOpen source ↗

    13GitHub starsstable
  8. 8
    Active

    Amazon Bedrock AgentCore enterprise platform accelerator with AWS CDK, Terraform organization guardrails, MCP/A2A agents, security, memory, and observability.

    githubmeasured growthOpen source ↗

    6GitHub stars+2 (+50.0 %)
  9. 9
    Active

    Query Spanlens LLM observability from Cursor, Claude Desktop, or Continue via MCP.

    mcpmeasured growthOpen source ↗

    Install claude mcp add mcp-server -- npx @spanlens/mcp-server

    12GitHub starsstable
  10. 10

    Agent
    Active

    Provide clear documentation for AgentStack’s MCP protocol, plugins, and ecosystem API with usage examples and tool references.

    githubmeasured growthOpen source ↗

    0GitHub starsstable
  11. 11

    Agent
    Active

    MemroOS / memroos: memory OS and governance layer for AI agents, agent workflows, dispatch, proof, and context continuity.

    githubmeasured growthOpen source ↗

    7GitHub starsstable
  12. 12

    Skill
    Active

    Production-grade AI coding rules for Cursor and Claude Code. 15 rules + 9 doc templates + skills + agents + MCP setup. Drop into any project.

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/aiagentwithdhruv/ai-dev-stack ~/.claude/skills/ai-dev-stack

    10GitHub stars+1 (+11.1 %)
  13. 13
    Dormant

    A self-improving harness router for Claude Code.

    githubmeasured growthOpen source ↗

    Install /plugin marketplace add SeongwoongCho/adaptive-harness

    8GitHub starsstable
  14. 14
    Active

    Audit which Claude Code skills you actually use — surface dead installs and hallucinated invocations from your session logs.

    githubmeasured growthOpen source ↗

    Install /plugin marketplace add sfrangulov/skill-graveyard

    10GitHub stars+1 (+11.1 %)
  15. 15

    Skill
    Active

    Anchor — the production-grade AGENTS.md template for AI coding agents. 51 sections of battle-tested rules. Keep your agents grounded.

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/peva3/anchor ~/.claude/skills/anchor

    8GitHub starsstable
  16. 16

    Skill
    Active

    Turn one decision into a judged tournament of solutions, then pick the best — a Claude Code skill that generates candidates, auto-derives the rubric, judges independently, and returns a defensible winner.

    githubmeasured growthOpen source ↗

    Install /plugin marketplace add CoriChui/bakeoff

    10GitHub starsstable
  17. 17

    Other
    Active

    A Go-native framework for LLM agents, with OpenTelemetry observability built in.

    githubmeasured growthOpen source ↗

    10GitHub starsstable
  18. 18
    Active

    Architecture-first Python scaffold for an auditable multi-agent corporate credit desk using A2A, MCP, deterministic credit policies, model routing, and OpenTelemetry.

    githubmeasured growthOpen source ↗

    0GitHub starsstable
  19. 19

    Agent
    Dormant

    Contract-based testing for LLM agents: hybrid evaluation, multi-agent + A2A, deterministic replay.

    githubmeasured growthOpen source ↗

    0GitHub starsstable
  20. 20

    Skill
    Active

    Marketing Mix Model expert agent plugin for Claude Code - pymc-marketing v0.18.2+

    githubmeasured growthOpen source ↗

    Install /plugin marketplace add Yakoub-ai/agent-mmm

    4GitHub starsstable
  21. 21
    Dormant

    A Python proof-of-concept for tracing multi-turn Agent-to-Agent (A2A) conversations as a single unified MLflow trace for LLM observability and evaluation.

    githubmeasured growthOpen source ↗

    0GitHub starsstable
  22. 22

    Agent
    Active

    A2A green-agent orchestrator for evaluating agents on the AppWorld benchmark, built on the AgentBeats SDK

    githubmeasured growthOpen source ↗

    0GitHub starsstable
  23. 23

    Agent
    Dormant

    AgentOps: Multi-agent infrastructure remediation platform. A2A protocol, agent coordination, HITL approval, auto-rollback.

    githubmeasured growthOpen source ↗

    0GitHub starsstable
  24. 24

    Other
    Active

    🔑 HomeStream · 家园·流 — 零成本自托管多Agent协作框架,通往AI世界的那把钥匙 | Zero-cost self-hosted multi-agent framework — The key to AI world

    githubmeasured growthOpen source ↗

    0GitHub starsstable
  25. 25
    Active

    Governed local-LLM observability: model policy, prompt scanner, capture proxy, 21 tools.

    mcpmeasured growthOpen source ↗

    Install claude mcp add ai-guardian -- uvx ai-guardian-aiops

    0GitHub starsstable