LLM observability: tools for developers

Monitoring, traces, evaluation and quality tracking for LLM applications. An entry can appear under several use cases when its topics or description provide several explicit signals.

Use case

Activity

Sort by

122 entries in this view.

Tool ranking

Ranked by measured growth, normalized across sources. Raw values keep their source-specific window, shown under each tool; a change requires at least two snapshots from the last 7 days.

  1. 1
    Dormant

    Two Claude Code skills for adversarial code review (single-model PAR and multi-model MMAR with cross-critique to catch hallucinations), plus a fixture-based eval suite.

    githubmeasured growthOpen source ↗

    Install /plugin marketplace add prime-radiant-inc/parallel-adversarial-review

    17GitHub starsstable
  2. 2

    Skill
    Dormant

    An assembly line for AI software development. 35 skills, 11 agent personas, 29 commands. From raw idea to shipped code.

    githubmeasured growthOpen source ↗

    Install /plugin marketplace add aneja5/forge-skills

    3GitHub starsstable
  3. 3
    Active

    Make Claude write clearly, for everyone.

    githubmeasured growthOpen source ↗

    Install /plugin marketplace add stefanobaghino/simple-output-styles

    16GitHub starsstable
  4. 4

    Agent
    Active

    面向长程研发任务的 Java Agent Harness,支持 A2A 跨语言协作、可恢复执行、上下文工程与 Eval 驱动开发。

    githubmeasured growthOpen source ↗

    0GitHub starsstable
  5. 5
    Active

    Production reliability, benchmarking, and evaluation layer for MCP agent runtimes — protocol negotiation, A2A interop, circuit breakers, connection pooling, adaptive streaming, and autoscaling, with measured throughput/latency and a…

    githubmeasured growthOpen source ↗

    3GitHub starsstable
  6. 6

    Agent
    Dormant

    Real-time debugging proxy for Agent2Agent (A2A) multi-agent systems

    githubmeasured growthOpen source ↗

    3GitHub starsstable
  7. 7

    Skill
    Active

    Axiom is a curated marketplace of shared plugins for Claude Code and Codex.

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/netopsengineer/axiom ~/.claude/skills/axiom

    5GitHub starsstable
  8. 8
    Active

    Claude Code plugin marketplace — agentic-engineering (spec-driven shape→decide→execute→measure→eval with adversarial review) + github-keeper (audit/elevate READMEs and make a repo open-source-ready).

    githubmeasured growthOpen source ↗

    Install /plugin marketplace add GiustoPiedimonte/agentic-engineering-marketplace

    13GitHub starsstable
  9. 9

    MCP
    Dormant

    Open-source Python framework to deploy AI agents via HTTP, A2A, and MCP with built-in observability

    githubmeasured growthOpen source ↗

    5GitHub starsstable
  10. 10

    Other
    Dormant

    A Claude Code skill that adds a rubric-based eval layer to any agent project. Framework-agnostic — generates rubric, test cases, judge prompt, and harness. Returns a weighted score plus a judge-leniency signal.

    githubmeasured growthOpen source ↗

    13GitHub starsstable
  11. 11
    Active

    Query Spanlens LLM observability from Cursor, Claude Desktop, or Continue via MCP.

    mcpmeasured growthOpen source ↗

    Install claude mcp add mcp-server -- npx @spanlens/mcp-server

    12GitHub starsstable
  12. 12

    Agent
    Active

    Provide clear documentation for AgentStack’s MCP protocol, plugins, and ecosystem API with usage examples and tool references.

    githubmeasured growthOpen source ↗

    0GitHub starsstable
  13. 13

    Agent
    Active

    MemroOS / memroos: memory OS and governance layer for AI agents, agent workflows, dispatch, proof, and context continuity.

    githubmeasured growthOpen source ↗

    7GitHub starsstable
  14. 14
    Dormant

    A self-improving harness router for Claude Code.

    githubmeasured growthOpen source ↗

    Install /plugin marketplace add SeongwoongCho/adaptive-harness

    8GitHub starsstable
  15. 15

    Skill
    Active

    Anchor — the production-grade AGENTS.md template for AI coding agents. 51 sections of battle-tested rules. Keep your agents grounded.

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/peva3/anchor ~/.claude/skills/anchor

    8GitHub starsstable
  16. 16

    Skill
    Active

    Turn one decision into a judged tournament of solutions, then pick the best — a Claude Code skill that generates candidates, auto-derives the rubric, judges independently, and returns a defensible winner.

    githubmeasured growthOpen source ↗

    Install /plugin marketplace add CoriChui/bakeoff

    10GitHub starsstable
  17. 17
    Dormant

    Local open-source dev tool to debug, secure, and evaluate LLM agents. Provides static analysis, dynamic security checks, and runtime monitoring - integrates with Cursor and Claude Code.

    githubmeasured growthOpen source ↗

    Install /plugin marketplace add cylestio/agent-inspector

    9GitHub starsstable
  18. 18

    Other
    Active

    A Go-native framework for LLM agents, with OpenTelemetry observability built in.

    githubmeasured growthOpen source ↗

    10GitHub starsstable
  19. 19
    Active

    Architecture-first Python scaffold for an auditable multi-agent corporate credit desk using A2A, MCP, deterministic credit policies, model routing, and OpenTelemetry.

    githubmeasured growthOpen source ↗

    0GitHub starsstable
  20. 20

    Agent
    Dormant

    Contract-based testing for LLM agents: hybrid evaluation, multi-agent + A2A, deterministic replay.

    githubmeasured growthOpen source ↗

    0GitHub starsstable
  21. 21

    Skill
    Active

    Marketing Mix Model expert agent plugin for Claude Code - pymc-marketing v0.18.2+

    githubmeasured growthOpen source ↗

    Install /plugin marketplace add Yakoub-ai/agent-mmm

    4GitHub starsstable
  22. 22
    Dormant

    A Python proof-of-concept for tracing multi-turn Agent-to-Agent (A2A) conversations as a single unified MLflow trace for LLM observability and evaluation.

    githubmeasured growthOpen source ↗

    0GitHub starsstable
  23. 23

    Agent
    Active

    A2A green-agent orchestrator for evaluating agents on the AppWorld benchmark, built on the AgentBeats SDK

    githubmeasured growthOpen source ↗

    0GitHub starsstable
  24. 24
    Active

    🟪 Open-source runtime that ships any LangGraph or Google ADK agent as a production-ready FastAPI service. Bundled , AG-UI copilotkit API, chat UI, 15+ guardrails, MCP, OpenTelemetry, OIDC. One pip install. Self-hosted, no vendor lock-in.

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/Idun-Group/idun-agent-platform ~/.claude/skills/idun-agent-platform

    199GitHub starsstable
  25. 25

    Agent
    Dormant

    AgentOps: Multi-agent infrastructure remediation platform. A2A protocol, agent coordination, HITL approval, auto-rollback.

    githubmeasured growthOpen source ↗

    0GitHub starsstable