LLM observability: tools for developers

Monitoring, traces, evaluation and quality tracking for LLM applications. An entry can appear under several use cases when its topics or description provide several explicit signals.

Use case

Activity

Sort by

117 entries in this view.

Tool ranking

Ranked by measured growth, normalized across sources. Raw values keep their source-specific window, shown under each tool; a change requires at least two snapshots from the last 7 days.

  1. 1

    Skill
    Dormant

    An assembly line for AI software development. 35 skills, 11 agent personas, 29 commands. From raw idea to shipped code.

    githubmeasured growthOpen source ↗

    Install /plugin marketplace add aneja5/forge-skills

    3GitHub starsstable
  2. 2
    Active

    Audit which Claude Code skills you actually use — surface dead installs and hallucinated invocations from your session logs.

    githubmeasured growthOpen source ↗

    Install /plugin marketplace add sfrangulov/skill-graveyard

    10GitHub starsstable
  3. 3
    Dormant

    A Python proof-of-concept for tracing multi-turn Agent-to-Agent (A2A) conversations as a single unified MLflow trace for LLM observability and evaluation.

    githubmeasured growthOpen source ↗

    0GitHub starsstable
  4. 4

    Skill
    Active

    Production-grade AI coding rules for Cursor and Claude Code. 15 rules + 9 doc templates + skills + agents + MCP setup. Drop into any project.

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/aiagentwithdhruv/ai-dev-stack ~/.claude/skills/ai-dev-stack

    10GitHub starsstable
  5. 5
    Active

    Agent 降智检测与自愈公评网络 — an immune system for the AI agent society

    githubmeasured growthOpen source ↗

    1GitHub starsstable
  6. 6

    Other
    Active

    A Go-native framework for LLM agents, with OpenTelemetry observability built in.

    githubmeasured growthOpen source ↗

    10GitHub starsstable
  7. 7

    MCP
    Dormant

    Open-source Python framework to deploy AI agents via HTTP, A2A, and MCP with built-in observability

    githubmeasured growthOpen source ↗

    5GitHub starsstable
  8. 8
    Active

    Deterministic crash tests for agent control planes: paired worlds, hard stops, budgets, scopes, approvals, and hash-linked receipts.

    githubmeasured growthOpen source ↗

    1GitHub starsstable
  9. 9

    Agent
    Active

    Generic DeepSeek Harness (dsh) plugins: A2A protocol server, session storage mirror, and Langfuse observability.

    githubmeasured growthOpen source ↗

    1GitHub starsstable
  10. 10
    Active

    Multi-agent quality gate skill for Claude Code that researches, reviews, tests and challenges AI-generated work before the final answer.

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/ma-nucho-pro/supervisor-skill-claude ~/.claude/skills/supervisor-skill-claude

    2GitHub starsstable
  11. 11

    Agent
    Active

    A local A2A event and continuity service for agent applications, with structured history, provenance-safe transcripts, literal search, and natural-language queries.

    githubmeasured growthOpen source ↗

    1GitHub starsstable
  12. 12
    Active

    Local-first Agentic AI Infrastructure Platform with LLM Gateway, Agent DAG Runtime, MCP Tool Hub, A2A Agent Mesh, Local RAG, Tool Sandbox and Observability.

    githubmeasured growthOpen source ↗

    1GitHub starsstable
  13. 13

    Agent
    Active

    Multi-agent SRE on-call investigator that auto-triages Slack/Discord infrastructure alerts via AWS Bedrock AgentCore, fanning out to specialized agents (CloudWatch, EKS, Slack/Discord scanners) for parallel investigation.

    githubmeasured growthOpen source ↗

    3GitHub starsstable
  14. 14
    Active

    Query Spanlens LLM observability from Cursor, Claude Desktop, or Continue via MCP.

    mcpmeasured growthOpen source ↗

    Install claude mcp add mcp-server -- npx @spanlens/mcp-server

    12GitHub starsstable
  15. 15
    Active

    Governed local-LLM observability: model policy, prompt scanner, capture proxy, 21 tools.

    mcpmeasured growthOpen source ↗

    Install claude mcp add ai-guardian -- uvx ai-guardian-aiops

    0GitHub starsstable
  16. 16

    Agent
    Active

    A lightweight Python library that decouples agentic runtime from applications it builds

    githubmeasured growthOpen source ↗

    1GitHub starsstable
  17. 17
    Active

    Amazon Bedrock AgentCore enterprise platform accelerator with AWS CDK, Terraform organization guardrails, MCP/A2A agents, security, memory, and observability.

    githubmeasured growthOpen source ↗

    6GitHub stars+1 (+20.0 %)
  18. 18
    Active

    Claude Code plugin marketplace — agentic-engineering (spec-driven shape→decide→execute→measure→eval with adversarial review) + github-keeper (audit/elevate READMEs and make a repo open-source-ready).

    githubmeasured growthOpen source ↗

    Install /plugin marketplace add GiustoPiedimonte/agentic-engineering-marketplace

    13GitHub starsstable
  19. 19

    Skill
    Active

    Measure prompt and skill improvements with blind A/B comparison.

    githubmeasured growthOpen source ↗

    Install /plugin marketplace add shinpr/rashomon

    18GitHub starsstable
  20. 20
    Active

    Terminal-first validation harness for reviewing PRDs, RFCs, strategy docs, and experiment plans.

    githubmeasured growthOpen source ↗

    1GitHub starsstable
  21. 21

    Agent
    Dormant

    The reliability and interoperability layer for AI agents. A2A-native. MCP-bridged.

    githubmeasured growthOpen source ↗

    1GitHub starsstable
  22. 22

    MCP
    Active

    A typed, replayable, budget-aware programming language for LLM agents. Compiles to Python & TypeScript.

    githubmeasured growthOpen source ↗

    2GitHub starsstable
  23. 23

    Other
    Dormant

    A Claude Code skill that adds a rubric-based eval layer to any agent project. Framework-agnostic — generates rubric, test cases, judge prompt, and harness. Returns a weighted score plus a judge-leniency signal.

    githubmeasured growthOpen source ↗

    13GitHub starsstable
  24. 24

    Agent
    Active

    Local-first control plane for durable, governed agent teams: DAG orchestration, approvals, memory, signed skills, trace contracts, and realtime.

    githubmeasured growthOpen source ↗

    0GitHub starsstable
  25. 25

    Agent
    Active

    MemroOS / memroos: memory OS and governance layer for AI agents, agent workflows, dispatch, proof, and context continuity.

    githubmeasured growthOpen source ↗

    7GitHub starsstable