LLM observability: tools for developers

Monitoring, traces, evaluation and quality tracking for LLM applications. An entry can appear under several use cases when its topics or description provide several explicit signals.

Use case

Activity

Sort by

123 entries in this view.

Tool ranking

Ranked by creation date, newest first; undated entries come last. Raw values keep their source-specific window, shown under each tool; a change requires at least two snapshots from the last 7 days.

  1. 1

    Skill
    Active

    Anchor — the production-grade AGENTS.md template for AI coding agents. 51 sections of battle-tested rules. Keep your agents grounded.

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/peva3/anchor ~/.claude/skills/anchor

    8GitHub starsstable
  2. 2

    Other
    Active

    Open-source observability tool that uses AI agents to self-heal your software

    githubmeasured growthOpen source ↗

    1 404GitHub stars+1 (+0.07 %)
  3. 3

    Skill
    Active

    Glanceable Claude Code and Codex state in your terminal tabs: white=idle, blue=working, orange=waiting. Multi-terminal (iTerm2, WezTerm, AI Power Term), a live status bar with Anthropic usage limits, /sfl and /nil window save-and-restore,…

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/wasulajr/headsup ~/.claude/skills/headsup

    1GitHub starsstable
  4. 4
    Active

    74 open-source Agent Skills for Claude Code and Codex: AI SEO, AEO and GEO, code review with an A-F ship grade, CI gates, AI evals, design systems, conversion copy, Instagram growth, iOS and Android app shipping, creator rights, and…

    githubmeasured growthOpen source ↗

    129GitHub stars-4 (-3.0 %)
  5. 5

    Other
    Active

    A Go-native framework for LLM agents, with OpenTelemetry observability built in.

    githubmeasured growthOpen source ↗

    10GitHub starsstable
  6. 6

    Agent
    Dormant

    The reliability and interoperability layer for AI agents. A2A-native. MCP-bridged.

    githubmeasured growthOpen source ↗

    1GitHub starsstable
  7. 7

    Agent
    Active

    Multi-agent SRE on-call investigator that auto-triages Slack/Discord infrastructure alerts via AWS Bedrock AgentCore, fanning out to specialized agents (CloudWatch, EKS, Slack/Discord scanners) for parallel investigation.

    githubmeasured growthOpen source ↗

    3GitHub starsstable
  8. 8

    Skill
    Dormant

    Don't make LLMs honest. Make every factual claim auditable. — An LLM Claim Auditing Layer with T1-T7 truth gradients. 98.1% business effectiveness on LiarBench v0.2.

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/lc198707/anti-lie ~/.claude/skills/anti-lie

    89GitHub starsstable
  9. 9
    Active

    Two Claude Code skills for adversarial code review (single-model PAR and multi-model MMAR with cross-critique to catch hallucinations), plus a fixture-based eval suite.

    githubmeasured growthOpen source ↗

    Install /plugin marketplace add prime-radiant-inc/parallel-adversarial-review

    17GitHub starsstable
  10. 10
    Active

    A test runner for agentskills.io-style AI agent skills

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/darkrishabh/agent-skills-eval ~/.claude/skills/agent-skills-eval

    723GitHub stars+15 (+2.1 %)
  11. 11
    Active

    8 MCP SMB products — standalone AI servers for SMBs: Guardrails, FinOps, Observability, Router, Trust Score, Memory, ThinkSecure, A2A Lite. 37 tools, TNC credits billing.

    githubmeasured growthOpen source ↗

    1GitHub starsstable
  12. 12
    Active

    Audit which Claude Code skills you actually use — surface dead installs and hallucinated invocations from your session logs.

    githubmeasured growthOpen source ↗

    Install /plugin marketplace add sfrangulov/skill-graveyard

    10GitHub stars+1 (+11.1 %)
  13. 13

    Skill
    Active

    Real-time execution trace and cost intelligence for Claude Code

    githubmeasured growthOpen source ↗

    Install /plugin marketplace add DeibyGS/claudestat

    34GitHub starsstable
  14. 14

    Other
    Dormant

    A Claude Code skill that adds a rubric-based eval layer to any agent project. Framework-agnostic — generates rubric, test cases, judge prompt, and harness. Returns a weighted score plus a judge-leniency signal.

    githubmeasured growthOpen source ↗

    13GitHub starsstable
  15. 15
    Active

    Query Spanlens LLM observability from Cursor, Claude Desktop, or Continue via MCP.

    mcpmeasured growthOpen source ↗

    Install claude mcp add mcp-server -- npx @spanlens/mcp-server

    12GitHub starsstable
  16. 16

    Agent
    Dormant

    Contract-based testing for LLM agents: hybrid evaluation, multi-agent + A2A, deterministic replay.

    githubmeasured growthOpen source ↗

    0GitHub starsstable
  17. 17
    Active

    Sentry instrumentation skill for system-behavior tracking

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/tortastudios/sentry-instrumentation ~/.claude/skills/sentry-instrumentation

    24GitHub starsstable
  18. 18

    Agent
    Active

    MemroOS / memroos: memory OS and governance layer for AI agents, agent workflows, dispatch, proof, and context continuity.

    githubmeasured growthOpen source ↗

    7GitHub starsstable
  19. 19

    Agent
    Dormant

    Real-time debugging proxy for Agent2Agent (A2A) multi-agent systems

    githubmeasured growthOpen source ↗

    3GitHub starsstable
  20. 20
    Active

    YAO = Yielding AI Outcomes. A rigorous engineering, evaluation, governance, and portability system for reusable agent skills.

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/yaojingang/yao-meta-skill ~/.claude/skills/yao-meta-skill

    2 589GitHub stars+38 (+1.5 %)
  21. 21
    Active

    Skills, prompts, and instructions for building AI agents on top of Dynatrace production context

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/Dynatrace/dynatrace-for-ai ~/.claude/skills/dynatrace-for-ai

    135GitHub stars+4 (+3.1 %)
  22. 22
    Active

    AgentGateway — independent third-party profile of a public API surface, by API Evangelist. AgentGateway is an open-source, AI-native proxy and gateway for routing, observing, and governing traffic to and from AI agents, LLM providers, and…

    githubmeasured growthOpen source ↗

    0GitHub starsstable
  23. 23

    Skill
    Active

    Marketing Mix Model expert agent plugin for Claude Code - pymc-marketing v0.18.2+

    githubmeasured growthOpen source ↗

    Install /plugin marketplace add Yakoub-ai/agent-mmm

    4GitHub starsstable
  24. 24

    Agent
    Active

    Provide clear documentation for AgentStack’s MCP protocol, plugins, and ecosystem API with usage examples and tool references.

    githubmeasured growthOpen source ↗

    0GitHub starsstable

Learning resources

Ranked by creation date, newest first; undated entries come last. These resources remain available separately and do not take part in the main tool ranking.

  1. 1
    ActiveOpen source ↗

    Handbook técnico aberto sobre engenharia de IA em produção: ML tradicional, LLMs, RAG, agentes, segurança, observabilidade, FinOps e deployment.

    githubResourcemeasured growth
    1GitHub starsstable