LLM observability: tools for developers

Monitoring, traces, evaluation and quality tracking for LLM applications. An entry can appear under several use cases when its topics or description provide several explicit signals.

Use case

Activity

Sort by

123 entries in this view.

Tool ranking

Ranked by creation date, newest first; undated entries come last. Raw values keep their source-specific window, shown under each tool; a change requires at least two snapshots from the last 7 days.

  1. 1

    Skill
    Dormant

    An assembly line for AI software development. 35 skills, 11 agent personas, 29 commands. From raw idea to shipped code.

    githubmeasured growthOpen source ↗

    Install /plugin marketplace add aneja5/forge-skills

    3GitHub starsstable
  2. 2

    MCP
    Active

    Open-source, self-hosted control plane for OpenClaw and Hermes AI-agent fleets on Docker/Kubernetes — REST, CLI, and MCP.

    githubmeasured growthOpen source ↗

    50GitHub stars+2 (+4.2 %)
  3. 3
    Dormant

    A self-improving harness router for Claude Code.

    githubmeasured growthOpen source ↗

    Install /plugin marketplace add SeongwoongCho/adaptive-harness

    8GitHub starsstable
  4. 4
    Dormant

    OpenTelemetry semantic conventions and instrumentation for agent provenance, derivation lineage, and acceptance criteria evaluation. Fills the Microsoft AI stack observability gap.

    githubmeasured growthOpen source ↗

    0GitHub starsstable
  5. 5

    Skill
    Active

    Production-grade AI coding rules for Cursor and Claude Code. 15 rules + 9 doc templates + skills + agents + MCP setup. Drop into any project.

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/aiagentwithdhruv/ai-dev-stack ~/.claude/skills/ai-dev-stack

    10GitHub stars+1 (+11.1 %)
  6. 6
    Active

    Vendor-neutral OpenTelemetry skills for AI coding agents, grounded in upstream sources

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/ollygarden/opentelemetry-agent-skills ~/.claude/skills/opentelemetry-agent-skills

    97GitHub stars+1 (+1.0 %)
  7. 7

    Skill
    Active

    Agent skills for Arize — datasets, experiments, and traces via the ax CLI

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/Arize-ai/arize-skills ~/.claude/skills/arize-skills

    47GitHub starsstable
  8. 8

    MCP
    Active

    AURA is a production-tested SRE agent platform you can deploy in minutes. AURA handles the guardrails, APIs, state management, streaming, and failure handling required to put AI to work safely on production infrastructure.

    githubmeasured growthOpen source ↗

    291GitHub stars+33 (+12.8 %)
  9. 9

    MCP
    Dormant

    Open-source Python framework to deploy AI agents via HTTP, A2A, and MCP with built-in observability

    githubmeasured growthOpen source ↗

    5GitHub starsstable
  10. 10

    Agent
    Dormant

    AgentOps: Multi-agent infrastructure remediation platform. A2A protocol, agent coordination, HITL approval, auto-rollback.

    githubmeasured growthOpen source ↗

    0GitHub starsstable
  11. 11

    Skill
    Active

    local-first analytics for AI agent skills

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/crafter-station/skill-kit ~/.claude/skills/skill-kit

    77GitHub stars+1 (+1.3 %)
  12. 12

    Other
    Dormant

    Official Python SDK for GT8004 — AI agent observability with MCP, A2A, x402 payment tracking. FastAPI, Flask, FastMCP middleware included.

    githubmeasured growthOpen source ↗

    1GitHub starsstable
  13. 13
    Active

    Log Claude Code sessions to Opik, the open-source LLM observability and evaluation platform, built by Comet. Tracing, evaluation, and skills for observable AI applications.

    githubmeasured growthOpen source ↗

    22GitHub stars+1 (+4.8 %)
  14. 14

    Agent
    Dormant

    Policy-enforced observability and fail-closed guardrails for MCP/A2A multi-agent systems.

    githubmeasured growthOpen source ↗

    26GitHub starsstable
  15. 15
    Dormant

    A Python proof-of-concept for tracing multi-turn Agent-to-Agent (A2A) conversations as a single unified MLflow trace for LLM observability and evaluation.

    githubmeasured growthOpen source ↗

    0GitHub starsstable
  16. 16

    Skill
    Active

    Measure prompt and skill improvements with blind A/B comparison.

    githubmeasured growthOpen source ↗

    Install /plugin marketplace add shinpr/rashomon

    18GitHub starsstable
  17. 17
    Active

    Dashboard for monitoring claude code sessions.

    githubmeasured growthOpen source ↗

    Install /plugin marketplace add JayantDevkar/claude-code-karma

    323GitHub stars+2 (+0.62 %)
  18. 18

    Skill
    Active

    A skill creator that proves its skills work. Evidence-driven skill creation for Claude Code and Codex: baseline-tested generation, per-skill regression evals, ecosystem doctor, cross-runtime compile, and an opt-in proactive advisor.

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/tripleyak/SkillForge ~/.claude/skills/SkillForge

    884GitHub stars-1 (-0.11 %)
  19. 19
    Dormant

    🔍 AI observability skill for Claude Code. Debug LangChain/LangGraph agents by fetching execution traces from LangSmith Studio directly in your terminal.

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/OthmanAdi/langsmith-fetch-skill ~/.claude/skills/langsmith-fetch-skill

    27GitHub starsstable
  20. 20
    Active

    Self improving agents through iterations

    githubmeasured growthOpen source ↗

    105GitHub stars+1 (+0.96 %)
  21. 21
    Dormant

    An AI-powered multi-agent system that demonstrates clinical triage, OTC medication recommendations, and e-pharmacy integration for respiratory conditions. Built with modular agents that collaborate to provide safe, intelligent healthcare…

    githubmeasured growthOpen source ↗

    0GitHub starsstable
  22. 22
    Dormant

    A safety-first multi-agent mental health companion with real-time distress tracking, triple-layer guardrails, and evidence-based grounding techniques. Built for Kaggle × Google Agents Intensive 2025 Capstone (Agents for Good Track)

    githubmeasured growthOpen source ↗

    1GitHub starsstable
  23. 23

    Other
    Dormant

    Multi-agent customer support system with Google ADK & Gemini 2.5 Flash Lite. Kaggle capstone demonstrating 11+ concepts. Automates 80%+ queries, <10s response time.

    githubmeasured growthOpen source ↗

    2GitHub starsstable
  24. 24
    Dormant

    Local open-source dev tool to debug, secure, and evaluate LLM agents. Provides static analysis, dynamic security checks, and runtime monitoring - integrates with Cursor and Claude Code.

    githubmeasured growthOpen source ↗

    Install /plugin marketplace add cylestio/agent-inspector

    9GitHub starsstable
  25. 25
    Active

    Build tested agent skills and govern their lifecycle through a user-defined marketplace: evidence, discovery, updates, rollback, quarantine, and 17-platform distribution.

    githubestimated momentumOpen source ↗

    Install git clone https://github.com/FrancyJGLisboa/agent-skills-platform ~/.claude/skills/agent-skills-platform

    2 376GitHub stars