LLM observability: tools for developers

Monitoring, traces, evaluation and quality tracking for LLM applications. An entry can appear under several use cases when its topics or description provide several explicit signals.

Use case

Activity

Sort by

122 entries in this view.

Tool ranking

Ranked by measured growth, normalized across sources. Raw values keep their source-specific window, shown under each tool; a change requires at least two snapshots from the last 7 days.

  1. 1
    Active

    Claude Code skills where every entry ships receipts — accuracy-gated benchmarks against baseline and placebo, rejects published

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/sjh9714/skill-receipts ~/.claude/skills/skill-receipts

    2GitHub starsstable
  2. 2

    Skill
    Active

    Production patterns for agent orchestrators, harnesses, evals, MCP/A2A interoperability, memory, provider routing, and LLM usage metering.

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/ssheleg/agent-stack ~/.claude/skills/agent-stack

    2GitHub stars+1 (+100.0 %)
  3. 3

    Agent
    Active

    A lightweight Python library that decouples agentic runtime from applications it builds

    githubmeasured growthOpen source ↗

    1GitHub starsstable
  4. 4
    Active

    Multi-agent quality gate skill for Claude Code that researches, reviews, tests and challenges AI-generated work before the final answer.

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/ma-nucho-pro/supervisor-skill-claude ~/.claude/skills/supervisor-skill-claude

    2GitHub starsstable
  5. 5

    Other
    Dormant

    Multi-agent customer support system with Google ADK & Gemini 2.5 Flash Lite. Kaggle capstone demonstrating 11+ concepts. Automates 80%+ queries, <10s response time.

    githubmeasured growthOpen source ↗

    2GitHub starsstable
  6. 6
    Active

    🪢 Langfuse documentation -- Langfuse is the open source LLM Engineering Platform. Observability, evals, prompt management, playground and metrics to debug and improve LLM apps

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/langfuse/langfuse-docs ~/.claude/skills/langfuse-docs

    239GitHub stars+3 (+1.3 %)
  7. 7

    Agent
    Dormant

    A Multi-Agent System (MAS) evaluation framework using PydanticAI that generates and evaluates scientific paper reviews through a three-tiered assessment approach: traditional metrics, LLM-as-a-Judge, and graph-based complexity analysis.

    githubmeasured growthOpen source ↗

    2GitHub starsstable
  8. 8
    Active

    Deterministic crash tests for agent control planes: paired worlds, hard stops, budgets, scopes, approvals, and hash-linked receipts.

    githubmeasured growthOpen source ↗

    1GitHub starsstable
  9. 9
    Active

    Agent 降智检测与自愈公评网络 — an immune system for the AI agent society

    githubmeasured growthOpen source ↗

    1GitHub starsstable
  10. 10
    Dormant

    A safety-first multi-agent mental health companion with real-time distress tracking, triple-layer guardrails, and evidence-based grounding techniques. Built for Kaggle × Google Agents Intensive 2025 Capstone (Agents for Good Track)

    githubmeasured growthOpen source ↗

    1GitHub starsstable
  11. 11
    Active

    Local-first Agentic AI Infrastructure Platform with LLM Gateway, Agent DAG Runtime, MCP Tool Hub, A2A Agent Mesh, Local RAG, Tool Sandbox and Observability.

    githubmeasured growthOpen source ↗

    1GitHub starsstable
  12. 12

    Agent
    Dormant

    The reliability and interoperability layer for AI agents. A2A-native. MCP-bridged.

    githubmeasured growthOpen source ↗

    1GitHub starsstable
  13. 13
    Active

    🟪 Open-source runtime that ships any LangGraph or Google ADK agent as a production-ready FastAPI service. Bundled , AG-UI copilotkit API, chat UI, 15+ guardrails, MCP, OpenTelemetry, OIDC. One pip install. Self-hosted, no vendor lock-in.

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/Idun-Group/idun-agent-platform ~/.claude/skills/idun-agent-platform

    198GitHub stars-1 (-0.50 %)
  14. 14

    Agent
    Active

    A local A2A event and continuity service for agent applications, with structured history, provenance-safe transcripts, literal search, and natural-language queries.

    githubmeasured growthOpen source ↗

    1GitHub stars+1
  15. 15
    Active

    Skill instalable para Claude Code que mide si tu agente cumple las reglas establecidas. Lee el historial de sesiones de tu disco y devuelve un porcentaje por regla de cumplimiento. Sin instalar nada, sin red, biblioteca estándar.

    githubmeasured growthOpen source ↗

    Install /plugin marketplace add jleonceo/skill-adherencia-reglas

    1GitHub starsstable
  16. 16
    Active

    Mide si tu agente de IA cumple las reglas que le escribiste. Lee el historial local de Claude Code y devuelve un porcentaje por regla. Sin instalar nada, sin red, biblioteca estándar.

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/jleonceo/adherencia-reglas ~/.claude/skills/adherencia-reglas

    1GitHub starsstable
  17. 17

    Agent
    Dormant

    Contract-based testing for LLM agents: hybrid evaluation, multi-agent + A2A, deterministic replay.

    githubmeasured growthOpen source ↗

    0GitHub starsstable
  18. 18
    Active

    8 MCP SMB products — standalone AI servers for SMBs: Guardrails, FinOps, Observability, Router, Trust Score, Memory, ThinkSecure, A2A Lite. 37 tools, TNC credits billing.

    githubmeasured growthOpen source ↗

    1GitHub starsstable
  19. 19

    Skill
    Active

    The container platform tailored for Kubernetes multi-cloud, datacenter, and edge management ⎈ 🖥 ☁️

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/kubesphere/kubesphere ~/.claude/skills/kubesphere

    17 037GitHub stars+8 (+0.05 %)
  20. 20

    Skill
    Active

    Research-backed, eval-driven skills for AI agents

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/ZSeven-W/craft-skills ~/.claude/skills/craft-skills

    142GitHub stars+18 (+14.5 %)
  21. 21
    Active

    74 open-source Agent Skills for Claude Code and Codex: AI SEO, AEO and GEO, code review with an A-F ship grade, CI gates, AI evals, design systems, conversion copy, Instagram growth, iOS and Android app shipping, creator rights, and…

    githubmeasured growthOpen source ↗

    129GitHub stars-4 (-3.0 %)
  22. 22

    Agent
    Active

    面向长程研发任务的 Java Agent Harness,支持 A2A 跨语言协作、可恢复执行、上下文工程与 Eval 驱动开发。

    githubmeasured growthOpen source ↗

    0GitHub starsstable
  23. 23
    Active

    🐙 ADLC Team Skills — Agentic SDLC for Engineering Teams

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/tikalk/adlc-team-skills ~/.claude/skills/adlc-team-skills

    132GitHub stars+1 (+0.76 %)
  24. 24
    Active

    Skills, prompts, and instructions for building AI agents on top of Dynatrace production context

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/Dynatrace/dynatrace-for-ai ~/.claude/skills/dynatrace-for-ai

    132GitHub stars+4 (+3.1 %)
  25. 25
    Active

    A Model Context Protocol (MCP) server for Langfuse, enabling AI agents to query Langfuse trace data for enhanced debugging and observability

    githubmeasured growthOpen source ↗

    105GitHub starsstable