LLM observability: tools for developers

Monitoring, traces, evaluation and quality tracking for LLM applications. An entry can appear under several use cases when its topics or description provide several explicit signals.

Use case

Activity

Sort by

122 entries in this view.

Tool ranking

Mixed ranking: measured growth takes priority; entries without two snapshots are estimated. Raw values keep their source-specific window, shown under each tool; a change requires at least two snapshots from the last 7 days.

  1. 1
    Active

    Make Claude write clearly, for everyone.

    githubmeasured growthOpen source ↗

    Install /plugin marketplace add stefanobaghino/simple-output-styles

    16GitHub starsstable
  2. 2

    Skill
    Active

    Marketing Mix Model expert agent plugin for Claude Code - pymc-marketing v0.18.2+

    githubmeasured growthOpen source ↗

    Install /plugin marketplace add Yakoub-ai/agent-mmm

    4GitHub starsstable
  3. 3

    Agent
    Active

    The agent framework where the model never holds the trigger — every consequential action clears your policy first, waits for a human when it matters, and lands on a record you can verify. Build on it, or put it around the agent you already…

    githubmeasured growthOpen source ↗

    1GitHub starsstable
  4. 4

    Agent
    Dormant

    A Multi-Agent System (MAS) evaluation framework using PydanticAI that generates and evaluates scientific paper reviews through a three-tiered assessment approach: traditional metrics, LLM-as-a-Judge, and graph-based complexity analysis.

    githubmeasured growthOpen source ↗

    2GitHub starsstable
  5. 5
    Dormant

    MCP as a Judge: a behavioral MCP that strengthens AI coding assistants via explicit LLM evaluations

    mcpmeasured growthOpen source ↗

    Install claude mcp add mcp-as-a-judge -- uvx mcp-as-a-judge

    17GitHub starsstable
  6. 6
    Dormant

    An AI-powered multi-agent system that demonstrates clinical triage, OTC medication recommendations, and e-pharmacy integration for respiratory conditions. Built with modular agents that collaborate to provide safe, intelligent healthcare…

    githubmeasured growthOpen source ↗

    0GitHub starsstable
  7. 7
    Active

    A Model Context Protocol (MCP) server for Langfuse, enabling AI agents to query Langfuse trace data for enhanced debugging and observability

    githubmeasured growthOpen source ↗

    105GitHub starsstable
  8. 8

    Other
    Dormant

    Multi-agent customer support system with Google ADK & Gemini 2.5 Flash Lite. Kaggle capstone demonstrating 11+ concepts. Automates 80%+ queries, <10s response time.

    githubmeasured growthOpen source ↗

    2GitHub starsstable
  9. 9
    Active

    Architecture-first Python scaffold for an auditable multi-agent corporate credit desk using A2A, MCP, deterministic credit policies, model routing, and OpenTelemetry.

    githubmeasured growthOpen source ↗

    0GitHub starsstable
  10. 10

    Skill
    Active

    The container platform tailored for Kubernetes multi-cloud, datacenter, and edge management ⎈ 🖥 ☁️

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/kubesphere/kubesphere ~/.claude/skills/kubesphere

    17 035GitHub starsstable
  11. 11
    Dormant

    Two Claude Code skills for adversarial code review (single-model PAR and multi-model MMAR with cross-critique to catch hallucinations), plus a fixture-based eval suite.

    githubmeasured growthOpen source ↗

    Install /plugin marketplace add prime-radiant-inc/parallel-adversarial-review

    17GitHub starsstable
  12. 12
    Active

    Build tested agent skills and govern their lifecycle through a user-defined marketplace: evidence, discovery, updates, rollback, quarantine, and 17-platform distribution.

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/FrancyJGLisboa/agent-skills-platform ~/.claude/skills/agent-skills-platform

    2 376GitHub starsstable
  13. 13
    Active

    Run many AIs on one board and keep control of all of it. Deterministic code decides who acts — never a model. A privacy floor keeps sensitive work on your machine, your own tests decide what counts as done, and every action lands on a…

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/sandhusukhdeep2/sc-prism-releases ~/.claude/skills/sc-prism-releases

    1GitHub starsstable
  14. 14

    Skill
    Active

    Glanceable Claude Code and Codex state in your terminal tabs: white=idle, blue=working, orange=waiting. Multi-terminal (iTerm2, WezTerm, AI Power Term), a live status bar with Anthropic usage limits, /sfl and /nil window save-and-restore,…

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/wasulajr/headsup ~/.claude/skills/headsup

    1GitHub starsstable
  15. 15
    Active

    Self-hosted AI SRE for Kubernetes — zero-instrumentation eBPF observability plus a copilot that fixes issues through guardrailed, self-verifying actions. BYO-LLM, air-gapped capable.

    githubmeasured growthOpen source ↗

    131GitHub stars-4 (-3.0 %)
  16. 16
    Active

    74 open-source Agent Skills for Claude Code and Codex: AI SEO, AEO and GEO, code review with an A-F ship grade, CI gates, AI evals, design systems, conversion copy, Instagram growth, iOS and Android app shipping, creator rights, and…

    githubmeasured growthOpen source ↗

    129GitHub stars-4 (-3.0 %)
  17. 17

    Other
    Active

    Open-source observability tool that uses AI agents to self-heal your software

    githubestimated momentumOpen source ↗

    1 404GitHub stars

Learning resources

Ranked by measured growth, normalized across sources. These resources remain available separately and do not take part in the main tool ranking.

  1. 1
    ActiveOpen source ↗

    My complete journey to becoming an Agentic AI Engineer through structured learning, projects, experiments, and production-ready implementations of modern AI systems.

    githubResourcemeasured growth
    18GitHub starsstable
  2. 2
    ActiveOpen source ↗

    Curated + scored map of official MCP servers and agents for DevOps, Cloud, SRE, and Platform Engineering — every entry rated on production access, approval gates, and audit evidence.

    githubResourcemeasured growth
    77GitHub starsstable
  3. 3
    ActiveOpen source ↗

    Handbook técnico aberto sobre engenharia de IA em produção: ML tradicional, LLMs, RAG, agentes, segurança, observabilidade, FinOps e deployment.

    githubResourcemeasured growth
    1GitHub starsstable
  4. 4
    ActiveOpen source ↗

    Documentation-discovery telemetry for Claude Code — heat/cold maps, health grade, evidence-backed router fixes. 100% local, zero tokens. /tt

    github

    Install /plugin marketplace add Hedde/trigger_tree

    Resourcemeasured growth
    14GitHub starsstable
  5. 5

    tunelab

    Skill
    ActiveOpen source ↗

    Claude Code plugin for LLM fine-tuning, distillation, and evaluation — decide whether you need fine-tuning at all, distill your LLM logs into small local models (MLX/LoRA), evaluate with held-out discipline, and learn the why at every step.

    github

    Install /plugin marketplace add rchaz/tunelab

    Resourcemeasured growth
    6GitHub starsstable