LLM observability: tools for developers
Monitoring, traces, evaluation and quality tracking for LLM applications. An entry can appear under several use cases when its topics or description provide several explicit signals.
Use case
Activity
Sort by
122 entries in this view.
Tool ranking
Mixed ranking: measured growth takes priority; entries without two snapshots are estimated. Raw values keep their source-specific window, shown under each tool; a change requires at least two snapshots from the last 7 days.
- 1Active
skill-receipts
SkillClaude Code skills where every entry ships receipts — accuracy-gated benchmarks against baseline and placebo, rejects published
githubmeasured growthOpen source ↗
Install
git clone https://github.com/sjh9714/skill-receipts ~/.claude/skills/skill-receipts2GitHub starsstable - 2Active
Sentry instrumentation skill for system-behavior tracking
githubmeasured growthOpen source ↗
Install
git clone https://github.com/tortastudios/sentry-instrumentation ~/.claude/skills/sentry-instrumentation24GitHub starsstable - 3Dormant
CustoFlow
OtherMulti-agent customer support system with Google ADK & Gemini 2.5 Flash Lite. Kaggle capstone demonstrating 11+ concepts. Automates 80%+ queries, <10s response time.
githubmeasured growthOpen source ↗
2GitHub starsstable - 4Active
Deterministic, local-only audit reports for Claude Code AI agent sessions. Rust + SQLite. Zero network calls. MCP server for agents.
githubmeasured growthOpen source ↗
22GitHub stars+2 (+10.0 %) - 5Active
nudge
MCPA typed, replayable, budget-aware programming language for LLM agents. Compiles to Python & TypeScript.
githubmeasured growthOpen source ↗
2GitHub starsstable - 6Active
Log Claude Code sessions to Opik, the open-source LLM observability and evaluation platform, built by Comet. Tracing, evaluation, and skills for observable AI applications.
githubmeasured growthOpen source ↗
22GitHub stars+1 (+4.8 %) - 7Active
untell
OtherAI-detector auditing toolkit: measures how often a detector flags genuine human writing, how stable its verdict is across seeds, and whether that verdict survives meaning-preserving edits. Detector-in-the-loop measurement harness. Claude…
githubmeasured growthOpen source ↗
18GitHub starsstable - 8Dormant
OpenTelemetry semantic conventions and instrumentation for agent provenance, derivation lineage, and acceptance criteria evaluation. Fills the Microsoft AI stack observability gap.
githubmeasured growthOpen source ↗
0GitHub starsstable - 9Dormant
Agents-eval
AgentA Multi-Agent System (MAS) evaluation framework using PydanticAI that generates and evaluates scientific paper reviews through a three-tiered assessment approach: traditional metrics, LLM-as-a-Judge, and graph-based complexity analysis.
githubmeasured growthOpen source ↗
2GitHub starsstable - 10Active
Multi-agent quality gate skill for Claude Code that researches, reviews, tests and challenges AI-generated work before the final answer.
githubmeasured growthOpen source ↗
Install
git clone https://github.com/ma-nucho-pro/supervisor-skill-claude ~/.claude/skills/supervisor-skill-claude2GitHub starsstable - 11Active
rashomon
SkillMeasure prompt and skill improvements with blind A/B comparison.
githubmeasured growthOpen source ↗
Install
/plugin marketplace add shinpr/rashomon18GitHub starsstable - 12Active
agent-stack
SkillProduction patterns for agent orchestrators, harnesses, evals, MCP/A2A interoperability, memory, provider routing, and LLM usage metering.
githubmeasured growthOpen source ↗
Install
git clone https://github.com/ssheleg/agent-stack ~/.claude/skills/agent-stack2GitHub stars+1 (+100.0 %) - 13Dormant
MCP as a Judge: a behavioral MCP that strengthens AI coding assistants via explicit LLM evaluations
mcpmeasured growthOpen source ↗
Install
claude mcp add mcp-as-a-judge -- uvx mcp-as-a-judge17GitHub starsstable - 14Active
sre-on-call
AgentMulti-agent SRE on-call investigator that auto-triages Slack/Discord infrastructure alerts via AWS Bedrock AgentCore, fanning out to specialized agents (CloudWatch, EKS, Slack/Discord scanners) for parallel investigation.
githubmeasured growthOpen source ↗
3GitHub starsstable - 15Active
deslop-GPT
SkillDeletion-first Agent Skill for removing test bloat, verification theater, and speculative fallbacks while preserving behavior.
githubmeasured growthOpen source ↗
Install
git clone https://github.com/MrZoyo/deslop-GPT ~/.claude/skills/deslop-GPT76GitHub stars+48 (+171.4 %) - 16Active
Mini Program Engineering Skill Suite is an Agent Skill suite for evidence‑first mini‑program development. It helps agents bring WeChat or other mini‑program projects from vague intent to reliable engineering work: project intake, product…
githubmeasured growthOpen source ↗
Install
git clone https://github.com/NocodeMrLi/mini-program-engineering-skill-suite ~/.claude/skills/mini-program-engineering-skill-suite43GitHub stars+4 (+10.3 %) - 17Active
Build tested agent skills and govern their lifecycle through a user-defined marketplace: evidence, discovery, updates, rollback, quarantine, and 17-platform distribution.
githubestimated momentumOpen source ↗
Install
git clone https://github.com/FrancyJGLisboa/agent-skills-platform ~/.claude/skills/agent-skills-platform2 376GitHub stars—
Learning resources
Ranked by measured growth, normalized across sources. These resources remain available separately and do not take part in the main tool ranking.
- 1ActiveOpen source ↗
trigger_tree
SkillDocumentation-discovery telemetry for Claude Code — heat/cold maps, health grade, evidence-backed router fixes. 100% local, zero tokens. /tt
githubInstall
Resourcemeasured growth/plugin marketplace add Hedde/trigger_tree14GitHub stars+1 (+7.7 %) - 2ActiveOpen source ↗
Handbook técnico aberto sobre engenharia de IA em produção: ML tradicional, LLMs, RAG, agentes, segurança, observabilidade, FinOps e deployment.
githubResourcemeasured growth1GitHub starsstable - 3ActiveOpen source ↗
50+ curated LLM observability tools PLUS 26 Agent Skills (several with runnable, unit-tested scripts) to build, evaluate, debug, secure & monitor reliable LLM apps. Tracing, evals, guardrails, LLMOps.
githubInstall
Resourcemeasured growthgit clone https://github.com/ContextJet-ai/awesome-llm-observability ~/.claude/skills/awesome-llm-observability33GitHub stars+3 (+10.0 %) - 4ActiveOpen source ↗
Agentic_AI_Engineer
AgentMy complete journey to becoming an Agentic AI Engineer through structured learning, projects, experiments, and production-ready implementations of modern AI systems.
githubResourcemeasured growth18GitHub starsstable - 5ActiveOpen source ↗
tunelab
SkillClaude Code plugin for LLM fine-tuning, distillation, and evaluation — decide whether you need fine-tuning at all, distill your LLM logs into small local models (MLX/LoRA), evaluate with held-out discipline, and learn the why at every step.
githubInstall
Resourcemeasured growth/plugin marketplace add rchaz/tunelab6GitHub starsstable