LLM observability: tools for developers
Monitoring, traces, evaluation and quality tracking for LLM applications. An entry can appear under several use cases when its topics or description provide several explicit signals.
Use case
Activity
Sort by
117 entries in this view.
Tool ranking
Ranked by measured growth, normalized across sources. Raw values keep their source-specific window, shown under each tool; a change requires at least two snapshots from the last 7 days.
- 1Active
aura
MCPAURA is a production-tested SRE agent platform you can deploy in minutes. AURA handles the guardrails, APIs, state management, streaming, and failure handling required to put AI to work safely on production infrastructure.
githubmeasured growthOpen source ↗
254GitHub stars-3 (-1.2 %) - 2Active
idun-agent-platform
Skill🟪 Open-source runtime that ships any LangGraph or Google ADK agent as a production-ready FastAPI service. Bundled , AG-UI copilotkit API, chat UI, 15+ guardrails, MCP, OpenTelemetry, OIDC. One pip install. Self-hosted, no vendor lock-in.
githubmeasured growthOpen source ↗
Install
git clone https://github.com/Idun-Group/idun-agent-platform ~/.claude/skills/idun-agent-platform198GitHub stars-1 (-0.50 %) - 3Active
rocketplaneIO
OtherSelf-hosted AI SRE for Kubernetes — zero-instrumentation eBPF observability plus a copilot that fixes issues through guardrailed, self-verifying actions. BYO-LLM, air-gapped capable.
githubmeasured growthOpen source ↗
131GitHub stars-4 (-3.0 %) - 4Active
74 open-source Agent Skills for Claude Code and Codex: AI SEO, AEO and GEO, code review with an A-F ship grade, CI gates, AI evals, design systems, conversion copy, Instagram growth, iOS and Android app shipping, creator rights, and…
githubmeasured growthOpen source ↗
129GitHub stars-4 (-3.0 %) - 5Active
langfuse-mcp
MCPA Model Context Protocol (MCP) server for Langfuse, enabling AI agents to query Langfuse trace data for enhanced debugging and observability
githubmeasured growthOpen source ↗
105GitHub starsstable - 6Dormant
anti-lie
SkillDon't make LLMs honest. Make every factual claim auditable. — An LLM Claim Auditing Layer with T1-T7 truth gradients. 98.1% business effectiveness on LiarBench v0.2.
githubmeasured growthOpen source ↗
Install
git clone https://github.com/lc198707/anti-lie ~/.claude/skills/anti-lie89GitHub starsstable - 7Active
MCP server for Langfuse LLM observability — trace and observation analysis.
mcpmeasured growthOpen source ↗
Install
claude mcp add langfuse -- npx langfuse-observability-mcp-server78GitHub starsstable - 8Active
AgentX-Python
OtherAgentX python SDK. Build multi-agent AI workforce. Run evaluation. Trace your agent. Full Observability.
githubmeasured growthOpen source ↗
68GitHub starsstable - 9Active
seo-skill-bench
SkillOpen benchmark for Claude Code SEO skills — real headless execution against fixture sites with planted-defect answer keys. Deterministic scoring, pre-registered rubric.
githubmeasured growthOpen source ↗
Install
git clone https://github.com/aleclindz/seo-skill-bench ~/.claude/skills/seo-skill-bench51GitHub stars-5 (-8.9 %) - 10Active
arize-skills
SkillAgent skills for Arize — datasets, experiments, and traces via the ax CLI
githubmeasured growthOpen source ↗
Install
git clone https://github.com/Arize-ai/arize-skills ~/.claude/skills/arize-skills47GitHub starsstable - 11Active
opus-fable-playbook
SkillMake Claude Opus 4.8 behave like Claude Fable 5 — doctrine output style, drift-catching hooks, and an eval loop against golden Fable transcripts. Claude Code plugin.
githubmeasured growthOpen source ↗
Install
/plugin marketplace add rennf93/opus-fable-playbook33GitHub starsstable - 12Dormant
🔍 AI observability skill for Claude Code. Debug LangChain/LangGraph agents by fetching execution traces from LangSmith Studio directly in your terminal.
githubmeasured growthOpen source ↗
Install
git clone https://github.com/OthmanAdi/langsmith-fetch-skill ~/.claude/skills/langsmith-fetch-skill27GitHub starsstable - 13Dormant
astragraph
AgentPolicy-enforced observability and fail-closed guardrails for MCP/A2A multi-agent systems.
githubmeasured growthOpen source ↗
26GitHub starsstable - 14Active
Sentry instrumentation skill for system-behavior tracking
githubmeasured growthOpen source ↗
Install
git clone https://github.com/tortastudios/sentry-instrumentation ~/.claude/skills/sentry-instrumentation24GitHub starsstable - 15Active
untell
OtherAI-detector auditing toolkit: measures how often a detector flags genuine human writing, how stable its verdict is across seeds, and whether that verdict survives meaning-preserving edits. Detector-in-the-loop measurement harness. Claude…
githubmeasured growthOpen source ↗
18GitHub starsstable - 16Active
rashomon
SkillMeasure prompt and skill improvements with blind A/B comparison.
githubmeasured growthOpen source ↗
Install
/plugin marketplace add shinpr/rashomon18GitHub starsstable - 17Active
Two Claude Code skills for adversarial code review (single-model PAR and multi-model MMAR with cross-critique to catch hallucinations), plus a fixture-based eval suite.
githubmeasured growthOpen source ↗
Install
/plugin marketplace add prime-radiant-inc/parallel-adversarial-review17GitHub starsstable - 18Active
simple-output-styles
SkillMake Claude write clearly, for everyone.
githubmeasured growthOpen source ↗
Install
/plugin marketplace add stefanobaghino/simple-output-styles16GitHub starsstable - 19Active
Claude Code plugin marketplace — agentic-engineering (spec-driven shape→decide→execute→measure→eval with adversarial review) + github-keeper (audit/elevate READMEs and make a repo open-source-ready).
githubmeasured growthOpen source ↗
Install
/plugin marketplace add GiustoPiedimonte/agentic-engineering-marketplace13GitHub starsstable - 20Dormant
eval-layer
OtherA Claude Code skill that adds a rubric-based eval layer to any agent project. Framework-agnostic — generates rubric, test cases, judge prompt, and harness. Returns a weighted score plus a judge-leniency signal.
githubmeasured growthOpen source ↗
13GitHub starsstable - 21Active
Query Spanlens LLM observability from Cursor, Claude Desktop, or Continue via MCP.
mcpmeasured growthOpen source ↗
Install
claude mcp add mcp-server -- npx @spanlens/mcp-server12GitHub starsstable - 22Active
bakeoff
SkillTurn one decision into a judged tournament of solutions, then pick the best — a Claude Code skill that generates candidates, auto-derives the rubric, judges independently, and returns a defensible winner.
githubmeasured growthOpen source ↗
Install
/plugin marketplace add CoriChui/bakeoff10GitHub starsstable - 23Active
galdor
OtherA Go-native framework for LLM agents, with OpenTelemetry observability built in.
githubmeasured growthOpen source ↗
10GitHub starsstable
Learning resources
Ranked by measured growth, normalized across sources. These resources remain available separately and do not take part in the main tool ranking.
- 1ActiveOpen source ↗
Curated + scored map of official MCP servers and agents for DevOps, Cloud, SRE, and Platform Engineering — every entry rated on production access, approval gates, and audit evidence.
githubResourcemeasured growth77GitHub starsstable - 2ActiveOpen source ↗
Agentic_AI_Engineer
AgentMy complete journey to becoming an Agentic AI Engineer through structured learning, projects, experiments, and production-ready implementations of modern AI systems.
githubResourcemeasured growth18GitHub starsstable