LLM observability: tools for developers
Monitoring, traces, evaluation and quality tracking for LLM applications. An entry can appear under several use cases when its topics or description provide several explicit signals.
Use case
Activity
Sort by
117 entries in this view.
Tool ranking
Ranked by measured growth, normalized across sources. Raw values keep their source-specific window, shown under each tool; a change requires at least two snapshots from the last 7 days.
- 1Active
rocketplaneIO
OtherSelf-hosted AI SRE for Kubernetes — zero-instrumentation eBPF observability plus a copilot that fixes issues through guardrailed, self-verifying actions. BYO-LLM, air-gapped capable.
githubmeasured growthOpen source ↗
130GitHub stars-11 (-7.8 %) - 2Active
Open-source Agent Skills for Claude Code and Codex: ship-DAG orchestration, A-F code review, AI evals, CI gates, design, copy, SEO/AEO/GEO, Instagram growth, app shipping, creator rights, and consumer recovery.
githubmeasured growthOpen source ↗
129GitHub stars-10 (-7.2 %) - 3Active
langfuse-mcp
MCPA Model Context Protocol (MCP) server for Langfuse, enabling AI agents to query Langfuse trace data for enhanced debugging and observability
githubmeasured growthOpen source ↗
105GitHub starsstable - 4Dormant
anti-lie
SkillDon't make LLMs honest. Make every factual claim auditable. — An LLM Claim Auditing Layer with T1-T7 truth gradients. 98.1% business effectiveness on LiarBench v0.2.
githubmeasured growthOpen source ↗
Install
git clone https://github.com/lc198707/anti-lie ~/.claude/skills/anti-lie89GitHub starsstable - 5Active
AgentX-Python
OtherAgentX python SDK. Build multi-agent AI workforce. Run evaluation. Trace your agent. Full Observability.
githubmeasured growthOpen source ↗
68GitHub starsstable - 6Active
seo-skill-bench
SkillOpen benchmark for Claude Code SEO skills — real headless execution against fixture sites with planted-defect answer keys. Deterministic scoring, pre-registered rubric.
githubmeasured growthOpen source ↗
Install
git clone https://github.com/aleclindz/seo-skill-bench ~/.claude/skills/seo-skill-bench50GitHub stars-6 (-10.7 %) - 7Active
arize-skills
SkillAgent skills for Arize — datasets, experiments, and traces via the ax CLI
githubmeasured growthOpen source ↗
Install
git clone https://github.com/Arize-ai/arize-skills ~/.claude/skills/arize-skills47GitHub starsstable - 8Dormant
research-proof
SkillPressure-test research claims with falsifiable evidence plans, adversarial checks, frozen verifiers, and proof ledgers.
githubmeasured growthOpen source ↗
Install
git clone https://github.com/tonyblu331/research-proof ~/.claude/skills/research-proof45GitHub starsstable - 9Dormant
🔍 AI observability skill for Claude Code. Debug LangChain/LangGraph agents by fetching execution traces from LangSmith Studio directly in your terminal.
githubmeasured growthOpen source ↗
Install
git clone https://github.com/OthmanAdi/langsmith-fetch-skill ~/.claude/skills/langsmith-fetch-skill27GitHub starsstable - 10Dormant
astragraph
AgentPolicy-enforced observability and fail-closed guardrails for MCP/A2A multi-agent systems.
githubmeasured growthOpen source ↗
26GitHub starsstable - 11Active
Sentry instrumentation skill for system-behavior tracking
githubmeasured growthOpen source ↗
Install
git clone https://github.com/tortastudios/sentry-instrumentation ~/.claude/skills/sentry-instrumentation24GitHub starsstable - 12Active
Log Claude Code sessions to Opik, the open-source LLM observability and evaluation platform, built by Comet. Tracing, evaluation, and skills for observable AI applications.
githubmeasured growthOpen source ↗
21GitHub starsstable - 13Active
rashomon
SkillMeasure prompt and skill improvements with blind A/B comparison.
githubmeasured growthOpen source ↗
Install
/plugin marketplace add shinpr/rashomon18GitHub starsstable - 14Dormant
MCP as a Judge: a behavioral MCP that strengthens AI coding assistants via explicit LLM evaluations
mcpmeasured growthOpen source ↗
Install
claude mcp add mcp-as-a-judge -- uvx mcp-as-a-judge17GitHub starsstable - 15Active
Two Claude Code skills for adversarial code review (single-model PAR and multi-model MMAR with cross-critique to catch hallucinations), plus a fixture-based eval suite.
githubmeasured growthOpen source ↗
Install
/plugin marketplace add prime-radiant-inc/parallel-adversarial-review17GitHub starsstable - 16Active
simple-output-styles
SkillMake Claude write clearly, for everyone.
githubmeasured growthOpen source ↗
Install
/plugin marketplace add stefanobaghino/simple-output-styles16GitHub starsstable - 17Active
Claude Code plugin marketplace — agentic-engineering (spec-driven shape→decide→execute→measure→eval with adversarial review) + github-keeper (audit/elevate READMEs and make a repo open-source-ready).
githubmeasured growthOpen source ↗
Install
/plugin marketplace add GiustoPiedimonte/agentic-engineering-marketplace13GitHub starsstable - 18Active
adl-cli
OtherA command-line tool to scaffold and manage enterprise-ready AI Agents powered by the A2A (Agent-to-Agent) protocol
githubmeasured growthOpen source ↗
13GitHub starsstable - 19Active
Query Spanlens LLM observability from Cursor, Claude Desktop, or Continue via MCP.
mcpmeasured growthOpen source ↗
Install
claude mcp add mcp-server -- npx @spanlens/mcp-server12GitHub starsstable - 20Active
CLI for benchmarks & evals of AI coding agents — on tasks you already understand, using your Claude / Codex / Gemini individual subscriptions or API keys.
githubmeasured growthOpen source ↗
12GitHub starsstable - 21Active
bakeoff
SkillTurn one decision into a judged tournament of solutions, then pick the best — a Claude Code skill that generates candidates, auto-derives the rubric, judges independently, and returns a defensible winner.
githubmeasured growthOpen source ↗
Install
/plugin marketplace add CoriChui/bakeoff10GitHub starsstable - 22Active
galdor
OtherA Go-native framework for LLM agents, with OpenTelemetry observability built in.
githubmeasured growthOpen source ↗
10GitHub starsstable - 23Dormant
agent-inspector
SkillLocal open-source dev tool to debug, secure, and evaluate LLM agents. Provides static analysis, dynamic security checks, and runtime monitoring - integrates with Cursor and Claude Code.
githubmeasured growthOpen source ↗
Install
/plugin marketplace add cylestio/agent-inspector9GitHub starsstable - 24Active
anchor
SkillAnchor — the production-grade AGENTS.md template for AI coding agents. 51 sections of battle-tested rules. Keep your agents grounded.
githubmeasured growthOpen source ↗
Install
git clone https://github.com/peva3/anchor ~/.claude/skills/anchor8GitHub starsstable
Learning resources
Ranked by measured growth, normalized across sources. These resources remain available separately and do not take part in the main tool ranking.
- 1ActiveOpen source ↗
Agentic_AI_Engineer
AgentMy complete journey to becoming an Agentic AI Engineer through structured learning, projects, experiments, and production-ready implementations of modern AI systems.
githubResourcemeasured growth18GitHub starsstable