工具说明为英文。
LLM 可观测性
LLM 应用的监控、追踪、评估和质量管理。
用途
活跃度
排序方式
121
工具排名
Ranked by known GitHub stars, highest first; archived repositories come after maintained repositories. 原始值保留各自的数据窗口;计算变化需要 7 天内至少两次测量。
- 1活跃
Real-time execution trace and cost intelligence for Claude Code
github实测增长打开来源 ↗
安装
/plugin marketplace add DeibyGS/claudestat34GitHub 星标稳定 - 2休眠
🔍 AI observability skill for Claude Code. Debug LangChain/LangGraph agents by fetching execution traces from LangSmith Studio directly in your terminal.
github实测增长打开来源 ↗
安装
git clone https://github.com/OthmanAdi/langsmith-fetch-skill ~/.claude/skills/langsmith-fetch-skill27GitHub 星标稳定 - 3休眠
astragraph
智能体Policy-enforced observability and fail-closed guardrails for MCP/A2A multi-agent systems.
github实测增长打开来源 ↗
26GitHub 星标稳定 - 4活跃
Sentry instrumentation skill for system-behavior tracking
github实测增长打开来源 ↗
安装
git clone https://github.com/tortastudios/sentry-instrumentation ~/.claude/skills/sentry-instrumentation24GitHub 星标稳定 - 5活跃
Log Claude Code sessions to Opik, the open-source LLM observability and evaluation platform, built by Comet. Tracing, evaluation, and skills for observable AI applications.
github实测增长打开来源 ↗
22GitHub 星标+1 (+4.8 %) - 6活跃
Deterministic, local-only audit reports for Claude Code AI agent sessions. Rust + SQLite. Zero network calls. MCP server for agents.
github实测增长打开来源 ↗
22GitHub 星标稳定 - 7活跃
untell
其他AI-detector auditing toolkit: measures how often a detector flags genuine human writing, how stable its verdict is across seeds, and whether that verdict survives meaning-preserving edits. Detector-in-the-loop measurement harness. Claude…
github实测增长打开来源 ↗
18GitHub 星标-1 (-5.3 %) - 8活跃
rashomon
技能Measure prompt and skill improvements with blind A/B comparison.
github实测增长打开来源 ↗
安装
/plugin marketplace add shinpr/rashomon18GitHub 星标稳定 - 9休眠
Two Claude Code skills for adversarial code review (single-model PAR and multi-model MMAR with cross-critique to catch hallucinations), plus a fixture-based eval suite.
github实测增长打开来源 ↗
安装
/plugin marketplace add prime-radiant-inc/parallel-adversarial-review17GitHub 星标稳定 - 10休眠
MCP as a Judge: a behavioral MCP that strengthens AI coding assistants via explicit LLM evaluations
mcp实测增长打开来源 ↗
安装
claude mcp add mcp-as-a-judge -- uvx mcp-as-a-judge17GitHub 星标稳定 - 11活跃
Make Claude write clearly, for everyone.
github实测增长打开来源 ↗
安装
/plugin marketplace add stefanobaghino/simple-output-styles16GitHub 星标稳定 - 12活跃
adl-cli
其他A command-line tool to scaffold and manage enterprise-ready AI Agents powered by the A2A (Agent-to-Agent) protocol
github实测增长打开来源 ↗
14GitHub 星标+1 (+7.7 %) - 13活跃
Query Spanlens LLM observability from Cursor, Claude Desktop, or Continue via MCP.
mcp实测增长打开来源 ↗
安装
claude mcp add mcp-server -- npx @spanlens/mcp-server13GitHub 星标+1 (+8.3 %) - 14活跃
Claude Code plugin marketplace — agentic-engineering (spec-driven shape→decide→execute→measure→eval with adversarial review) + github-keeper (audit/elevate READMEs and make a repo open-source-ready).
github实测增长打开来源 ↗
安装
/plugin marketplace add GiustoPiedimonte/agentic-engineering-marketplace13GitHub 星标稳定 - 15休眠
A Claude Code skill that adds a rubric-based eval layer to any agent project. Framework-agnostic — generates rubric, test cases, judge prompt, and harness. Returns a weighted score plus a judge-leniency signal.
github实测增长打开来源 ↗
13GitHub 星标稳定 - 16活跃
Audit which Claude Code skills you actually use — surface dead installs and hallucinated invocations from your session logs.
github实测增长打开来源 ↗
安装
/plugin marketplace add sfrangulov/skill-graveyard10GitHub 星标稳定 - 17活跃
galdor
其他A Go-native framework for LLM agents, with OpenTelemetry observability built in.
github实测增长打开来源 ↗
10GitHub 星标稳定 - 18活跃
bakeoff
技能Turn one decision into a judged tournament of solutions, then pick the best — a Claude Code skill that generates candidates, auto-derives the rubric, judges independently, and returns a defensible winner.
github实测增长打开来源 ↗
安装
/plugin marketplace add CoriChui/bakeoff10GitHub 星标稳定 - 19活跃
Production-grade AI coding rules for Cursor and Claude Code. 15 rules + 9 doc templates + skills + agents + MCP setup. Drop into any project.
github实测增长打开来源 ↗
安装
git clone https://github.com/aiagentwithdhruv/ai-dev-stack ~/.claude/skills/ai-dev-stack10GitHub 星标稳定 - 20休眠
Local open-source dev tool to debug, secure, and evaluate LLM agents. Provides static analysis, dynamic security checks, and runtime monitoring - integrates with Cursor and Claude Code.
github实测增长打开来源 ↗
安装
/plugin marketplace add cylestio/agent-inspector9GitHub 星标稳定 - 21活跃
anchor
技能Anchor — the production-grade AGENTS.md template for AI coding agents. 51 sections of battle-tested rules. Keep your agents grounded.
github实测增长打开来源 ↗
安装
git clone https://github.com/peva3/anchor ~/.claude/skills/anchor8GitHub 星标稳定 - 22休眠
A self-improving harness router for Claude Code.
github实测增长打开来源 ↗
安装
/plugin marketplace add SeongwoongCho/adaptive-harness8GitHub 星标稳定
学习与参考资源
Ranked by known GitHub stars, highest first; archived repositories come after maintained repositories. 这些资源可单独访问,不参与主要排名。
- 1活跃打开来源 ↗
50+ curated LLM observability tools PLUS 26 Agent Skills (several with runnable, unit-tested scripts) to build, evaluate, debug, secure & monitor reliable LLM apps. Tracing, evals, guardrails, LLMOps.
github安装
资源实测增长git clone https://github.com/ContextJet-ai/awesome-llm-observability ~/.claude/skills/awesome-llm-observability33GitHub 星标+2 (+6.5 %) - 2活跃打开来源 ↗
My complete journey to becoming an Agentic AI Engineer through structured learning, projects, experiments, and production-ready implementations of modern AI systems.
github资源实测增长18GitHub 星标稳定 - 3活跃打开来源 ↗
Documentation-discovery telemetry for Claude Code — heat/cold maps, health grade, evidence-backed router fixes. 100% local, zero tokens. /tt
github安装
资源实测增长/plugin marketplace add Hedde/trigger_tree14GitHub 星标稳定