工具说明为英文。

LLM 可观测性

LLM 应用的监控、追踪、评估和质量管理。

用途

活跃度

排序方式

122

工具排名

Ranked by normalized popularity across sources. 原始值保留各自的数据窗口;计算变化需要 7 天内至少两次测量。

  1. 1
    活跃

    Mini Program Engineering Skill Suite is an Agent Skill suite for evidence‑first mini‑program development. It helps agents bring WeChat or other mini‑program projects from vague intent to reliable engineering work: project intake, product…

    github估算动量打开来源 ↗

    安装 git clone https://github.com/NocodeMrLi/mini-program-engineering-skill-suite ~/.claude/skills/mini-program-engineering-skill-suite

    39GitHub 星标
  2. 2

    技能
    活跃

    Real-time execution trace and cost intelligence for Claude Code

    github实测增长打开来源 ↗

    安装 /plugin marketplace add DeibyGS/claudestat

    34GitHub 星标+1 (+3.0 %)
  3. 3
    活跃

    Make Claude Opus 4.8 behave like Claude Fable 5 — doctrine output style, drift-catching hooks, and an eval loop against golden Fable transcripts. Claude Code plugin.

    github实测增长打开来源 ↗

    安装 /plugin marketplace add rennf93/opus-fable-playbook

    33GitHub 星标稳定
  4. 4
    休眠

    🔍 AI observability skill for Claude Code. Debug LangChain/LangGraph agents by fetching execution traces from LangSmith Studio directly in your terminal.

    github实测增长打开来源 ↗

    安装 git clone https://github.com/OthmanAdi/langsmith-fetch-skill ~/.claude/skills/langsmith-fetch-skill

    27GitHub 星标稳定
  5. 5

    智能体
    休眠

    Policy-enforced observability and fail-closed guardrails for MCP/A2A multi-agent systems.

    github实测增长打开来源 ↗

    26GitHub 星标稳定
  6. 6
    活跃

    Sentry instrumentation skill for system-behavior tracking

    github实测增长打开来源 ↗

    安装 git clone https://github.com/tortastudios/sentry-instrumentation ~/.claude/skills/sentry-instrumentation

    24GitHub 星标稳定
  7. 7
    活跃

    Deterministic, local-only audit reports for Claude Code AI agent sessions. Rust + SQLite. Zero network calls. MCP server for agents.

    github实测增长打开来源 ↗

    22GitHub 星标+2 (+10.0 %)
  8. 8
    活跃

    Log Claude Code sessions to Opik, the open-source LLM observability and evaluation platform, built by Comet. Tracing, evaluation, and skills for observable AI applications.

    github实测增长打开来源 ↗

    22GitHub 星标+1 (+4.8 %)
  9. 9

    其他
    活跃

    AI-detector auditing toolkit: measures how often a detector flags genuine human writing, how stable its verdict is across seeds, and whether that verdict survives meaning-preserving edits. Detector-in-the-loop measurement harness. Claude…

    github实测增长打开来源 ↗

    18GitHub 星标稳定
  10. 10

    技能
    活跃

    Measure prompt and skill improvements with blind A/B comparison.

    github实测增长打开来源 ↗

    安装 /plugin marketplace add shinpr/rashomon

    18GitHub 星标稳定
  11. 11
    活跃

    Two Claude Code skills for adversarial code review (single-model PAR and multi-model MMAR with cross-critique to catch hallucinations), plus a fixture-based eval suite.

    github实测增长打开来源 ↗

    安装 /plugin marketplace add prime-radiant-inc/parallel-adversarial-review

    17GitHub 星标稳定
  12. 12
    休眠

    MCP as a Judge: a behavioral MCP that strengthens AI coding assistants via explicit LLM evaluations

    mcp实测增长打开来源 ↗

    安装 claude mcp add mcp-as-a-judge -- uvx mcp-as-a-judge

    17GitHub 星标稳定
  13. 13
    活跃

    Make Claude write clearly, for everyone.

    github实测增长打开来源 ↗

    安装 /plugin marketplace add stefanobaghino/simple-output-styles

    16GitHub 星标稳定
  14. 14

    其他
    活跃

    A command-line tool to scaffold and manage enterprise-ready AI Agents powered by the A2A (Agent-to-Agent) protocol

    github实测增长打开来源 ↗

    14GitHub 星标+1 (+7.7 %)
  15. 15
    活跃

    Claude Code plugin marketplace — agentic-engineering (spec-driven shape→decide→execute→measure→eval with adversarial review) + github-keeper (audit/elevate READMEs and make a repo open-source-ready).

    github实测增长打开来源 ↗

    安装 /plugin marketplace add GiustoPiedimonte/agentic-engineering-marketplace

    13GitHub 星标稳定
  16. 16

    其他
    休眠

    A Claude Code skill that adds a rubric-based eval layer to any agent project. Framework-agnostic — generates rubric, test cases, judge prompt, and harness. Returns a weighted score plus a judge-leniency signal.

    github实测增长打开来源 ↗

    13GitHub 星标稳定
  17. 17
    活跃

    Query Spanlens LLM observability from Cursor, Claude Desktop, or Continue via MCP.

    mcp实测增长打开来源 ↗

    安装 claude mcp add mcp-server -- npx @spanlens/mcp-server

    12GitHub 星标稳定
  18. 18

    技能
    活跃

    Production-grade AI coding rules for Cursor and Claude Code. 15 rules + 9 doc templates + skills + agents + MCP setup. Drop into any project.

    github实测增长打开来源 ↗

    安装 git clone https://github.com/aiagentwithdhruv/ai-dev-stack ~/.claude/skills/ai-dev-stack

    10GitHub 星标+1 (+11.1 %)
  19. 19
    活跃

    Audit which Claude Code skills you actually use — surface dead installs and hallucinated invocations from your session logs.

    github实测增长打开来源 ↗

    安装 /plugin marketplace add sfrangulov/skill-graveyard

    10GitHub 星标+1 (+11.1 %)
  20. 20

    技能
    活跃

    Turn one decision into a judged tournament of solutions, then pick the best — a Claude Code skill that generates candidates, auto-derives the rubric, judges independently, and returns a defensible winner.

    github实测增长打开来源 ↗

    安装 /plugin marketplace add CoriChui/bakeoff

    10GitHub 星标稳定
  21. 21

    其他
    活跃

    A Go-native framework for LLM agents, with OpenTelemetry observability built in.

    github实测增长打开来源 ↗

    10GitHub 星标稳定
  22. 22
    休眠

    Local open-source dev tool to debug, secure, and evaluate LLM agents. Provides static analysis, dynamic security checks, and runtime monitoring - integrates with Cursor and Claude Code.

    github实测增长打开来源 ↗

    安装 /plugin marketplace add cylestio/agent-inspector

    9GitHub 星标稳定

学习与参考资源

Ranked by normalized popularity across sources. 这些资源可单独访问,不参与主要排名。

  1. 1
    活跃打开来源 ↗

    50+ curated LLM observability tools PLUS 26 Agent Skills (several with runnable, unit-tested scripts) to build, evaluate, debug, secure & monitor reliable LLM apps. Tracing, evals, guardrails, LLMOps.

    github

    安装 git clone https://github.com/ContextJet-ai/awesome-llm-observability ~/.claude/skills/awesome-llm-observability

    资源实测增长
    32GitHub 星标+2 (+6.7 %)
  2. 2
    活跃打开来源 ↗

    My complete journey to becoming an Agentic AI Engineer through structured learning, projects, experiments, and production-ready implementations of modern AI systems.

    github资源实测增长
    18GitHub 星标稳定
  3. 3
    活跃打开来源 ↗

    Documentation-discovery telemetry for Claude Code — heat/cold maps, health grade, evidence-backed router fixes. 100% local, zero tokens. /tt

    github

    安装 /plugin marketplace add Hedde/trigger_tree

    资源实测增长
    14GitHub 星标+1 (+7.7 %)