工具说明为英文。

LLM 可观测性

LLM 应用的监控、追踪、评估和质量管理。

用途

活跃度

排序方式

121

工具排名

按实测增长排序,并在不同来源间归一化。 原始值保留各自的数据窗口;计算变化需要 7 天内至少两次测量。

  1. 1

    智能体
    活跃

    A lightweight Python library that decouples agentic runtime from applications it builds

    github实测增长打开来源 ↗

    1GitHub 星标稳定
  2. 2
    活跃

    Multi-agent quality gate skill for Claude Code that researches, reviews, tests and challenges AI-generated work before the final answer.

    github实测增长打开来源 ↗

    安装 git clone https://github.com/ma-nucho-pro/supervisor-skill-claude ~/.claude/skills/supervisor-skill-claude

    2GitHub 星标稳定
  3. 3
    活跃

    Log Claude Code sessions to Opik, the open-source LLM observability and evaluation platform, built by Comet. Tracing, evaluation, and skills for observable AI applications.

    github实测增长打开来源 ↗

    22GitHub 星标稳定
  4. 4

    技能
    活跃

    Turn one decision into a judged tournament of solutions, then pick the best — a Claude Code skill that generates candidates, auto-derives the rubric, judges independently, and returns a defensible winner.

    github实测增长打开来源 ↗

    安装 /plugin marketplace add CoriChui/bakeoff

    10GitHub 星标稳定
  5. 5
    活跃

    Audit which Claude Code skills you actually use — surface dead installs and hallucinated invocations from your session logs.

    github实测增长打开来源 ↗

    安装 /plugin marketplace add sfrangulov/skill-graveyard

    10GitHub 星标稳定
  6. 6

    技能
    活跃

    Production patterns for agent orchestrators, harnesses, evals, MCP/A2A interoperability, memory, provider routing, and LLM usage metering.

    github实测增长打开来源 ↗

    安装 git clone https://github.com/ssheleg/agent-stack ~/.claude/skills/agent-stack

    2GitHub 星标稳定
  7. 7

    技能
    活跃

    Axiom is a curated marketplace of shared plugins for Claude Code and Codex.

    github实测增长打开来源 ↗

    安装 git clone https://github.com/netopsengineer/axiom ~/.claude/skills/axiom

    5GitHub 星标稳定
  8. 8

    智能体
    活跃

    Generic DeepSeek Harness (dsh) plugins: A2A protocol server, session storage mirror, and Langfuse observability.

    github实测增长打开来源 ↗

    1GitHub 星标稳定
  9. 9
    活跃

    Production reliability, benchmarking, and evaluation layer for MCP agent runtimes — protocol negotiation, A2A interop, circuit breakers, connection pooling, adaptive streaming, and autoscaling, with measured throughput/latency and a…

    github实测增长打开来源 ↗

    3GitHub 星标稳定
  10. 10

    技能
    活跃

    Measure prompt and skill improvements with blind A/B comparison.

    github实测增长打开来源 ↗

    安装 /plugin marketplace add shinpr/rashomon

    18GitHub 星标稳定
  11. 11
    活跃

    A Model Context Protocol (MCP) server for Langfuse, enabling AI agents to query Langfuse trace data for enhanced debugging and observability

    github实测增长打开来源 ↗

    105GitHub 星标稳定
  12. 12
    活跃

    Make Claude write clearly, for everyone.

    github实测增长打开来源 ↗

    安装 /plugin marketplace add stefanobaghino/simple-output-styles

    16GitHub 星标稳定
  13. 13

    智能体
    休眠

    Real-time debugging proxy for Agent2Agent (A2A) multi-agent systems

    github实测增长打开来源 ↗

    3GitHub 星标稳定
  14. 14

    技能
    活跃

    Glanceable Claude Code and Codex state in your terminal tabs: white=idle, blue=working, orange=waiting. Multi-terminal (iTerm2, WezTerm, AI Power Term), a live status bar with Anthropic usage limits, /sfl and /nil window save-and-restore,…

    github实测增长打开来源 ↗

    安装 git clone https://github.com/wasulajr/headsup ~/.claude/skills/headsup

    1GitHub 星标稳定
  15. 15
    活跃

    Run many AIs on one board and keep control of all of it. Deterministic code decides who acts — never a model. A privacy floor keeps sensitive work on your machine, your own tests decide what counts as done, and every action lands on a…

    github实测增长打开来源 ↗

    安装 git clone https://github.com/sandhusukhdeep2/sc-prism-releases ~/.claude/skills/sc-prism-releases

    1GitHub 星标稳定
  16. 16
    活跃

    Your LLM reviewer said APPROVE. Did it? A structured verdict contract: prompt rule + parser + exit-code gate in one stdlib file, with the 42 counterexamples that forced every line. Free, MIT.

    github实测增长打开来源 ↗

    安装 git clone https://github.com/tonydzi/verdict-contract ~/.claude/skills/verdict-contract

    2GitHub 星标稳定
  17. 17
    活跃

    Description Evidence-driven evaluation, benchmarking, and verified repair for Agent Skills across Codex, Claude Code, Gemini CLI, and Antigravity.

    github实测增长打开来源 ↗

    安装 git clone https://github.com/MaxLaurieHutchinson/skill-evaluation-graph ~/.claude/skills/skill-evaluation-graph

    2GitHub 星标稳定

学习与参考资源

混合排名:实测增长优先。 这些资源可单独访问,不参与主要排名。

  1. 1
    活跃打开来源 ↗

    Handbook técnico aberto sobre engenharia de IA em produção: ML tradicional, LLMs, RAG, agentes, segurança, observabilidade, FinOps e deployment.

    github资源实测增长
    1GitHub 星标稳定
  2. 2
    活跃打开来源 ↗

    Documentation-discovery telemetry for Claude Code — heat/cold maps, health grade, evidence-backed router fixes. 100% local, zero tokens. /tt

    github

    安装 /plugin marketplace add Hedde/trigger_tree

    资源实测增长
    14GitHub 星标稳定
  3. 3
    活跃打开来源 ↗

    My complete journey to becoming an Agentic AI Engineer through structured learning, projects, experiments, and production-ready implementations of modern AI systems.

    github资源实测增长
    18GitHub 星标稳定
  4. 4

    tunelab

    技能
    活跃打开来源 ↗

    Claude Code plugin for LLM fine-tuning, distillation, and evaluation — decide whether you need fine-tuning at all, distill your LLM logs into small local models (MLX/LoRA), evaluate with held-out discipline, and learn the why at every step.

    github

    安装 /plugin marketplace add rchaz/tunelab

    资源估算动量
    6GitHub 星标·