工具说明为英文。

LLM 可观测性

LLM 应用的监控、追踪、评估和质量管理。

用途

活跃度

排序方式

121

工具排名

按实测增长排序,并在不同来源间归一化。 原始值保留各自的数据窗口;计算变化需要 7 天内至少两次测量。

  1. 1

    智能体
    休眠

    Real-time debugging proxy for Agent2Agent (A2A) multi-agent systems

    github实测增长打开来源 ↗

    3GitHub 星标稳定
  2. 2

    技能
    活跃

    Axiom is a curated marketplace of shared plugins for Claude Code and Codex.

    github实测增长打开来源 ↗

    安装 git clone https://github.com/netopsengineer/axiom ~/.claude/skills/axiom

    5GitHub 星标稳定
  3. 3
    活跃

    Claude Code plugin marketplace — agentic-engineering (spec-driven shape→decide→execute→measure→eval with adversarial review) + github-keeper (audit/elevate READMEs and make a repo open-source-ready).

    github实测增长打开来源 ↗

    安装 /plugin marketplace add GiustoPiedimonte/agentic-engineering-marketplace

    13GitHub 星标稳定
  4. 4

    技能
    活跃

    The container platform tailored for Kubernetes multi-cloud, datacenter, and edge management ⎈ 🖥 ☁️

    github实测增长打开来源 ↗

    安装 git clone https://github.com/kubesphere/kubesphere ~/.claude/skills/kubesphere

    17 035GitHub 星标+2 (+0.01 %)
  5. 5

    MCP
    休眠

    Open-source Python framework to deploy AI agents via HTTP, A2A, and MCP with built-in observability

    github实测增长打开来源 ↗

    5GitHub 星标稳定
  6. 6

    其他
    休眠

    A Claude Code skill that adds a rubric-based eval layer to any agent project. Framework-agnostic — generates rubric, test cases, judge prompt, and harness. Returns a weighted score plus a judge-leniency signal.

    github实测增长打开来源 ↗

    13GitHub 星标稳定
  7. 7
    活跃

    Amazon Bedrock AgentCore enterprise platform accelerator with AWS CDK, Terraform organization guardrails, MCP/A2A agents, security, memory, and observability.

    github实测增长打开来源 ↗

    6GitHub 星标+2 (+50.0 %)
  8. 8

    智能体
    活跃

    Provide clear documentation for AgentStack’s MCP protocol, plugins, and ecosystem API with usage examples and tool references.

    github实测增长打开来源 ↗

    0GitHub 星标稳定
  9. 9

    智能体
    活跃

    MemroOS / memroos: memory OS and governance layer for AI agents, agent workflows, dispatch, proof, and context continuity.

    github实测增长打开来源 ↗

    7GitHub 星标稳定
  10. 10

    技能
    活跃

    Production-grade AI coding rules for Cursor and Claude Code. 15 rules + 9 doc templates + skills + agents + MCP setup. Drop into any project.

    github实测增长打开来源 ↗

    安装 git clone https://github.com/aiagentwithdhruv/ai-dev-stack ~/.claude/skills/ai-dev-stack

    10GitHub 星标+1 (+11.1 %)
  11. 11
    活跃

    Multi-tier framework for evaluating AI agent skills with quality gates, semantic overlap detection, synthetic evaluation dataset generation, and live agent evaluation that measures how skills affect agent behavior.

    github实测增长打开来源 ↗

    安装 git clone https://github.com/NVIDIA/SkillEvaluator ~/.claude/skills/SkillEvaluator

    399GitHub 星标+51 (+14.7 %)
  12. 12
    休眠

    A self-improving harness router for Claude Code.

    github实测增长打开来源 ↗

    安装 /plugin marketplace add SeongwoongCho/adaptive-harness

    8GitHub 星标稳定
  13. 13
    活跃

    Audit which Claude Code skills you actually use — surface dead installs and hallucinated invocations from your session logs.

    github实测增长打开来源 ↗

    安装 /plugin marketplace add sfrangulov/skill-graveyard

    10GitHub 星标+1 (+11.1 %)
  14. 14

    技能
    活跃

    Anchor — the production-grade AGENTS.md template for AI coding agents. 51 sections of battle-tested rules. Keep your agents grounded.

    github实测增长打开来源 ↗

    安装 git clone https://github.com/peva3/anchor ~/.claude/skills/anchor

    8GitHub 星标稳定
  15. 15

    技能
    活跃

    Turn one decision into a judged tournament of solutions, then pick the best — a Claude Code skill that generates candidates, auto-derives the rubric, judges independently, and returns a defensible winner.

    github实测增长打开来源 ↗

    安装 /plugin marketplace add CoriChui/bakeoff

    10GitHub 星标稳定
  16. 16
    休眠

    Local open-source dev tool to debug, secure, and evaluate LLM agents. Provides static analysis, dynamic security checks, and runtime monitoring - integrates with Cursor and Claude Code.

    github实测增长打开来源 ↗

    安装 /plugin marketplace add cylestio/agent-inspector

    9GitHub 星标稳定
  17. 17

    其他
    活跃

    A Go-native framework for LLM agents, with OpenTelemetry observability built in.

    github实测增长打开来源 ↗

    10GitHub 星标稳定
  18. 18
    活跃

    Architecture-first Python scaffold for an auditable multi-agent corporate credit desk using A2A, MCP, deterministic credit policies, model routing, and OpenTelemetry.

    github实测增长打开来源 ↗

    0GitHub 星标稳定
  19. 19
    活跃

    Dashboard for monitoring claude code sessions.

    github实测增长打开来源 ↗

    安装 /plugin marketplace add JayantDevkar/claude-code-karma

    323GitHub 星标+2 (+0.62 %)
  20. 20

    智能体
    休眠

    Contract-based testing for LLM agents: hybrid evaluation, multi-agent + A2A, deterministic replay.

    github实测增长打开来源 ↗

    0GitHub 星标稳定
  21. 21

    MCP
    活跃

    AURA is a production-tested SRE agent platform you can deploy in minutes. AURA handles the guardrails, APIs, state management, streaming, and failure handling required to put AI to work safely on production infrastructure.

    github实测增长打开来源 ↗

    291GitHub 星标+33 (+12.8 %)
  22. 22

    技能
    活跃

    Marketing Mix Model expert agent plugin for Claude Code - pymc-marketing v0.18.2+

    github实测增长打开来源 ↗

    安装 /plugin marketplace add Yakoub-ai/agent-mmm

    4GitHub 星标稳定
  23. 23
    休眠

    A Python proof-of-concept for tracing multi-turn Agent-to-Agent (A2A) conversations as a single unified MLflow trace for LLM observability and evaluation.

    github实测增长打开来源 ↗

    0GitHub 星标稳定
  24. 24

    技能
    活跃

    🪢 Langfuse documentation -- Langfuse is the open source LLM Engineering Platform. Observability, evals, prompt management, playground and metrics to debug and improve LLM apps

    github实测增长打开来源 ↗

    安装 git clone https://github.com/langfuse/langfuse-docs ~/.claude/skills/langfuse-docs

    239GitHub 星标+3 (+1.3 %)
  25. 25

    智能体
    活跃

    A2A green-agent orchestrator for evaluating agents on the AppWorld benchmark, built on the AgentBeats SDK

    github实测增长打开来源 ↗

    0GitHub 星标稳定