工具说明为英文。

LLM 可观测性

LLM 应用的监控、追踪、评估和质量管理。

用途

活跃度

排序方式

117

工具排名

按实测增长排序,并在不同来源间归一化。 原始值保留各自的数据窗口;计算变化需要 7 天内至少两次测量。

  1. 1

    MCP
    活跃

    AURA is a production-tested SRE agent platform you can deploy in minutes. AURA handles the guardrails, APIs, state management, streaming, and failure handling required to put AI to work safely on production infrastructure.

    github实测增长打开来源 ↗

    254GitHub 星标-3 (-1.2 %)
  2. 2
    活跃

    🟪 Open-source runtime that ships any LangGraph or Google ADK agent as a production-ready FastAPI service. Bundled , AG-UI copilotkit API, chat UI, 15+ guardrails, MCP, OpenTelemetry, OIDC. One pip install. Self-hosted, no vendor lock-in.

    github实测增长打开来源 ↗

    安装 git clone https://github.com/Idun-Group/idun-agent-platform ~/.claude/skills/idun-agent-platform

    198GitHub 星标-1 (-0.50 %)
  3. 3

    其他
    活跃

    Self-hosted AI SRE for Kubernetes — zero-instrumentation eBPF observability plus a copilot that fixes issues through guardrailed, self-verifying actions. BYO-LLM, air-gapped capable.

    github实测增长打开来源 ↗

    131GitHub 星标-4 (-3.0 %)
  4. 4
    活跃

    74 open-source Agent Skills for Claude Code and Codex: AI SEO, AEO and GEO, code review with an A-F ship grade, CI gates, AI evals, design systems, conversion copy, Instagram growth, iOS and Android app shipping, creator rights, and…

    github实测增长打开来源 ↗

    129GitHub 星标-4 (-3.0 %)
  5. 5
    活跃

    A Model Context Protocol (MCP) server for Langfuse, enabling AI agents to query Langfuse trace data for enhanced debugging and observability

    github实测增长打开来源 ↗

    105GitHub 星标稳定
  6. 6

    技能
    休眠

    Don't make LLMs honest. Make every factual claim auditable. — An LLM Claim Auditing Layer with T1-T7 truth gradients. 98.1% business effectiveness on LiarBench v0.2.

    github实测增长打开来源 ↗

    安装 git clone https://github.com/lc198707/anti-lie ~/.claude/skills/anti-lie

    89GitHub 星标稳定
  7. 7
    活跃

    MCP server for Langfuse LLM observability — trace and observation analysis.

    mcp实测增长打开来源 ↗

    安装 claude mcp add langfuse -- npx langfuse-observability-mcp-server

    78GitHub 星标稳定
  8. 8

    其他
    活跃

    AgentX python SDK. Build multi-agent AI workforce. Run evaluation. Trace your agent. Full Observability.

    github实测增长打开来源 ↗

    68GitHub 星标稳定
  9. 9
    活跃

    Open benchmark for Claude Code SEO skills — real headless execution against fixture sites with planted-defect answer keys. Deterministic scoring, pre-registered rubric.

    github实测增长打开来源 ↗

    安装 git clone https://github.com/aleclindz/seo-skill-bench ~/.claude/skills/seo-skill-bench

    51GitHub 星标-5 (-8.9 %)
  10. 10

    技能
    活跃

    Agent skills for Arize — datasets, experiments, and traces via the ax CLI

    github实测增长打开来源 ↗

    安装 git clone https://github.com/Arize-ai/arize-skills ~/.claude/skills/arize-skills

    47GitHub 星标稳定
  11. 11
    活跃

    Make Claude Opus 4.8 behave like Claude Fable 5 — doctrine output style, drift-catching hooks, and an eval loop against golden Fable transcripts. Claude Code plugin.

    github实测增长打开来源 ↗

    安装 /plugin marketplace add rennf93/opus-fable-playbook

    33GitHub 星标稳定
  12. 12
    休眠

    🔍 AI observability skill for Claude Code. Debug LangChain/LangGraph agents by fetching execution traces from LangSmith Studio directly in your terminal.

    github实测增长打开来源 ↗

    安装 git clone https://github.com/OthmanAdi/langsmith-fetch-skill ~/.claude/skills/langsmith-fetch-skill

    27GitHub 星标稳定
  13. 13

    智能体
    休眠

    Policy-enforced observability and fail-closed guardrails for MCP/A2A multi-agent systems.

    github实测增长打开来源 ↗

    26GitHub 星标稳定
  14. 14
    活跃

    Sentry instrumentation skill for system-behavior tracking

    github实测增长打开来源 ↗

    安装 git clone https://github.com/tortastudios/sentry-instrumentation ~/.claude/skills/sentry-instrumentation

    24GitHub 星标稳定
  15. 15

    其他
    活跃

    AI-detector auditing toolkit: measures how often a detector flags genuine human writing, how stable its verdict is across seeds, and whether that verdict survives meaning-preserving edits. Detector-in-the-loop measurement harness. Claude…

    github实测增长打开来源 ↗

    18GitHub 星标稳定
  16. 16

    技能
    活跃

    Measure prompt and skill improvements with blind A/B comparison.

    github实测增长打开来源 ↗

    安装 /plugin marketplace add shinpr/rashomon

    18GitHub 星标稳定
  17. 17
    活跃

    Two Claude Code skills for adversarial code review (single-model PAR and multi-model MMAR with cross-critique to catch hallucinations), plus a fixture-based eval suite.

    github实测增长打开来源 ↗

    安装 /plugin marketplace add prime-radiant-inc/parallel-adversarial-review

    17GitHub 星标稳定
  18. 18
    活跃

    Make Claude write clearly, for everyone.

    github实测增长打开来源 ↗

    安装 /plugin marketplace add stefanobaghino/simple-output-styles

    16GitHub 星标稳定
  19. 19
    活跃

    Claude Code plugin marketplace — agentic-engineering (spec-driven shape→decide→execute→measure→eval with adversarial review) + github-keeper (audit/elevate READMEs and make a repo open-source-ready).

    github实测增长打开来源 ↗

    安装 /plugin marketplace add GiustoPiedimonte/agentic-engineering-marketplace

    13GitHub 星标稳定
  20. 20

    其他
    休眠

    A Claude Code skill that adds a rubric-based eval layer to any agent project. Framework-agnostic — generates rubric, test cases, judge prompt, and harness. Returns a weighted score plus a judge-leniency signal.

    github实测增长打开来源 ↗

    13GitHub 星标稳定
  21. 21
    活跃

    Query Spanlens LLM observability from Cursor, Claude Desktop, or Continue via MCP.

    mcp实测增长打开来源 ↗

    安装 claude mcp add mcp-server -- npx @spanlens/mcp-server

    12GitHub 星标稳定
  22. 22

    技能
    活跃

    Turn one decision into a judged tournament of solutions, then pick the best — a Claude Code skill that generates candidates, auto-derives the rubric, judges independently, and returns a defensible winner.

    github实测增长打开来源 ↗

    安装 /plugin marketplace add CoriChui/bakeoff

    10GitHub 星标稳定
  23. 23

    其他
    活跃

    A Go-native framework for LLM agents, with OpenTelemetry observability built in.

    github实测增长打开来源 ↗

    10GitHub 星标稳定

学习与参考资源

按实测增长排序,并在不同来源间归一化。 这些资源可单独访问,不参与主要排名。

  1. 1
    活跃打开来源 ↗

    Curated + scored map of official MCP servers and agents for DevOps, Cloud, SRE, and Platform Engineering — every entry rated on production access, approval gates, and audit evidence.

    github资源实测增长
    77GitHub 星标稳定
  2. 2
    活跃打开来源 ↗

    My complete journey to becoming an Agentic AI Engineer through structured learning, projects, experiments, and production-ready implementations of modern AI systems.

    github资源实测增长
    18GitHub 星标稳定