工具说明为英文。

LLM 可观测性

LLM 应用的监控、追踪、评估和质量管理。

用途

活跃度

排序方式

122

工具排名

按实测增长排序,并在不同来源间归一化。 原始值保留各自的数据窗口;计算变化需要 7 天内至少两次测量。

  1. 1

    技能
    活跃

    A skill creator that proves its skills work. Evidence-driven skill creation for Claude Code and Codex: baseline-tested generation, per-skill regression evals, ecosystem doctor, cross-runtime compile, and an opt-in proactive advisor.

    github实测增长打开来源 ↗

    安装 git clone https://github.com/tripleyak/SkillForge ~/.claude/skills/SkillForge

    884GitHub 星标-1 (-0.11 %)
  2. 2
    活跃

    🟪 Open-source runtime that ships any LangGraph or Google ADK agent as a production-ready FastAPI service. Bundled , AG-UI copilotkit API, chat UI, 15+ guardrails, MCP, OpenTelemetry, OIDC. One pip install. Self-hosted, no vendor lock-in.

    github实测增长打开来源 ↗

    安装 git clone https://github.com/Idun-Group/idun-agent-platform ~/.claude/skills/idun-agent-platform

    199GitHub 星标稳定
  3. 3

    其他
    活跃

    Self-hosted AI SRE for Kubernetes — zero-instrumentation eBPF observability plus a copilot that fixes issues through guardrailed, self-verifying actions. BYO-LLM, air-gapped capable.

    github实测增长打开来源 ↗

    131GitHub 星标-4 (-3.0 %)
  4. 4
    活跃

    74 open-source Agent Skills for Claude Code and Codex: AI SEO, AEO and GEO, code review with an A-F ship grade, CI gates, AI evals, design systems, conversion copy, Instagram growth, iOS and Android app shipping, creator rights, and…

    github实测增长打开来源 ↗

    129GitHub 星标-4 (-3.0 %)
  5. 5
    活跃

    A Model Context Protocol (MCP) server for Langfuse, enabling AI agents to query Langfuse trace data for enhanced debugging and observability

    github实测增长打开来源 ↗

    105GitHub 星标稳定
  6. 6

    技能
    休眠

    Don't make LLMs honest. Make every factual claim auditable. — An LLM Claim Auditing Layer with T1-T7 truth gradients. 98.1% business effectiveness on LiarBench v0.2.

    github实测增长打开来源 ↗

    安装 git clone https://github.com/lc198707/anti-lie ~/.claude/skills/anti-lie

    89GitHub 星标稳定
  7. 7

    其他
    活跃

    AgentX python SDK. Build multi-agent AI workforce. Run evaluation. Trace your agent. Full Observability.

    github实测增长打开来源 ↗

    68GitHub 星标稳定
  8. 8
    活跃

    Open benchmark for Claude Code SEO skills — real headless execution against fixture sites with planted-defect answer keys. Deterministic scoring, pre-registered rubric.

    github实测增长打开来源 ↗

    安装 git clone https://github.com/aleclindz/seo-skill-bench ~/.claude/skills/seo-skill-bench

    51GitHub 星标-5 (-8.9 %)
  9. 9

    技能
    活跃

    Agent skills for Arize — datasets, experiments, and traces via the ax CLI

    github实测增长打开来源 ↗

    安装 git clone https://github.com/Arize-ai/arize-skills ~/.claude/skills/arize-skills

    47GitHub 星标稳定
  10. 10

    技能
    活跃

    Real-time execution trace and cost intelligence for Claude Code

    github实测增长打开来源 ↗

    安装 /plugin marketplace add DeibyGS/claudestat

    34GitHub 星标稳定
  11. 11
    休眠

    🔍 AI observability skill for Claude Code. Debug LangChain/LangGraph agents by fetching execution traces from LangSmith Studio directly in your terminal.

    github实测增长打开来源 ↗

    安装 git clone https://github.com/OthmanAdi/langsmith-fetch-skill ~/.claude/skills/langsmith-fetch-skill

    27GitHub 星标稳定
  12. 12

    智能体
    休眠

    Policy-enforced observability and fail-closed guardrails for MCP/A2A multi-agent systems.

    github实测增长打开来源 ↗

    26GitHub 星标稳定
  13. 13
    活跃

    Sentry instrumentation skill for system-behavior tracking

    github实测增长打开来源 ↗

    安装 git clone https://github.com/tortastudios/sentry-instrumentation ~/.claude/skills/sentry-instrumentation

    24GitHub 星标稳定
  14. 14

    其他
    活跃

    AI-detector auditing toolkit: measures how often a detector flags genuine human writing, how stable its verdict is across seeds, and whether that verdict survives meaning-preserving edits. Detector-in-the-loop measurement harness. Claude…

    github实测增长打开来源 ↗

    18GitHub 星标稳定
  15. 15

    技能
    活跃

    Measure prompt and skill improvements with blind A/B comparison.

    github实测增长打开来源 ↗

    安装 /plugin marketplace add shinpr/rashomon

    18GitHub 星标稳定
  16. 16
    活跃

    Two Claude Code skills for adversarial code review (single-model PAR and multi-model MMAR with cross-critique to catch hallucinations), plus a fixture-based eval suite.

    github实测增长打开来源 ↗

    安装 /plugin marketplace add prime-radiant-inc/parallel-adversarial-review

    17GitHub 星标稳定
  17. 17
    休眠

    MCP as a Judge: a behavioral MCP that strengthens AI coding assistants via explicit LLM evaluations

    mcp实测增长打开来源 ↗

    安装 claude mcp add mcp-as-a-judge -- uvx mcp-as-a-judge

    17GitHub 星标稳定
  18. 18
    活跃

    Make Claude write clearly, for everyone.

    github实测增长打开来源 ↗

    安装 /plugin marketplace add stefanobaghino/simple-output-styles

    16GitHub 星标稳定
  19. 19
    活跃

    Claude Code plugin marketplace — agentic-engineering (spec-driven shape→decide→execute→measure→eval with adversarial review) + github-keeper (audit/elevate READMEs and make a repo open-source-ready).

    github实测增长打开来源 ↗

    安装 /plugin marketplace add GiustoPiedimonte/agentic-engineering-marketplace

    13GitHub 星标稳定
  20. 20

    其他
    休眠

    A Claude Code skill that adds a rubric-based eval layer to any agent project. Framework-agnostic — generates rubric, test cases, judge prompt, and harness. Returns a weighted score plus a judge-leniency signal.

    github实测增长打开来源 ↗

    13GitHub 星标稳定
  21. 21
    活跃

    Query Spanlens LLM observability from Cursor, Claude Desktop, or Continue via MCP.

    mcp实测增长打开来源 ↗

    安装 claude mcp add mcp-server -- npx @spanlens/mcp-server

    12GitHub 星标稳定
  22. 22

    技能
    活跃

    Turn one decision into a judged tournament of solutions, then pick the best — a Claude Code skill that generates candidates, auto-derives the rubric, judges independently, and returns a defensible winner.

    github实测增长打开来源 ↗

    安装 /plugin marketplace add CoriChui/bakeoff

    10GitHub 星标稳定
  23. 23

    其他
    活跃

    A Go-native framework for LLM agents, with OpenTelemetry observability built in.

    github实测增长打开来源 ↗

    10GitHub 星标稳定

学习与参考资源

按实测增长排序,并在不同来源间归一化。 这些资源可单独访问,不参与主要排名。

  1. 1
    活跃打开来源 ↗

    Curated + scored map of official MCP servers and agents for DevOps, Cloud, SRE, and Platform Engineering — every entry rated on production access, approval gates, and audit evidence.

    github资源实测增长
    77GitHub 星标稳定
  2. 2
    活跃打开来源 ↗

    My complete journey to becoming an Agentic AI Engineer through structured learning, projects, experiments, and production-ready implementations of modern AI systems.

    github资源实测增长
    18GitHub 星标稳定