工具说明为英文。

LLM 可观测性

LLM 应用的监控、追踪、评估和质量管理。

用途

活跃度

排序方式

118

工具排名

Ranked by creation date, newest first; undated entries come last. 原始值保留各自的数据窗口;计算变化需要 7 天内至少两次测量。

  1. 1
    活跃

    Description Evidence-driven evaluation, benchmarking, and verified repair for Agent Skills across Codex, Claude Code, Gemini CLI, and Antigravity.

    github估算动量打开来源 ↗

    安装 git clone https://github.com/MaxLaurieHutchinson/skill-evaluation-graph ~/.claude/skills/skill-evaluation-graph

    2GitHub 星标
  2. 2
    活跃

    Run many AIs on one board and keep control of all of it. Deterministic code decides who acts — never a model. A privacy floor keeps sensitive work on your machine, your own tests decide what counts as done, and every action lands on a…

    github实测增长打开来源 ↗

    安装 git clone https://github.com/sandhusukhdeep2/sc-prism-releases ~/.claude/skills/sc-prism-releases

    1GitHub 星标稳定
  3. 3

    智能体
    活跃

    面向长程研发任务的 Java Agent Harness,支持 A2A 跨语言协作、可恢复执行、上下文工程与 Eval 驱动开发。

    github实测增长打开来源 ↗

    0GitHub 星标稳定
  4. 4

    技能
    活跃

    Deletion-first Agent Skill for removing test bloat, verification theater, and speculative fallbacks while preserving behavior.

    github实测增长打开来源 ↗

    安装 git clone https://github.com/MrZoyo/deslop-GPT ~/.claude/skills/deslop-GPT

    87GitHub 星标+59 (+210.7 %)
  5. 5

    智能体
    活跃

    Generic DeepSeek Harness (dsh) plugins: A2A protocol server, session storage mirror, and Langfuse observability.

    github实测增长打开来源 ↗

    1GitHub 星标稳定
  6. 6
    活跃

    Production reliability, benchmarking, and evaluation layer for MCP agent runtimes — protocol negotiation, A2A interop, circuit breakers, connection pooling, adaptive streaming, and autoscaling, with measured throughput/latency and a…

    github实测增长打开来源 ↗

    3GitHub 星标稳定
  7. 7
    活跃

    Multi-agent quality gate skill for Claude Code that researches, reviews, tests and challenges AI-generated work before the final answer.

    github实测增长打开来源 ↗

    安装 git clone https://github.com/ma-nucho-pro/supervisor-skill-claude ~/.claude/skills/supervisor-skill-claude

    2GitHub 星标稳定
  8. 8
    活跃

    Mini Program Engineering Skill Suite is an Agent Skill suite for evidence‑first mini‑program development. It helps agents bring WeChat or other mini‑program projects from vague intent to reliable engineering work: project intake, product…

    github实测增长打开来源 ↗

    安装 git clone https://github.com/NocodeMrLi/mini-program-engineering-skill-suite ~/.claude/skills/mini-program-engineering-skill-suite

    45GitHub 星标+6 (+15.4 %)
  9. 9
    活跃

    Deterministic, local-only audit reports for Claude Code AI agent sessions. Rust + SQLite. Zero network calls. MCP server for agents.

    github实测增长打开来源 ↗

    22GitHub 星标+2 (+10.0 %)
  10. 10
    活跃

    Amazon Bedrock AgentCore enterprise platform accelerator with AWS CDK, Terraform organization guardrails, MCP/A2A agents, security, memory, and observability.

    github实测增长打开来源 ↗

    6GitHub 星标+1 (+20.0 %)
  11. 11

    技能
    活跃

    Open-source infrastructure that turns scattered SKILL.md files into curated, retrieval-ready agent-skill corpora—with retrieval and evaluation tooling included.

    github实测增长打开来源 ↗

    安装 git clone https://github.com/EverMind-AI/SkillCorpus ~/.claude/skills/SkillCorpus

    525GitHub 星标+199 (+61.0 %)
  12. 12

    智能体
    活跃

    Local-first control plane for durable, governed agent teams: DAG orchestration, approvals, memory, signed skills, trace contracts, and realtime.

    github实测增长打开来源 ↗

    0GitHub 星标稳定
  13. 13

    技能
    活跃

    Production patterns for agent orchestrators, harnesses, evals, MCP/A2A interoperability, memory, provider routing, and LLM usage metering.

    github实测增长打开来源 ↗

    安装 git clone https://github.com/ssheleg/agent-stack ~/.claude/skills/agent-stack

    2GitHub 星标+1 (+100.0 %)
  14. 14
    活跃

    Your LLM reviewer said APPROVE. Did it? A structured verdict contract: prompt rule + parser + exit-code gate in one stdlib file, with the 42 counterexamples that forced every line. Free, MIT.

    github估算动量打开来源 ↗

    安装 git clone https://github.com/tonydzi/verdict-contract ~/.claude/skills/verdict-contract

    2GitHub 星标
  15. 15

    技能
    活跃

    Research-backed, eval-driven skills for AI agents

    github实测增长打开来源 ↗

    安装 git clone https://github.com/ZSeven-W/craft-skills ~/.claude/skills/craft-skills

    145GitHub 星标+9 (+6.6 %)
  16. 16
    活跃

    Make Claude write clearly, for everyone.

    github实测增长打开来源 ↗

    安装 /plugin marketplace add stefanobaghino/simple-output-styles

    16GitHub 星标稳定
  17. 17
    活跃

    Skill instalable para Claude Code que mide si tu agente cumple las reglas establecidas. Lee el historial de sesiones de tu disco y devuelve un porcentaje por regla de cumplimiento. Sin instalar nada, sin red, biblioteca estándar.

    github实测增长打开来源 ↗

    安装 /plugin marketplace add jleonceo/skill-adherencia-reglas

    1GitHub 星标稳定
  18. 18
    活跃

    Mide si tu agente de IA cumple las reglas que le escribiste. Lee el historial local de Claude Code y devuelve un porcentaje por regla. Sin instalar nada, sin red, biblioteca estándar.

    github实测增长打开来源 ↗

    安装 git clone https://github.com/jleonceo/adherencia-reglas ~/.claude/skills/adherencia-reglas

    1GitHub 星标稳定
  19. 19

    其他
    活跃

    Evidence-first infrastructure for reproducible, isolated executable evaluation and online rewards.

    github实测增长打开来源 ↗

    197GitHub 星标+81 (+69.8 %)
  20. 20
    活跃

    Claude Code skills where every entry ships receipts — accuracy-gated benchmarks against baseline and placebo, rejects published

    github实测增长打开来源 ↗

    安装 git clone https://github.com/sjh9714/skill-receipts ~/.claude/skills/skill-receipts

    2GitHub 星标稳定
  21. 21

    智能体
    活跃

    Deterministic crash tests for agent control planes: paired worlds, hard stops, budgets, scopes, approvals, and hash-linked receipts.

    github实测增长打开来源 ↗

    1GitHub 星标稳定
  22. 22
    活跃

    🐙 ADLC Team Skills — Agentic SDLC for Engineering Teams

    github实测增长打开来源 ↗

    安装 git clone https://github.com/tikalk/adlc-team-skills ~/.claude/skills/adlc-team-skills

    132GitHub 星标+1 (+0.76 %)
  23. 23

    MCP
    活跃

    A typed, replayable, budget-aware programming language for LLM agents. Compiles to Python & TypeScript.

    github实测增长打开来源 ↗

    2GitHub 星标稳定

学习与参考资源

Ranked by creation date, newest first; undated entries come last. 这些资源可单独访问,不参与主要排名。

  1. 1
    活跃打开来源 ↗

    My complete journey to becoming an Agentic AI Engineer through structured learning, projects, experiments, and production-ready implementations of modern AI systems.

    github资源实测增长
    18GitHub 星标稳定
  2. 2
    活跃打开来源 ↗

    Documentation-discovery telemetry for Claude Code — heat/cold maps, health grade, evidence-backed router fixes. 100% local, zero tokens. /tt

    github

    安装 /plugin marketplace add Hedde/trigger_tree

    资源实测增长
    14GitHub 星标稳定