工具说明为英文。
LLM 可观测性
LLM 应用的监控、追踪、评估和质量管理。
用途
活跃度
排序方式
122
工具排名
按实测增长排序,并在不同来源间归一化。 原始值保留各自的数据窗口;计算变化需要 7 天内至少两次测量。
- 1活跃
Multi-tier framework for evaluating AI agent skills with quality gates, semantic overlap detection, synthetic evaluation dataset generation, and live agent evaluation that measures how skills affect agent behavior.
github实测增长打开来源 ↗
安装
git clone https://github.com/NVIDIA/SkillEvaluator ~/.claude/skills/SkillEvaluator404GitHub 星标+53 (+15.1 %) - 2活跃
Multi-agent quality gate skill for Claude Code that researches, reviews, tests and challenges AI-generated work before the final answer.
github实测增长打开来源 ↗
安装
git clone https://github.com/ma-nucho-pro/supervisor-skill-claude ~/.claude/skills/supervisor-skill-claude2GitHub 星标稳定 - 3活跃
historian
智能体A local A2A event and continuity service for agent applications, with structured history, provenance-safe transcripts, literal search, and natural-language queries.
github实测增长打开来源 ↗
1GitHub 星标稳定 - 4活跃
Local-first Agentic AI Infrastructure Platform with LLM Gateway, Agent DAG Runtime, MCP Tool Hub, A2A Agent Mesh, Local RAG, Tool Sandbox and Observability.
github实测增长打开来源 ↗
1GitHub 星标稳定 - 5活跃
sre-on-call
智能体Multi-agent SRE on-call investigator that auto-triages Slack/Discord infrastructure alerts via AWS Bedrock AgentCore, fanning out to specialized agents (CloudWatch, EKS, Slack/Discord scanners) for parallel investigation.
github实测增长打开来源 ↗
3GitHub 星标稳定 - 6活跃
AI Guardian
MCPGoverned local-LLM observability: model policy, prompt scanner, capture proxy, 21 tools.
mcp实测增长打开来源 ↗
安装
claude mcp add ai-guardian -- uvx ai-guardian-aiops0GitHub 星标稳定 - 7活跃
pyxen
智能体A lightweight Python library that decouples agentic runtime from applications it builds
github实测增长打开来源 ↗
1GitHub 星标稳定 - 8活跃
Amazon Bedrock AgentCore enterprise platform accelerator with AWS CDK, Terraform organization guardrails, MCP/A2A agents, security, memory, and observability.
github实测增长打开来源 ↗
6GitHub 星标+1 (+20.0 %) - 9活跃
Claude Code plugin marketplace — agentic-engineering (spec-driven shape→decide→execute→measure→eval with adversarial review) + github-keeper (audit/elevate READMEs and make a repo open-source-ready).
github实测增长打开来源 ↗
安装
/plugin marketplace add GiustoPiedimonte/agentic-engineering-marketplace13GitHub 星标稳定 - 10104GitHub 星标稳定
- 11活跃
rashomon
技能Measure prompt and skill improvements with blind A/B comparison.
github实测增长打开来源 ↗
安装
/plugin marketplace add shinpr/rashomon18GitHub 星标稳定 - 12活跃
Terminal-first validation harness for reviewing PRDs, RFCs, strategy docs, and experiment plans.
github实测增长打开来源 ↗
1GitHub 星标稳定 - 13休眠
primal-core
智能体The reliability and interoperability layer for AI agents. A2A-native. MCP-bridged.
github实测增长打开来源 ↗
1GitHub 星标稳定 - 14活跃
nudge
MCPA typed, replayable, budget-aware programming language for LLM agents. Compiles to Python & TypeScript.
github实测增长打开来源 ↗
2GitHub 星标稳定 - 15活跃
Research-backed, eval-driven skills for AI agents
github实测增长打开来源 ↗
安装
git clone https://github.com/ZSeven-W/craft-skills ~/.claude/skills/craft-skills145GitHub 星标+9 (+6.6 %) - 16活跃
🪢 Langfuse documentation -- Langfuse is the open source LLM Engineering Platform. Observability, evals, prompt management, playground and metrics to debug and improve LLM apps
github实测增长打开来源 ↗
安装
git clone https://github.com/langfuse/langfuse-docs ~/.claude/skills/langfuse-docs240GitHub 星标+4 (+1.7 %) - 17休眠
A Claude Code skill that adds a rubric-based eval layer to any agent project. Framework-agnostic — generates rubric, test cases, judge prompt, and harness. Returns a weighted score plus a judge-leniency signal.
github实测增长打开来源 ↗
13GitHub 星标稳定 - 18活跃
evoagent-os
智能体Local-first control plane for durable, governed agent teams: DAG orchestration, approvals, memory, signed skills, trace contracts, and realtime.
github实测增长打开来源 ↗
0GitHub 星标稳定 - 19活跃
Skills, prompts, and instructions for building AI agents on top of Dynatrace production context
github实测增长打开来源 ↗
安装
git clone https://github.com/Dynatrace/dynatrace-for-ai ~/.claude/skills/dynatrace-for-ai135GitHub 星标+4 (+3.1 %) - 20活跃
memroos
智能体MemroOS / memroos: memory OS and governance layer for AI agents, agent workflows, dispatch, proof, and context continuity.
github实测增长打开来源 ↗
7GitHub 星标稳定 - 21活跃
🔑 HomeStream · 家园·流 — 零成本自托管多Agent协作框架,通往AI世界的那把钥匙 | Zero-cost self-hosted multi-agent framework — The key to AI world
github实测增长打开来源 ↗
0GitHub 星标稳定 - 22活跃
adl-cli
其他A command-line tool to scaffold and manage enterprise-ready AI Agents powered by the A2A (Agent-to-Agent) protocol
github实测增长打开来源 ↗
14GitHub 星标+1 (+7.7 %) - 23活跃
Open-source infrastructure that turns scattered SKILL.md files into curated, retrieval-ready agent-skill corpora—with retrieval and evaluation tooling included.
github实测增长打开来源 ↗
安装
git clone https://github.com/EverMind-AI/SkillCorpus ~/.claude/skills/SkillCorpus525GitHub 星标+199 (+61.0 %) - 24活跃
74 open-source Agent Skills for Claude Code and Codex: AI SEO, AEO and GEO, code review with an A-F ship grade, CI gates, AI evals, design systems, conversion copy, Instagram growth, iOS and Android app shipping, creator rights, and…
github实测增长打开来源 ↗
129GitHub 星标-4 (-3.0 %) - 25活跃
Make Claude write clearly, for everyone.
github实测增长打开来源 ↗
安装
/plugin marketplace add stefanobaghino/simple-output-styles16GitHub 星标稳定