工具说明为英文。
LLM 可观测性
LLM 应用的监控、追踪、评估和质量管理。
用途
活跃度
排序方式
123
工具排名
按实测增长排序,并在不同来源间归一化。 原始值保留各自的数据窗口;计算变化需要 7 天内至少两次测量。
- 1活跃
Open-source infrastructure that turns scattered SKILL.md files into curated, retrieval-ready agent-skill corpora—with retrieval and evaluation tooling included.
github实测增长打开来源 ↗
安装
git clone https://github.com/EverMind-AI/SkillCorpus ~/.claude/skills/SkillCorpus360GitHub 星标+168 (+87.5 %) - 2活跃
Multi-tier framework for evaluating AI agent skills with quality gates, semantic overlap detection, synthetic evaluation dataset generation, and live agent evaluation that measures how skills affect agent behavior.
github实测增长打开来源 ↗
安装
git clone https://github.com/NVIDIA/SkillEvaluator ~/.claude/skills/SkillEvaluator363GitHub 星标+58 (+19.0 %) - 3活跃
Research-backed, eval-driven skills for AI agents
github实测增长打开来源 ↗
安装
git clone https://github.com/ZSeven-W/craft-skills ~/.claude/skills/craft-skills138GitHub 星标+21 (+17.9 %) - 4活跃
agent-kernel
智能体The Operating System for Scalable Enterprise AI Agents - Run, orchestrate, and deploy Compliant Enterprise AI Agents at scale across frameworks, without lock-in, rewrites or fragile glue code. Native support for MCP, A2A. Interface with all…
github实测增长打开来源 ↗
146GitHub 星标+13 (+9.8 %) - 5活跃
Agent Skill evaluation harness for paired variants, trace artifacts, and runner adapters
github实测增长打开来源 ↗
安装
git clone https://github.com/adewale/skill-eval-harness ~/.claude/skills/skill-eval-harness72GitHub 星标+4 (+5.9 %) - 6活跃
Skills, prompts, and instructions for building AI agents on top of Dynatrace production context
github实测增长打开来源 ↗
安装
git clone https://github.com/Dynatrace/dynatrace-for-ai ~/.claude/skills/dynatrace-for-ai131GitHub 星标+5 (+4.0 %) - 7活跃
databuff
智能体DataBuff is an AI-native APM built on Opentelemetry,with multi-agent troubleshooting out of the box.
github实测增长打开来源 ↗
627GitHub 星标+22 (+3.6 %) - 8活跃
Vendor-neutral OpenTelemetry skills for AI coding agents, grounded in upstream sources
github实测增长打开来源 ↗
安装
git clone https://github.com/ollygarden/opentelemetry-agent-skills ~/.claude/skills/opentelemetry-agent-skills96GitHub 星标+3 (+3.2 %) - 9活跃
YAO = Yielding AI Outcomes. A rigorous engineering, evaluation, governance, and portability system for reusable agent skills.
github实测增长打开来源 ↗
安装
git clone https://github.com/yaojingang/yao-meta-skill ~/.claude/skills/yao-meta-skill2 565GitHub 星标+58 (+2.3 %) - 10活跃
A test runner for agentskills.io-style AI agent skills
github实测增长打开来源 ↗
安装
git clone https://github.com/darkrishabh/agent-skills-eval ~/.claude/skills/agent-skills-eval713GitHub 星标+12 (+1.7 %) - 11活跃
local-first analytics for AI agent skills
github实测增长打开来源 ↗
安装
git clone https://github.com/crafter-station/skill-kit ~/.claude/skills/skill-kit77GitHub 星标+1 (+1.3 %) - 12活跃
MCP server for Langfuse LLM observability — trace and observation analysis.
mcp实测增长打开来源 ↗
安装
claude mcp add langfuse -- npx langfuse-observability-mcp-server78GitHub 星标+1 (+1.3 %) - 13活跃
A Go framework for building production agent systems with graph workflows, tools, memory, A2A, AG-UI, MCP, evaluation, and observability.
github实测增长打开来源 ↗
1 752GitHub 星标+22 (+1.3 %) - 14活跃
Build tested agent skills and govern their lifecycle through a user-defined marketplace: evidence, discovery, updates, rollback, quarantine, and 17-platform distribution.
github实测增长打开来源 ↗
安装
git clone https://github.com/FrancyJGLisboa/agent-skill-creator ~/.claude/skills/agent-skill-creator2 358GitHub 星标+29 (+1.2 %) - 15活跃
The Fable Workflow: how Claude Fable 5 worked, distilled into skills any model can run, with the eval that keeps it honest. Think / act / prove.
github实测增长打开来源 ↗
安装
git clone https://github.com/Sahir619/fable-method ~/.claude/skills/fable-method2 269GitHub 星标+22 (+0.98 %) - 16105GitHub 星标+1 (+0.96 %)
- 17活跃
OpenJudge: A Unified Framework for Holistic Evaluation and Quality Rewards
github实测增长打开来源 ↗
安装
git clone https://github.com/agentscope-ai/OpenJudge ~/.claude/skills/OpenJudge807GitHub 星标+7 (+0.88 %) - 18活跃
🪢 Langfuse documentation -- Langfuse is the open source LLM Engineering Platform. Observability, evals, prompt management, playground and metrics to debug and improve LLM apps
github实测增长打开来源 ↗
安装
git clone https://github.com/langfuse/langfuse-docs ~/.claude/skills/langfuse-docs237GitHub 星标+2 (+0.85 %) - 19活跃
aura
MCPAURA is a production-tested SRE agent platform you can deploy in minutes. AURA handles the guardrails, APIs, state management, streaming, and failure handling required to put AI to work safely on production infrastructure.
github实测增长打开来源 ↗
258GitHub 星标+2 (+0.78 %) - 20活跃
Dashboard for monitoring claude code sessions.
github实测增长打开来源 ↗
安装
/plugin marketplace add JayantDevkar/claude-code-karma321GitHub 星标+2 (+0.63 %) - 21活跃
A skill creator that proves its skills work. Evidence-driven skill creation for Claude Code and Codex: baseline-tested generation, per-skill regression evals, ecosystem doctor, cross-runtime compile, and an opt-in proactive advisor.
github实测增长打开来源 ↗
安装
git clone https://github.com/tripleyak/SkillForge ~/.claude/skills/SkillForge884GitHub 星标+5 (+0.57 %) - 22活跃
Test your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare performance of GPT, Claude, Gemini, DeepSeek, and more. Simple declarative configs with command line and CI/CD integration. Used by OpenAI…
github实测增长打开来源 ↗
24 710GitHub 星标+137 (+0.56 %) - 23活跃
Open source observability platform for logs, metrics, traces, frontend monitoring, pipelines and LLM observability. A sophisticated, simple and highly performant alternative to Datadog, Splunk, and Elasticsearch with 140x lower storage…
github实测增长打开来源 ↗
21 594GitHub 星标+119 (+0.55 %) - 249 079GitHub 星标+46 (+0.51 %)
学习与参考资源
按实测增长排序,并在不同来源间归一化。 这些资源可单独访问,不参与主要排名。
- 1活跃打开来源 ↗
Curated + scored map of official MCP servers and agents for DevOps, Cloud, SRE, and Platform Engineering — every entry rated on production access, approval gates, and audit evidence.
github资源实测增长77GitHub 星标+1 (+1.3 %)