工具说明为英文。
LLM 可观测性
LLM 应用的监控、追踪、评估和质量管理。
用途
活跃度
排序方式
75
工具排名
按实测增长排序,并在不同来源间归一化。 原始值保留各自的数据窗口;计算变化需要 7 天内至少两次测量。
- 1活跃
Open-source infrastructure that turns scattered SKILL.md files into curated, retrieval-ready agent-skill corpora—with retrieval and evaluation tooling included.
github实测增长打开来源 ↗
安装
git clone https://github.com/EverMind-AI/SkillCorpus ~/.claude/skills/SkillCorpus424GitHub 星标+170 (+66.9 %) - 2活跃
VeriRun
其他Evidence-first infrastructure for reproducible, isolated executable evaluation and online rewards.
github实测增长打开来源 ↗
140GitHub 星标+24 (+20.7 %) - 3活跃
Multi-tier framework for evaluating AI agent skills with quality gates, semantic overlap detection, synthetic evaluation dataset generation, and live agent evaluation that measures how skills affect agent behavior.
github实测增长打开来源 ↗
安装
git clone https://github.com/NVIDIA/SkillEvaluator ~/.claude/skills/SkillEvaluator389GitHub 星标+57 (+17.2 %) - 4活跃
Research-backed, eval-driven skills for AI agents
github实测增长打开来源 ↗
安装
git clone https://github.com/ZSeven-W/craft-skills ~/.claude/skills/craft-skills142GitHub 星标+18 (+14.5 %) - 5活跃
agent-kernel
智能体The Operating System for Scalable Enterprise AI Agents - Run, orchestrate, and deploy Compliant Enterprise AI Agents at scale across frameworks, without lock-in, rewrites or fragile glue code. Native support for MCP, A2A. Interface with all…
github实测增长打开来源 ↗
156GitHub 星标+19 (+13.9 %) - 6活跃
Agent Skill evaluation harness for paired variants, trace artifacts, and runner adapters
github实测增长打开来源 ↗
安装
git clone https://github.com/adewale/skill-eval-harness ~/.claude/skills/skill-eval-harness73GitHub 星标+4 (+5.8 %) - 7活跃
databuff
智能体DataBuff is an AI-native APM built on Opentelemetry,with multi-agent troubleshooting out of the box.
github实测增长打开来源 ↗
642GitHub 星标+32 (+5.2 %) - 8活跃
Vendor-neutral OpenTelemetry skills for AI coding agents, grounded in upstream sources
github实测增长打开来源 ↗
安装
git clone https://github.com/ollygarden/opentelemetry-agent-skills ~/.claude/skills/opentelemetry-agent-skills97GitHub 星标+4 (+4.3 %) - 9活跃
Skills, prompts, and instructions for building AI agents on top of Dynatrace production context
github实测增长打开来源 ↗
安装
git clone https://github.com/Dynatrace/dynatrace-for-ai ~/.claude/skills/dynatrace-for-ai132GitHub 星标+4 (+3.1 %) - 10活跃
A test runner for agentskills.io-style AI agent skills
github实测增长打开来源 ↗
安装
git clone https://github.com/darkrishabh/agent-skills-eval ~/.claude/skills/agent-skills-eval719GitHub 星标+13 (+1.8 %) - 11活跃
YAO = Yielding AI Outcomes. A rigorous engineering, evaluation, governance, and portability system for reusable agent skills.
github实测增长打开来源 ↗
安装
git clone https://github.com/yaojingang/yao-meta-skill ~/.claude/skills/yao-meta-skill2 577GitHub 星标+35 (+1.4 %) - 12活跃
local-first analytics for AI agent skills
github实测增长打开来源 ↗
安装
git clone https://github.com/crafter-station/skill-kit ~/.claude/skills/skill-kit77GitHub 星标+1 (+1.3 %) - 13活跃
🪢 Langfuse documentation -- Langfuse is the open source LLM Engineering Platform. Observability, evals, prompt management, playground and metrics to debug and improve LLM apps
github实测增长打开来源 ↗
安装
git clone https://github.com/langfuse/langfuse-docs ~/.claude/skills/langfuse-docs239GitHub 星标+3 (+1.3 %) - 14活跃
Build tested agent skills and govern their lifecycle through a user-defined marketplace: evidence, discovery, updates, rollback, quarantine, and 17-platform distribution.
github实测增长打开来源 ↗
安装
git clone https://github.com/FrancyJGLisboa/agent-skill-creator ~/.claude/skills/agent-skill-creator2 367GitHub 星标+28 (+1.2 %) - 15活跃
A Go framework for building production agent systems with graph workflows, tools, memory, A2A, AG-UI, MCP, evaluation, and observability.
github实测增长打开来源 ↗
1 759GitHub 星标+18 (+1.0 %) - 16105GitHub 星标+1 (+0.96 %)
- 17活跃
🐙 ADLC Team Skills — Agentic SDLC for Engineering Teams
github实测增长打开来源 ↗
安装
git clone https://github.com/tikalk/adlc-team-skills ~/.claude/skills/adlc-team-skills132GitHub 星标+1 (+0.76 %) - 18活跃
Dashboard for monitoring claude code sessions.
github实测增长打开来源 ↗
安装
/plugin marketplace add JayantDevkar/claude-code-karma323GitHub 星标+2 (+0.62 %) - 19活跃
Test your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare performance of GPT, Claude, Gemini, DeepSeek, and more. Simple declarative configs with command line and CI/CD integration. Used by OpenAI…
github实测增长打开来源 ↗
24 768GitHub 星标+142 (+0.58 %) - 20活跃
mastra
其他Mastra is the modern TypeScript framework for AI-powered applications and agents.
github实测增长打开来源 ↗
27 658GitHub 星标+128 (+0.46 %) - 21活跃
Open source observability platform for logs, metrics, traces, frontend monitoring, pipelines and LLM observability. A sophisticated, simple and highly performant alternative to Datadog, Splunk, and Elasticsearch with 140x lower storage…
github实测增长打开来源 ↗
21 615GitHub 星标+96 (+0.45 %) - 229 056GitHub 星标+31 (+0.34 %)
- 239 079GitHub 星标+31 (+0.34 %)
- 24活跃
mlflow
智能体The open source AI engineering platform for agents, LLMs, and ML models. MLflow enables teams of all sizes to debug, evaluate, monitor, and optimize production-quality AI applications while controlling costs and managing access to models…
github实测增长打开来源 ↗
27 783GitHub 星标+82 (+0.30 %) - 25活跃
Prefect is a workflow orchestration framework for building resilient data pipelines in Python.
github实测增长打开来源 ↗
23 766GitHub 星标+66 (+0.28 %)