工具说明为英文。
LLM 可观测性
LLM 应用的监控、追踪、评估和质量管理。
用途
活跃度
排序方式
117
工具排名
按实测增长排序,并在不同来源间归一化。 原始值保留各自的数据窗口;计算变化需要 7 天内至少两次测量。
- 1活跃
Open-source infrastructure that turns scattered SKILL.md files into curated, retrieval-ready agent-skill corpora—with retrieval and evaluation tooling included.
github实测增长打开来源 ↗
安装
git clone https://github.com/EverMind-AI/SkillCorpus ~/.claude/skills/SkillCorpus424GitHub 星标+170 (+66.9 %) - 2活跃
Test your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare performance of GPT, Claude, Gemini, DeepSeek, and more. Simple declarative configs with command line and CI/CD integration. Used by OpenAI…
github实测增长打开来源 ↗
24 768GitHub 星标+142 (+0.58 %) - 3活跃
mastra
其他Mastra is the modern TypeScript framework for AI-powered applications and agents.
github实测增长打开来源 ↗
27 658GitHub 星标+128 (+0.46 %) - 4活跃
Open source observability platform for logs, metrics, traces, frontend monitoring, pipelines and LLM observability. A sophisticated, simple and highly performant alternative to Datadog, Splunk, and Elasticsearch with 140x lower storage…
github实测增长打开来源 ↗
21 615GitHub 星标+96 (+0.45 %) - 5活跃
netdata
其他The fastest path to AI-powered full stack observability, even for lean teams.
github实测增长打开来源 ↗
80 412GitHub 星标+85 (+0.11 %) - 6活跃
mlflow
智能体The open source AI engineering platform for agents, LLMs, and ML models. MLflow enables teams of all sizes to debug, evaluate, monitor, and optimize production-quality AI applications while controlling costs and managing access to models…
github实测增长打开来源 ↗
27 783GitHub 星标+82 (+0.30 %) - 7活跃
Prefect is a workflow orchestration framework for building resilient data pipelines in Python.
github实测增长打开来源 ↗
23 766GitHub 星标+66 (+0.28 %) - 8活跃
Multi-tier framework for evaluating AI agent skills with quality gates, semantic overlap detection, synthetic evaluation dataset generation, and live agent evaluation that measures how skills affect agent behavior.
github实测增长打开来源 ↗
安装
git clone https://github.com/NVIDIA/SkillEvaluator ~/.claude/skills/SkillEvaluator389GitHub 星标+57 (+17.2 %) - 9活跃
signoz
其他SigNoz is an open-source, OpenTelemetry-native observability platform for your team and their AI agents. Get logs, metrics, and traces in one tool with features like APM, distributed tracing, log management, infra monitoring, etc. Combined…
github实测增长打开来源 ↗
32 003GitHub 星标+50 (+0.16 %) - 1022 509GitHub 星标+40 (+0.18 %)
- 11活跃
YAO = Yielding AI Outcomes. A rigorous engineering, evaluation, governance, and portability system for reusable agent skills.
github实测增长打开来源 ↗
安装
git clone https://github.com/yaojingang/yao-meta-skill ~/.claude/skills/yao-meta-skill2 577GitHub 星标+35 (+1.4 %) - 12活跃
databuff
智能体DataBuff is an AI-native APM built on Opentelemetry,with multi-agent troubleshooting out of the box.
github实测增长打开来源 ↗
642GitHub 星标+32 (+5.2 %) - 139 056GitHub 星标+31 (+0.34 %)
- 149 079GitHub 星标+31 (+0.34 %)
- 1525 047GitHub 星标+29 (+0.12 %)
- 16活跃
Build tested agent skills and govern their lifecycle through a user-defined marketplace: evidence, discovery, updates, rollback, quarantine, and 17-platform distribution.
github实测增长打开来源 ↗
安装
git clone https://github.com/FrancyJGLisboa/agent-skill-creator ~/.claude/skills/agent-skill-creator2 367GitHub 星标+28 (+1.2 %) - 17活跃
Deletion-first Agent Skill for removing test bloat, verification theater, and speculative fallbacks while preserving behavior.
github实测增长打开来源 ↗
安装
git clone https://github.com/MrZoyo/deslop-GPT ~/.claude/skills/deslop-GPT55GitHub 星标+27 (+96.4 %) - 18活跃
VeriRun
其他Evidence-first infrastructure for reproducible, isolated executable evaluation and online rewards.
github实测增长打开来源 ↗
140GitHub 星标+24 (+20.7 %) - 19活跃
agent-kernel
智能体The Operating System for Scalable Enterprise AI Agents - Run, orchestrate, and deploy Compliant Enterprise AI Agents at scale across frameworks, without lock-in, rewrites or fragile glue code. Native support for MCP, A2A. Interface with all…
github实测增长打开来源 ↗
156GitHub 星标+19 (+13.9 %) - 20活跃
Research-backed, eval-driven skills for AI agents
github实测增长打开来源 ↗
安装
git clone https://github.com/ZSeven-W/craft-skills ~/.claude/skills/craft-skills142GitHub 星标+18 (+14.5 %) - 21活跃
A Go framework for building production agent systems with graph workflows, tools, memory, A2A, AG-UI, MCP, evaluation, and observability.
github实测增长打开来源 ↗
1 759GitHub 星标+18 (+1.0 %) - 22活跃
The Fable Workflow: how Claude Fable 5 worked, distilled into skills any model can run, with the eval that keeps it honest. Think / act / prove.
github实测增长打开来源 ↗
安装
git clone https://github.com/Sahir619/fable-method ~/.claude/skills/fable-method2 272GitHub 星标+16 (+0.71 %) - 23活跃
A test runner for agentskills.io-style AI agent skills
github实测增长打开来源 ↗
安装
git clone https://github.com/darkrishabh/agent-skills-eval ~/.claude/skills/agent-skills-eval719GitHub 星标+13 (+1.8 %) - 24活跃
kubeshark
智能体eBPF-powered network observability for Kubernetes. Indexes L4/L7 traffic with full K8s context, decrypts TLS without keys. Queryable by AI agents via MCP and humans via dashboard.
github实测增长打开来源 ↗
12 068GitHub 星标+8 (+0.07 %) - 25活跃
The container platform tailored for Kubernetes multi-cloud, datacenter, and edge management ⎈ 🖥 ☁️
github实测增长打开来源 ↗
安装
git clone https://github.com/kubesphere/kubesphere ~/.claude/skills/kubesphere17 037GitHub 星标+8 (+0.05 %)