工具说明为英文。

LLM 可观测性

LLM 应用的监控、追踪、评估和质量管理。

用途

活跃度

排序方式

75

工具排名

按实测增长排序,并在不同来源间归一化。 原始值保留各自的数据窗口;计算变化需要 7 天内至少两次测量。

  1. 1

    技能
    活跃

    Open-source infrastructure that turns scattered SKILL.md files into curated, retrieval-ready agent-skill corpora—with retrieval and evaluation tooling included.

    github实测增长打开来源 ↗

    安装 git clone https://github.com/EverMind-AI/SkillCorpus ~/.claude/skills/SkillCorpus

    424GitHub 星标+170 (+66.9 %)
  2. 2

    其他
    活跃

    Evidence-first infrastructure for reproducible, isolated executable evaluation and online rewards.

    github实测增长打开来源 ↗

    140GitHub 星标+24 (+20.7 %)
  3. 3
    活跃

    Multi-tier framework for evaluating AI agent skills with quality gates, semantic overlap detection, synthetic evaluation dataset generation, and live agent evaluation that measures how skills affect agent behavior.

    github实测增长打开来源 ↗

    安装 git clone https://github.com/NVIDIA/SkillEvaluator ~/.claude/skills/SkillEvaluator

    389GitHub 星标+57 (+17.2 %)
  4. 4

    技能
    活跃

    Research-backed, eval-driven skills for AI agents

    github实测增长打开来源 ↗

    安装 git clone https://github.com/ZSeven-W/craft-skills ~/.claude/skills/craft-skills

    142GitHub 星标+18 (+14.5 %)
  5. 5

    智能体
    活跃

    The Operating System for Scalable Enterprise AI Agents - Run, orchestrate, and deploy Compliant Enterprise AI Agents at scale across frameworks, without lock-in, rewrites or fragile glue code. Native support for MCP, A2A. Interface with all…

    github实测增长打开来源 ↗

    156GitHub 星标+19 (+13.9 %)
  6. 6
    活跃

    Agent Skill evaluation harness for paired variants, trace artifacts, and runner adapters

    github实测增长打开来源 ↗

    安装 git clone https://github.com/adewale/skill-eval-harness ~/.claude/skills/skill-eval-harness

    73GitHub 星标+4 (+5.8 %)
  7. 7

    智能体
    活跃

    DataBuff is an AI-native APM built on Opentelemetry,with multi-agent troubleshooting out of the box.

    github实测增长打开来源 ↗

    642GitHub 星标+32 (+5.2 %)
  8. 8
    活跃

    Vendor-neutral OpenTelemetry skills for AI coding agents, grounded in upstream sources

    github实测增长打开来源 ↗

    安装 git clone https://github.com/ollygarden/opentelemetry-agent-skills ~/.claude/skills/opentelemetry-agent-skills

    97GitHub 星标+4 (+4.3 %)
  9. 9
    活跃

    Skills, prompts, and instructions for building AI agents on top of Dynatrace production context

    github实测增长打开来源 ↗

    安装 git clone https://github.com/Dynatrace/dynatrace-for-ai ~/.claude/skills/dynatrace-for-ai

    132GitHub 星标+4 (+3.1 %)
  10. 10
    活跃

    A test runner for agentskills.io-style AI agent skills

    github实测增长打开来源 ↗

    安装 git clone https://github.com/darkrishabh/agent-skills-eval ~/.claude/skills/agent-skills-eval

    719GitHub 星标+13 (+1.8 %)
  11. 11
    活跃

    YAO = Yielding AI Outcomes. A rigorous engineering, evaluation, governance, and portability system for reusable agent skills.

    github实测增长打开来源 ↗

    安装 git clone https://github.com/yaojingang/yao-meta-skill ~/.claude/skills/yao-meta-skill

    2 577GitHub 星标+35 (+1.4 %)
  12. 12

    技能
    活跃

    local-first analytics for AI agent skills

    github实测增长打开来源 ↗

    安装 git clone https://github.com/crafter-station/skill-kit ~/.claude/skills/skill-kit

    77GitHub 星标+1 (+1.3 %)
  13. 13

    技能
    活跃

    🪢 Langfuse documentation -- Langfuse is the open source LLM Engineering Platform. Observability, evals, prompt management, playground and metrics to debug and improve LLM apps

    github实测增长打开来源 ↗

    安装 git clone https://github.com/langfuse/langfuse-docs ~/.claude/skills/langfuse-docs

    239GitHub 星标+3 (+1.3 %)
  14. 14
    活跃

    Build tested agent skills and govern their lifecycle through a user-defined marketplace: evidence, discovery, updates, rollback, quarantine, and 17-platform distribution.

    github实测增长打开来源 ↗

    安装 git clone https://github.com/FrancyJGLisboa/agent-skill-creator ~/.claude/skills/agent-skill-creator

    2 367GitHub 星标+28 (+1.2 %)
  15. 15
    活跃

    A Go framework for building production agent systems with graph workflows, tools, memory, A2A, AG-UI, MCP, evaluation, and observability.

    github实测增长打开来源 ↗

    1 759GitHub 星标+18 (+1.0 %)
  16. 16
    活跃

    Self improving agents through iterations

    github实测增长打开来源 ↗

    105GitHub 星标+1 (+0.96 %)
  17. 17
    活跃

    🐙 ADLC Team Skills — Agentic SDLC for Engineering Teams

    github实测增长打开来源 ↗

    安装 git clone https://github.com/tikalk/adlc-team-skills ~/.claude/skills/adlc-team-skills

    132GitHub 星标+1 (+0.76 %)
  18. 18
    活跃

    Dashboard for monitoring claude code sessions.

    github实测增长打开来源 ↗

    安装 /plugin marketplace add JayantDevkar/claude-code-karma

    323GitHub 星标+2 (+0.62 %)
  19. 19

    其他
    活跃

    Test your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare performance of GPT, Claude, Gemini, DeepSeek, and more. Simple declarative configs with command line and CI/CD integration. Used by OpenAI…

    github实测增长打开来源 ↗

    24 768GitHub 星标+142 (+0.58 %)
  20. 20

    其他
    活跃

    Mastra is the modern TypeScript framework for AI-powered applications and agents.

    github实测增长打开来源 ↗

    27 658GitHub 星标+128 (+0.46 %)
  21. 21

    其他
    活跃

    Open source observability platform for logs, metrics, traces, frontend monitoring, pipelines and LLM observability. A sophisticated, simple and highly performant alternative to Datadog, Splunk, and Elasticsearch with 140x lower storage…

    github实测增长打开来源 ↗

    21 615GitHub 星标+96 (+0.45 %)
  22. 22
    活跃

    🫖 Status page with uptime monitoring & API monitoring as code 🫖

    github实测增长打开来源 ↗

    9 056GitHub 星标+31 (+0.34 %)
  23. 23

    其他
    活跃

    the LLM vulnerability scanner

    github实测增长打开来源 ↗

    9 079GitHub 星标+31 (+0.34 %)
  24. 24

    智能体
    活跃

    The open source AI engineering platform for agents, LLMs, and ML models. MLflow enables teams of all sizes to debug, evaluate, monitor, and optimize production-quality AI applications while controlling costs and managing access to models…

    github实测增长打开来源 ↗

    27 783GitHub 星标+82 (+0.30 %)
  25. 25

    活跃

    Prefect is a workflow orchestration framework for building resilient data pipelines in Python.

    github实测增长打开来源 ↗

    23 766GitHub 星标+66 (+0.28 %)