工具说明为英文。

LLM 可观测性

LLM 应用的监控、追踪、评估和质量管理。

用途

活跃度

排序方式

123

工具排名

Ranked by normalized popularity across sources. 原始值保留各自的数据窗口;计算变化需要 7 天内至少两次测量。

  1. 1

    其他
    活跃

    Test your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare performance of GPT, Claude, Gemini, DeepSeek, and more. Simple declarative configs with command line and CI/CD integration. Used by OpenAI…

    github实测增长打开来源 ↗

    24 738GitHub 星标+131 (+0.53 %)
  2. 2

    其他
    活跃

    SigNoz is an open-source, OpenTelemetry-native observability platform for your team and their AI agents. Get logs, metrics, and traces in one tool with features like APM, distributed tracing, log management, infra monitoring, etc. Combined…

    github实测增长打开来源 ↗

    31 992GitHub 星标+62 (+0.19 %)
  3. 3

    智能体
    活跃

    The open source AI engineering platform for agents, LLMs, and ML models. MLflow enables teams of all sizes to debug, evaluate, monitor, and optimize production-quality AI applications while controlling costs and managing access to models…

    github实测增长打开来源 ↗

    27 768GitHub 星标+80 (+0.29 %)
  4. 4

    活跃

    Prefect is a workflow orchestration framework for building resilient data pipelines in Python.

    github实测增长打开来源 ↗

    23 759GitHub 星标+66 (+0.28 %)
  5. 5

    其他
    活跃

    A high-performance observability data pipeline.

    github实测增长打开来源 ↗

    22 507GitHub 星标+45 (+0.20 %)
  6. 6

    其他
    活跃

    Open source observability platform for logs, metrics, traces, frontend monitoring, pipelines and LLM observability. A sophisticated, simple and highly performant alternative to Datadog, Splunk, and Elasticsearch with 140x lower storage…

    github实测增长打开来源 ↗

    21 608GitHub 星标+115 (+0.54 %)
  7. 7

    其他
    活跃

    Mastra is the modern TypeScript framework for AI-powered applications and agents.

    github实测增长打开来源 ↗

    27 624GitHub 星标+121 (+0.44 %)
  8. 8

    其他
    活跃

    The fastest path to AI-powered full stack observability, even for lean teams.

    github实测增长打开来源 ↗

    80 402GitHub 星标+91 (+0.11 %)
  9. 9

    其他
    活跃

    eBPF-based Networking, Security, and Observability

    github实测增长打开来源 ↗

    25 041GitHub 星标+27 (+0.11 %)
  10. 10

    智能体
    活跃

    eBPF-powered network observability for Kubernetes. Indexes L4/L7 traffic with full K8s context, decrypts TLS without keys. Queryable by AI agents via MCP and humans via dashboard.

    github实测增长打开来源 ↗

    12 066GitHub 星标+7 (+0.06 %)
  11. 11

    技能
    活跃

    The container platform tailored for Kubernetes multi-cloud, datacenter, and edge management ⎈ 🖥 ☁️

    github实测增长打开来源 ↗

    安装 git clone https://github.com/kubesphere/kubesphere ~/.claude/skills/kubesphere

    17 037GitHub 星标+9 (+0.05 %)
  12. 12

    其他
    活跃

    the LLM vulnerability scanner

    github实测增长打开来源 ↗

    9 079GitHub 星标+43 (+0.48 %)
  13. 13
    活跃

    🫖 Status page with uptime monitoring & API monitoring as code 🫖

    github实测增长打开来源 ↗

    9 052GitHub 星标+28 (+0.31 %)
  14. 14
    活跃

    YAO = Yielding AI Outcomes. A rigorous engineering, evaluation, governance, and portability system for reusable agent skills.

    github实测增长打开来源 ↗

    安装 git clone https://github.com/yaojingang/yao-meta-skill ~/.claude/skills/yao-meta-skill

    2 577GitHub 星标+54 (+2.1 %)
  15. 15
    活跃

    Build tested agent skills and govern their lifecycle through a user-defined marketplace: evidence, discovery, updates, rollback, quarantine, and 17-platform distribution.

    github实测增长打开来源 ↗

    安装 git clone https://github.com/FrancyJGLisboa/agent-skill-creator ~/.claude/skills/agent-skill-creator

    2 367GitHub 星标+30 (+1.3 %)
  16. 16

    技能
    活跃

    The Fable Workflow: how Claude Fable 5 worked, distilled into skills any model can run, with the eval that keeps it honest. Think / act / prove.

    github实测增长打开来源 ↗

    安装 git clone https://github.com/Sahir619/fable-method ~/.claude/skills/fable-method

    2 272GitHub 星标+19 (+0.84 %)
  17. 17
    活跃

    A Go framework for building production agent systems with graph workflows, tools, memory, A2A, AG-UI, MCP, evaluation, and observability.

    github实测增长打开来源 ↗

    1 759GitHub 星标+24 (+1.4 %)
  18. 18

    其他
    活跃

    Open-source observability tool that uses AI agents to self-heal your software

    github实测增长打开来源 ↗

    1 404GitHub 星标+2 (+0.14 %)
  19. 19

    技能
    活跃

    A skill creator that proves its skills work. Evidence-driven skill creation for Claude Code and Codex: baseline-tested generation, per-skill regression evals, ecosystem doctor, cross-runtime compile, and an opt-in proactive advisor.

    github实测增长打开来源 ↗

    安装 git clone https://github.com/tripleyak/SkillForge ~/.claude/skills/SkillForge

    884GitHub 星标+3 (+0.34 %)
  20. 20

    技能
    活跃

    OpenJudge: A Unified Framework for Holistic Evaluation and Quality Rewards

    github实测增长打开来源 ↗

    安装 git clone https://github.com/agentscope-ai/OpenJudge ~/.claude/skills/OpenJudge

    809GitHub 星标+10 (+1.3 %)
  21. 21

    其他
    活跃

    AI SRE AgenticOps for Kubernetes and cloud infrastructure.

    github实测增长打开来源 ↗

    782GitHub 星标+1 (+0.13 %)
  22. 22
    活跃

    A test runner for agentskills.io-style AI agent skills

    github实测增长打开来源 ↗

    安装 git clone https://github.com/darkrishabh/agent-skills-eval ~/.claude/skills/agent-skills-eval

    719GitHub 星标+15 (+2.1 %)
  23. 23

    智能体
    活跃

    DataBuff is an AI-native APM built on Opentelemetry,with multi-agent troubleshooting out of the box.

    github实测增长打开来源 ↗

    634GitHub 星标+26 (+4.3 %)
  24. 24

    技能
    活跃

    Open-source infrastructure that turns scattered SKILL.md files into curated, retrieval-ready agent-skill corpora—with retrieval and evaluation tooling included.

    github实测增长打开来源 ↗

    安装 git clone https://github.com/EverMind-AI/SkillCorpus ~/.claude/skills/SkillCorpus

    424GitHub 星标+204 (+92.7 %)
  25. 25
    活跃

    Multi-tier framework for evaluating AI agent skills with quality gates, semantic overlap detection, synthetic evaluation dataset generation, and live agent evaluation that measures how skills affect agent behavior.

    github实测增长打开来源 ↗

    安装 git clone https://github.com/NVIDIA/SkillEvaluator ~/.claude/skills/SkillEvaluator

    389GitHub 星标+63 (+19.3 %)