工具说明为英文。

LLM 可观测性

LLM 应用的监控、追踪、评估和质量管理。

用途

活跃度

排序方式

122

工具排名

按实测增长排序,并在不同来源间归一化。 原始值保留各自的数据窗口;计算变化需要 7 天内至少两次测量。

  1. 1
    活跃

    Query Spanlens LLM observability from Cursor, Claude Desktop, or Continue via MCP.

    mcp实测增长打开来源 ↗

    安装 claude mcp add mcp-server -- npx @spanlens/mcp-server

    12GitHub 星标稳定
  2. 2

    其他
    活跃

    Test your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare performance of GPT, Claude, Gemini, DeepSeek, and more. Simple declarative configs with command line and CI/CD integration. Used by OpenAI…

    github实测增长打开来源 ↗

    24 797GitHub 星标+146 (+0.59 %)
  3. 3

    其他
    活跃

    the LLM vulnerability scanner

    github实测增长打开来源 ↗

    9 106GitHub 星标+40 (+0.44 %)
  4. 4

    其他
    活跃

    AI SRE AgenticOps for Kubernetes and cloud infrastructure.

    github实测增长打开来源 ↗

    782GitHub 星标+1 (+0.13 %)
  5. 5

    智能体
    活跃

    The open source AI engineering platform for agents, LLMs, and ML models. MLflow enables teams of all sizes to debug, evaluate, monitor, and optimize production-quality AI applications while controlling costs and managing access to models…

    github实测增长打开来源 ↗

    27 801GitHub 星标+82 (+0.30 %)
  6. 6

    活跃

    Prefect is a workflow orchestration framework for building resilient data pipelines in Python.

    github实测增长打开来源 ↗

    23 772GitHub 星标+61 (+0.26 %)
  7. 7

    其他
    活跃

    A high-performance observability data pipeline.

    github实测增长打开来源 ↗

    22 510GitHub 星标+34 (+0.15 %)
  8. 8

    其他
    活跃

    Open source observability platform for logs, metrics, traces, frontend monitoring, pipelines and LLM observability. A sophisticated, simple and highly performant alternative to Datadog, Splunk, and Elasticsearch with 140x lower storage…

    github实测增长打开来源 ↗

    21 630GitHub 星标+88 (+0.41 %)
  9. 9

    其他
    活跃

    Mastra is the modern TypeScript framework for AI-powered applications and agents.

    github实测增长打开来源 ↗

    27 680GitHub 星标+132 (+0.48 %)
  10. 10

    其他
    活跃

    SigNoz is an open-source, OpenTelemetry-native observability platform for your team and their AI agents. Get logs, metrics, and traces in one tool with features like APM, distributed tracing, log management, infra monitoring, etc. Combined…

    github实测增长打开来源 ↗

    32 010GitHub 星标+52 (+0.16 %)
  11. 11

    其他
    活跃

    The fastest path to AI-powered full stack observability, even for lean teams.

    github实测增长打开来源 ↗

    80 424GitHub 星标+82 (+0.10 %)
  12. 12
    活跃

    🫖 Status page with uptime monitoring & API monitoring as code 🫖

    github实测增长打开来源 ↗

    9 059GitHub 星标+31 (+0.34 %)
  13. 13

    智能体
    活跃

    DataBuff is an AI-native APM built on Opentelemetry,with multi-agent troubleshooting out of the box.

    github实测增长打开来源 ↗

    649GitHub 星标+39 (+6.4 %)
  14. 14

    其他
    活跃

    Self-hosted AI SRE for Kubernetes — zero-instrumentation eBPF observability plus a copilot that fixes issues through guardrailed, self-verifying actions. BYO-LLM, air-gapped capable.

    github实测增长打开来源 ↗

    131GitHub 星标-4 (-3.0 %)
  15. 15

    其他
    活跃

    eBPF-based Networking, Security, and Observability

    github实测增长打开来源 ↗

    25 058GitHub 星标+36 (+0.14 %)
  16. 16

    智能体
    活跃

    eBPF-powered network observability for Kubernetes. Indexes L4/L7 traffic with full K8s context, decrypts TLS without keys. Queryable by AI agents via MCP and humans via dashboard.

    github实测增长打开来源 ↗

    12 068GitHub 星标+7 (+0.06 %)
  17. 17

    其他
    活跃

    Evidence-first infrastructure for reproducible, isolated executable evaluation and online rewards.

    github实测增长打开来源 ↗

    163GitHub 星标+47 (+40.5 %)
  18. 18

    技能
    活跃

    The Fable Workflow: how Claude Fable 5 worked, distilled into skills any model can run, with the eval that keeps it honest. Think / act / prove.

    github实测增长打开来源 ↗

    安装 git clone https://github.com/Sahir619/fable-method ~/.claude/skills/fable-method

    2 273GitHub 星标+11 (+0.49 %)
  19. 19
    活跃

    Two Claude Code skills for adversarial code review (single-model PAR and multi-model MMAR with cross-critique to catch hallucinations), plus a fixture-based eval suite.

    github实测增长打开来源 ↗

    安装 /plugin marketplace add prime-radiant-inc/parallel-adversarial-review

    17GitHub 星标稳定
  20. 20

    技能
    休眠

    An assembly line for AI software development. 35 skills, 11 agent personas, 29 commands. From raw idea to shipped code.

    github实测增长打开来源 ↗

    安装 /plugin marketplace add aneja5/forge-skills

    3GitHub 星标稳定
  21. 21
    活跃

    Make Claude write clearly, for everyone.

    github实测增长打开来源 ↗

    安装 /plugin marketplace add stefanobaghino/simple-output-styles

    16GitHub 星标稳定
  22. 22

    智能体
    活跃

    面向长程研发任务的 Java Agent Harness,支持 A2A 跨语言协作、可恢复执行、上下文工程与 Eval 驱动开发。

    github实测增长打开来源 ↗

    0GitHub 星标稳定
  23. 23
    活跃

    Production reliability, benchmarking, and evaluation layer for MCP agent runtimes — protocol negotiation, A2A interop, circuit breakers, connection pooling, adaptive streaming, and autoscaling, with measured throughput/latency and a…

    github实测增长打开来源 ↗

    3GitHub 星标稳定
  24. 24

    其他
    活跃

    A command-line tool to scaffold and manage enterprise-ready AI Agents powered by the A2A (Agent-to-Agent) protocol

    github实测增长打开来源 ↗

    14GitHub 星标+1 (+7.7 %)
  25. 25

    技能
    活跃

    Open-source infrastructure that turns scattered SKILL.md files into curated, retrieval-ready agent-skill corpora—with retrieval and evaluation tooling included.

    github实测增长打开来源 ↗

    安装 git clone https://github.com/EverMind-AI/SkillCorpus ~/.claude/skills/SkillCorpus

    493GitHub 星标+205 (+71.2 %)