工具说明为英文。

LLM 可观测性

LLM 应用的监控、追踪、评估和质量管理。

用途

活跃度

排序方式

122

工具排名

混合排名:实测增长优先。 原始值保留各自的数据窗口;计算变化需要 7 天内至少两次测量。

  1. 1

    其他
    休眠

    A Claude Code skill that adds a rubric-based eval layer to any agent project. Framework-agnostic — generates rubric, test cases, judge prompt, and harness. Returns a weighted score plus a judge-leniency signal.

    github实测增长打开来源 ↗

    13GitHub 星标稳定
  2. 2

    智能体
    活跃

    Local-first control plane for durable, governed agent teams: DAG orchestration, approvals, memory, signed skills, trace contracts, and realtime.

    github实测增长打开来源 ↗

    0GitHub 星标稳定
  3. 3

    智能体
    活跃

    MemroOS / memroos: memory OS and governance layer for AI agents, agent workflows, dispatch, proof, and context continuity.

    github实测增长打开来源 ↗

    7GitHub 星标稳定
  4. 4

    其他
    活跃

    🔑 HomeStream · 家园·流 — 零成本自托管多Agent协作框架,通往AI世界的那把钥匙 | Zero-cost self-hosted multi-agent framework — The key to AI world

    github实测增长打开来源 ↗

    0GitHub 星标稳定
  5. 5

    其他
    活跃

    A command-line tool to scaffold and manage enterprise-ready AI Agents powered by the A2A (Agent-to-Agent) protocol

    github实测增长打开来源 ↗

    14GitHub 星标+1 (+7.7 %)
  6. 6
    活跃

    Make Claude write clearly, for everyone.

    github实测增长打开来源 ↗

    安装 /plugin marketplace add stefanobaghino/simple-output-styles

    16GitHub 星标稳定
  7. 7

    技能
    活跃

    Marketing Mix Model expert agent plugin for Claude Code - pymc-marketing v0.18.2+

    github实测增长打开来源 ↗

    安装 /plugin marketplace add Yakoub-ai/agent-mmm

    4GitHub 星标稳定
  8. 8

    智能体
    活跃

    The agent framework where the model never holds the trigger — every consequential action clears your policy first, waits for a human when it matters, and lands on a record you can verify. Build on it, or put it around the agent you already…

    github实测增长打开来源 ↗

    1GitHub 星标稳定
  9. 9

    智能体
    休眠

    A Multi-Agent System (MAS) evaluation framework using PydanticAI that generates and evaluates scientific paper reviews through a three-tiered assessment approach: traditional metrics, LLM-as-a-Judge, and graph-based complexity analysis.

    github实测增长打开来源 ↗

    2GitHub 星标稳定
  10. 10
    休眠

    MCP as a Judge: a behavioral MCP that strengthens AI coding assistants via explicit LLM evaluations

    mcp实测增长打开来源 ↗

    安装 claude mcp add mcp-as-a-judge -- uvx mcp-as-a-judge

    17GitHub 星标稳定
  11. 11
    休眠

    An AI-powered multi-agent system that demonstrates clinical triage, OTC medication recommendations, and e-pharmacy integration for respiratory conditions. Built with modular agents that collaborate to provide safe, intelligent healthcare…

    github实测增长打开来源 ↗

    0GitHub 星标稳定
  12. 12

    其他
    休眠

    Multi-agent customer support system with Google ADK & Gemini 2.5 Flash Lite. Kaggle capstone demonstrating 11+ concepts. Automates 80%+ queries, <10s response time.

    github实测增长打开来源 ↗

    2GitHub 星标稳定
  13. 13
    活跃

    Architecture-first Python scaffold for an auditable multi-agent corporate credit desk using A2A, MCP, deterministic credit policies, model routing, and OpenTelemetry.

    github实测增长打开来源 ↗

    0GitHub 星标稳定
  14. 14
    休眠

    Two Claude Code skills for adversarial code review (single-model PAR and multi-model MMAR with cross-critique to catch hallucinations), plus a fixture-based eval suite.

    github实测增长打开来源 ↗

    安装 /plugin marketplace add prime-radiant-inc/parallel-adversarial-review

    17GitHub 星标稳定
  15. 15
    活跃

    Mini Program Engineering Skill Suite is an Agent Skill suite for evidence‑first mini‑program development. It helps agents bring WeChat or other mini‑program projects from vague intent to reliable engineering work: project intake, product…

    github实测增长打开来源 ↗

    安装 git clone https://github.com/NocodeMrLi/mini-program-engineering-skill-suite ~/.claude/skills/mini-program-engineering-skill-suite

    45GitHub 星标+6 (+15.4 %)
  16. 16
    活跃

    Run many AIs on one board and keep control of all of it. Deterministic code decides who acts — never a model. A privacy floor keeps sensitive work on your machine, your own tests decide what counts as done, and every action lands on a…

    github实测增长打开来源 ↗

    安装 git clone https://github.com/sandhusukhdeep2/sc-prism-releases ~/.claude/skills/sc-prism-releases

    1GitHub 星标稳定
  17. 17

    其他
    活跃

    Open-source observability tool that uses AI agents to self-heal your software

    github估算动量打开来源 ↗

    1 404GitHub 星标

学习与参考资源

按实测增长排序,并在不同来源间归一化。 这些资源可单独访问,不参与主要排名。

  1. 1
    活跃打开来源 ↗

    My complete journey to becoming an Agentic AI Engineer through structured learning, projects, experiments, and production-ready implementations of modern AI systems.

    github资源实测增长
    18GitHub 星标稳定
  2. 2
    活跃打开来源 ↗

    50+ curated LLM observability tools PLUS 26 Agent Skills (several with runnable, unit-tested scripts) to build, evaluate, debug, secure & monitor reliable LLM apps. Tracing, evals, guardrails, LLMOps.

    github

    安装 git clone https://github.com/ContextJet-ai/awesome-llm-observability ~/.claude/skills/awesome-llm-observability

    资源实测增长
    33GitHub 星标+2 (+6.5 %)
  3. 3
    活跃打开来源 ↗

    Handbook técnico aberto sobre engenharia de IA em produção: ML tradicional, LLMs, RAG, agentes, segurança, observabilidade, FinOps e deployment.

    github资源实测增长
    1GitHub 星标稳定
  4. 4
    活跃打开来源 ↗

    Documentation-discovery telemetry for Claude Code — heat/cold maps, health grade, evidence-backed router fixes. 100% local, zero tokens. /tt

    github

    安装 /plugin marketplace add Hedde/trigger_tree

    资源实测增长
    14GitHub 星标稳定
  5. 5

    tunelab

    技能
    活跃打开来源 ↗

    Claude Code plugin for LLM fine-tuning, distillation, and evaluation — decide whether you need fine-tuning at all, distill your LLM logs into small local models (MLX/LoRA), evaluate with held-out discipline, and learn the why at every step.

    github

    安装 /plugin marketplace add rchaz/tunelab

    资源实测增长
    6GitHub 星标稳定