工具说明为英文。
LLM 可观测性
LLM 应用的监控、追踪、评估和质量管理。
用途
活跃度
排序方式
75
工具排名
混合排名:实测增长优先。 原始值保留各自的数据窗口;计算变化需要 7 天内至少两次测量。
- 1活跃
Architecture-first Python scaffold for an auditable multi-agent corporate credit desk using A2A, MCP, deterministic credit policies, model routing, and OpenTelemetry.
github实测增长打开来源 ↗
0GitHub 星标稳定 - 2活跃
Marketing Mix Model expert agent plugin for Claude Code - pymc-marketing v0.18.2+
github实测增长打开来源 ↗
安装
/plugin marketplace add Yakoub-ai/agent-mmm4GitHub 星标稳定 - 3活跃
AI Guardian
MCPGoverned local-LLM observability: model policy, prompt scanner, capture proxy, 21 tools.
mcp实测增长打开来源 ↗
安装
claude mcp add ai-guardian -- uvx ai-guardian-aiops0GitHub 星标稳定 - 4活跃
evoagent-os
智能体Local-first control plane for durable, governed agent teams: DAG orchestration, approvals, memory, signed skills, trace contracts, and realtime.
github实测增长打开来源 ↗
0GitHub 星标稳定 - 5活跃
Deterministic crash tests for agent control planes: paired worlds, hard stops, budgets, scopes, approvals, and hash-linked receipts.
github实测增长打开来源 ↗
1GitHub 星标稳定 - 6活跃
headsup
技能Glanceable Claude Code and Codex state in your terminal tabs: white=idle, blue=working, orange=waiting. Multi-terminal (iTerm2, WezTerm, AI Power Term), a live status bar with Anthropic usage limits, /sfl and /nil window save-and-restore,…
github实测增长打开来源 ↗
安装
git clone https://github.com/wasulajr/headsup ~/.claude/skills/headsup1GitHub 星标稳定 - 7活跃
dsh-plugins
智能体Generic DeepSeek Harness (dsh) plugins: A2A protocol server, session storage mirror, and Langfuse observability.
github实测增长打开来源 ↗
1GitHub 星标稳定 - 8活跃
a2a-otel-kit
MCPVendor-neutral OpenTelemetry tracing for A2A agents and MCP services, with W3C context propagation and privacy-safe telemetry.
github实测增长打开来源 ↗
1GitHub 星标稳定 - 9活跃
nora
MCPOpen-source, self-hosted control plane for OpenClaw and Hermes AI-agent fleets on Docker/Kubernetes — REST, CLI, and MCP.
github实测增长打开来源 ↗
50GitHub 星标+2 (+4.2 %) - 10活跃
Real-time execution trace and cost intelligence for Claude Code
github实测增长打开来源 ↗
安装
/plugin marketplace add DeibyGS/claudestat34GitHub 星标稳定 - 11活跃
tulip-agents
智能体The agent framework where the model never holds the trigger — every consequential action clears your policy first, waits for a human when it matters, and lands on a record you can verify. Build on it, or put it around the agent you already…
github实测增长打开来源 ↗
1GitHub 星标稳定 - 12活跃
Local-first Agentic AI Infrastructure Platform with LLM Gateway, Agent DAG Runtime, MCP Tool Hub, A2A Agent Mesh, Local RAG, Tool Sandbox and Observability.
github实测增长打开来源 ↗
1GitHub 星标稳定 - 13活跃
agentgateway
MCPAgentGateway — independent third-party profile of a public API surface, by API Evangelist. AgentGateway is an open-source, AI-native proxy and gateway for routing, observing, and governing traffic to and from AI agents, LLM providers, and…
github实测增长打开来源 ↗
0GitHub 星标稳定 - 14活跃
Deterministic, local-only audit reports for Claude Code AI agent sessions. Rust + SQLite. Zero network calls. MCP server for agents.
github实测增长打开来源 ↗
22GitHub 星标+2 (+10.0 %) - 15活跃
nudge
MCPA typed, replayable, budget-aware programming language for LLM agents. Compiles to Python & TypeScript.
github实测增长打开来源 ↗
2GitHub 星标稳定 - 16活跃
untell
其他AI-detector auditing toolkit: measures how often a detector flags genuine human writing, how stable its verdict is across seeds, and whether that verdict survives meaning-preserving edits. Detector-in-the-loop measurement harness. Claude…
github实测增长打开来源 ↗
18GitHub 星标稳定 - 17活跃
Multi-agent quality gate skill for Claude Code that researches, reviews, tests and challenges AI-generated work before the final answer.
github实测增长打开来源 ↗
安装
git clone https://github.com/ma-nucho-pro/supervisor-skill-claude ~/.claude/skills/supervisor-skill-claude2GitHub 星标稳定 - 18活跃
rashomon
技能Measure prompt and skill improvements with blind A/B comparison.
github实测增长打开来源 ↗
安装
/plugin marketplace add shinpr/rashomon18GitHub 星标稳定 - 19活跃
Production patterns for agent orchestrators, harnesses, evals, MCP/A2A interoperability, memory, provider routing, and LLM usage metering.
github实测增长打开来源 ↗
安装
git clone https://github.com/ssheleg/agent-stack ~/.claude/skills/agent-stack2GitHub 星标+1 (+100.0 %) - 20活跃
Deletion-first Agent Skill for removing test bloat, verification theater, and speculative fallbacks while preserving behavior.
github实测增长打开来源 ↗
安装
git clone https://github.com/MrZoyo/deslop-GPT ~/.claude/skills/deslop-GPT76GitHub 星标+48 (+171.4 %) - 21活跃
Mini Program Engineering Skill Suite is an Agent Skill suite for evidence‑first mini‑program development. It helps agents bring WeChat or other mini‑program projects from vague intent to reliable engineering work: project intake, product…
github实测增长打开来源 ↗
安装
git clone https://github.com/NocodeMrLi/mini-program-engineering-skill-suite ~/.claude/skills/mini-program-engineering-skill-suite43GitHub 星标+4 (+10.3 %) - 22活跃
Build tested agent skills and govern their lifecycle through a user-defined marketplace: evidence, discovery, updates, rollback, quarantine, and 17-platform distribution.
github估算动量打开来源 ↗
安装
git clone https://github.com/FrancyJGLisboa/agent-skills-platform ~/.claude/skills/agent-skills-platform2 376GitHub 星标—
学习与参考资源
按实测增长排序,并在不同来源间归一化。 这些资源可单独访问,不参与主要排名。
- 1活跃打开来源 ↗
Documentation-discovery telemetry for Claude Code — heat/cold maps, health grade, evidence-backed router fixes. 100% local, zero tokens. /tt
github安装
资源实测增长/plugin marketplace add Hedde/trigger_tree14GitHub 星标+1 (+7.7 %) - 2活跃打开来源 ↗
50+ curated LLM observability tools PLUS 26 Agent Skills (several with runnable, unit-tested scripts) to build, evaluate, debug, secure & monitor reliable LLM apps. Tracing, evals, guardrails, LLMOps.
github安装
资源实测增长git clone https://github.com/ContextJet-ai/awesome-llm-observability ~/.claude/skills/awesome-llm-observability33GitHub 星标+3 (+10.0 %) - 3活跃打开来源 ↗
My complete journey to becoming an Agentic AI Engineer through structured learning, projects, experiments, and production-ready implementations of modern AI systems.
github资源实测增长18GitHub 星标稳定