工具说明为英文。
LLM 可观测性
LLM 应用的监控、追踪、评估和质量管理。
用途
活跃度
排序方式
122
工具排名
按实测增长排序,并在不同来源间归一化。 原始值保留各自的数据窗口;计算变化需要 7 天内至少两次测量。
- 1休眠
anti-lie
技能Don't make LLMs honest. Make every factual claim auditable. — An LLM Claim Auditing Layer with T1-T7 truth gradients. 98.1% business effectiveness on LiarBench v0.2.
github实测增长打开来源 ↗
安装
git clone https://github.com/lc198707/anti-lie ~/.claude/skills/anti-lie89GitHub 星标稳定 - 2活跃
Production patterns for agent orchestrators, harnesses, evals, MCP/A2A interoperability, memory, provider routing, and LLM usage metering.
github实测增长打开来源 ↗
安装
git clone https://github.com/ssheleg/agent-stack ~/.claude/skills/agent-stack2GitHub 星标+1 (+100.0 %) - 3活跃
untell
其他AI-detector auditing toolkit: measures how often a detector flags genuine human writing, how stable its verdict is across seeds, and whether that verdict survives meaning-preserving edits. Detector-in-the-loop measurement harness. Claude…
github实测增长打开来源 ↗
18GitHub 星标稳定 - 4活跃
Log Claude Code sessions to Opik, the open-source LLM observability and evaluation platform, built by Comet. Tracing, evaluation, and skills for observable AI applications.
github实测增长打开来源 ↗
22GitHub 星标+1 (+4.8 %) - 5活跃
Deterministic, local-only audit reports for Claude Code AI agent sessions. Rust + SQLite. Zero network calls. MCP server for agents.
github实测增长打开来源 ↗
22GitHub 星标+2 (+10.0 %) - 6活跃
Deletion-first Agent Skill for removing test bloat, verification theater, and speculative fallbacks while preserving behavior.
github实测增长打开来源 ↗
安装
git clone https://github.com/MrZoyo/deslop-GPT ~/.claude/skills/deslop-GPT87GitHub 星标+59 (+210.7 %) - 73GitHub 星标稳定
- 8休眠
agentanvil
智能体Contract-based testing for LLM agents: hybrid evaluation, multi-agent + A2A, deterministic replay.
github实测增长打开来源 ↗
0GitHub 星标稳定 - 9活跃
Sentry instrumentation skill for system-behavior tracking
github实测增长打开来源 ↗
安装
git clone https://github.com/tortastudios/sentry-instrumentation ~/.claude/skills/sentry-instrumentation24GitHub 星标稳定 - 10休眠
A self-improving harness router for Claude Code.
github实测增长打开来源 ↗
安装
/plugin marketplace add SeongwoongCho/adaptive-harness8GitHub 星标稳定 - 11活跃
MCP server for Langfuse LLM observability — trace and observation analysis.
mcp实测增长打开来源 ↗
安装
claude mcp add langfuse -- npx langfuse-observability-mcp-server80GitHub 星标+2 (+2.6 %) - 12活跃
axiom
技能Axiom is a curated marketplace of shared plugins for Claude Code and Codex.
github实测增长打开来源 ↗
安装
git clone https://github.com/netopsengineer/axiom ~/.claude/skills/axiom5GitHub 星标稳定 - 13休眠
astragraph
智能体Policy-enforced observability and fail-closed guardrails for MCP/A2A multi-agent systems.
github实测增长打开来源 ↗
26GitHub 星标稳定 - 14活跃
A Go framework for building production agent systems with graph workflows, tools, memory, A2A, AG-UI, MCP, evaluation, and observability.
github实测增长打开来源 ↗
1 763GitHub 星标+15 (+0.86 %) - 150GitHub 星标稳定
- 16休眠
A safety-first multi-agent mental health companion with real-time distress tracking, triple-layer guardrails, and evidence-based grounding techniques. Built for Kaggle × Google Agents Intensive 2025 Capstone (Agents for Good Track)
github实测增长打开来源 ↗
1GitHub 星标稳定 - 17活跃
OpenJudge: A Unified Framework for Holistic Evaluation and Quality Rewards
github实测增长打开来源 ↗
安装
git clone https://github.com/agentscope-ai/OpenJudge ~/.claude/skills/OpenJudge816GitHub 星标+11 (+1.4 %) - 18休眠
Official Python SDK for GT8004 — AI agent observability with MCP, A2A, x402 payment tracking. FastAPI, Flask, FastMCP middleware included.
github实测增长打开来源 ↗
1GitHub 星标稳定 - 19休眠
🔍 AI observability skill for Claude Code. Debug LangChain/LangGraph agents by fetching execution traces from LangSmith Studio directly in your terminal.
github实测增长打开来源 ↗
安装
git clone https://github.com/OthmanAdi/langsmith-fetch-skill ~/.claude/skills/langsmith-fetch-skill27GitHub 星标稳定 - 20活跃
local-first analytics for AI agent skills
github实测增长打开来源 ↗
安装
git clone https://github.com/crafter-station/skill-kit ~/.claude/skills/skill-kit77GitHub 星标+1 (+1.3 %) - 21活跃
8 MCP SMB products — standalone AI servers for SMBs: Guardrails, FinOps, Observability, Router, Trust Score, Memory, ThinkSecure, A2A Lite. 37 tools, TNC credits billing.
github实测增长打开来源 ↗
1GitHub 星标稳定 - 22休眠
agentops
智能体AgentOps: Multi-agent infrastructure remediation platform. A2A protocol, agent coordination, HITL approval, auto-rollback.
github实测增长打开来源 ↗
0GitHub 星标稳定 - 23活跃
aura
MCPAURA is a production-tested SRE agent platform you can deploy in minutes. AURA handles the guardrails, APIs, state management, streaming, and failure handling required to put AI to work safely on production infrastructure.
github实测增长打开来源 ↗
339GitHub 星标+81 (+31.4 %) - 24活跃
Agent Skill evaluation harness for paired variants, trace artifacts, and runner adapters
github实测增长打开来源 ↗
安装
git clone https://github.com/adewale/skill-eval-harness ~/.claude/skills/skill-eval-harness73GitHub 星标+2 (+2.8 %) - 25休眠
Local open-source dev tool to debug, secure, and evaluate LLM agents. Provides static analysis, dynamic security checks, and runtime monitoring - integrates with Cursor and Claude Code.
github实测增长打开来源 ↗
安装
/plugin marketplace add cylestio/agent-inspector9GitHub 星标稳定