工具说明为英文。
LLM 可观测性
LLM 应用的监控、追踪、评估和质量管理。
用途
活跃度
排序方式
121
工具排名
按实测增长排序,并在不同来源间归一化。 原始值保留各自的数据窗口;计算变化需要 7 天内至少两次测量。
- 1活跃
Production-grade AI coding rules for Cursor and Claude Code. 15 rules + 9 doc templates + skills + agents + MCP setup. Drop into any project.
github实测增长打开来源 ↗
安装
git clone https://github.com/aiagentwithdhruv/ai-dev-stack ~/.claude/skills/ai-dev-stack10GitHub 星标稳定 - 2活跃
Claude Code plugin marketplace — agentic-engineering (spec-driven shape→decide→execute→measure→eval with adversarial review) + github-keeper (audit/elevate READMEs and make a repo open-source-ready).
github实测增长打开来源 ↗
安装
/plugin marketplace add GiustoPiedimonte/agentic-engineering-marketplace13GitHub 星标稳定 - 3活跃
nudge
MCPA typed, replayable, budget-aware programming language for LLM agents. Compiles to Python & TypeScript.
github实测增长打开来源 ↗
2GitHub 星标稳定 - 4活跃
Marketing Mix Model expert agent plugin for Claude Code - pymc-marketing v0.18.2+
github实测增长打开来源 ↗
安装
/plugin marketplace add Yakoub-ai/agent-mmm4GitHub 星标稳定 - 5活跃
rashomon
技能Measure prompt and skill improvements with blind A/B comparison.
github实测增长打开来源 ↗
安装
/plugin marketplace add shinpr/rashomon18GitHub 星标稳定 - 6活跃
agentgateway
MCPAgentGateway — independent third-party profile of a public API surface, by API Evangelist. AgentGateway is an open-source, AI-native proxy and gateway for routing, observing, and governing traffic to and from AI agents, LLM providers, and…
github实测增长打开来源 ↗
0GitHub 星标稳定 - 7活跃
evoagent-os
智能体Local-first control plane for durable, governed agent teams: DAG orchestration, approvals, memory, signed skills, trace contracts, and realtime.
github实测增长打开来源 ↗
0GitHub 星标稳定 - 8活跃
Production reliability, benchmarking, and evaluation layer for MCP agent runtimes — protocol negotiation, A2A interop, circuit breakers, connection pooling, adaptive streaming, and autoscaling, with measured throughput/latency and a…
github实测增长打开来源 ↗
3GitHub 星标稳定 - 9活跃
dsh-plugins
智能体Generic DeepSeek Harness (dsh) plugins: A2A protocol server, session storage mirror, and Langfuse observability.
github实测增长打开来源 ↗
1GitHub 星标稳定 - 10休眠
anti-lie
技能Don't make LLMs honest. Make every factual claim auditable. — An LLM Claim Auditing Layer with T1-T7 truth gradients. 98.1% business effectiveness on LiarBench v0.2.
github实测增长打开来源 ↗
安装
git clone https://github.com/lc198707/anti-lie ~/.claude/skills/anti-lie89GitHub 星标稳定 - 11活跃
Audit which Claude Code skills you actually use — surface dead installs and hallucinated invocations from your session logs.
github实测增长打开来源 ↗
安装
/plugin marketplace add sfrangulov/skill-graveyard10GitHub 星标稳定 - 12活跃
Mide si tu agente de IA cumple las reglas que le escribiste. Lee el historial local de Claude Code y devuelve un porcentaje por regla. Sin instalar nada, sin red, biblioteca estándar.
github实测增长打开来源 ↗
安装
git clone https://github.com/jleonceo/adherencia-reglas ~/.claude/skills/adherencia-reglas1GitHub 星标稳定 - 13活跃
Local-first Agentic AI Infrastructure Platform with LLM Gateway, Agent DAG Runtime, MCP Tool Hub, A2A Agent Mesh, Local RAG, Tool Sandbox and Observability.
github实测增长打开来源 ↗
1GitHub 星标稳定 - 14休眠
primal-core
智能体The reliability and interoperability layer for AI agents. A2A-native. MCP-bridged.
github实测增长打开来源 ↗
1GitHub 星标稳定 - 15休眠
An AI-powered multi-agent system that demonstrates clinical triage, OTC medication recommendations, and e-pharmacy integration for respiratory conditions. Built with modular agents that collaborate to provide safe, intelligent healthcare…
github实测增长打开来源 ↗
0GitHub 星标稳定 - 16休眠
Local open-source dev tool to debug, secure, and evaluate LLM agents. Provides static analysis, dynamic security checks, and runtime monitoring - integrates with Cursor and Claude Code.
github实测增长打开来源 ↗
安装
/plugin marketplace add cylestio/agent-inspector9GitHub 星标稳定 - 17休眠
MCP as a Judge: a behavioral MCP that strengthens AI coding assistants via explicit LLM evaluations
mcp实测增长打开来源 ↗
安装
claude mcp add mcp-as-a-judge -- uvx mcp-as-a-judge17GitHub 星标稳定 - 18活跃
8 MCP SMB products — standalone AI servers for SMBs: Guardrails, FinOps, Observability, Router, Trust Score, Memory, ThinkSecure, A2A Lite. 37 tools, TNC credits billing.
github实测增长打开来源 ↗
1GitHub 星标稳定 - 19活跃
Claude Code skills where every entry ships receipts — accuracy-gated benchmarks against baseline and placebo, rejects published
github实测增长打开来源 ↗
安装
git clone https://github.com/sjh9714/skill-receipts ~/.claude/skills/skill-receipts2GitHub 星标稳定 - 20休眠
A Claude Code skill that adds a rubric-based eval layer to any agent project. Framework-agnostic — generates rubric, test cases, judge prompt, and harness. Returns a weighted score plus a judge-leniency signal.
github实测增长打开来源 ↗
13GitHub 星标稳定 - 21休眠
A safety-first multi-agent mental health companion with real-time distress tracking, triple-layer guardrails, and evidence-based grounding techniques. Built for Kaggle × Google Agents Intensive 2025 Capstone (Agents for Good Track)
github实测增长打开来源 ↗
1GitHub 星标稳定 - 22活跃
anchor
技能Anchor — the production-grade AGENTS.md template for AI coding agents. 51 sections of battle-tested rules. Keep your agents grounded.
github实测增长打开来源 ↗
安装
git clone https://github.com/peva3/anchor ~/.claude/skills/anchor8GitHub 星标稳定 - 23休眠
Two Claude Code skills for adversarial code review (single-model PAR and multi-model MMAR with cross-critique to catch hallucinations), plus a fixture-based eval suite.
github实测增长打开来源 ↗
安装
/plugin marketplace add prime-radiant-inc/parallel-adversarial-review17GitHub 星标稳定 - 24休眠
agentops
智能体AgentOps: Multi-agent infrastructure remediation platform. A2A protocol, agent coordination, HITL approval, auto-rollback.
github实测增长打开来源 ↗
0GitHub 星标稳定 - 25活跃
Production patterns for agent orchestrators, harnesses, evals, MCP/A2A interoperability, memory, provider routing, and LLM usage metering.
github实测增长打开来源 ↗
安装
git clone https://github.com/ssheleg/agent-stack ~/.claude/skills/agent-stack2GitHub 星标稳定