工具说明为英文。
LLM 可观测性
LLM 应用的监控、追踪、评估和质量管理。
用途
活跃度
排序方式
121
工具排名
按实测增长排序,并在不同来源间归一化。 原始值保留各自的数据窗口;计算变化需要 7 天内至少两次测量。
- 13GitHub 星标稳定
- 2活跃
axiom
技能Axiom is a curated marketplace of shared plugins for Claude Code and Codex.
github实测增长打开来源 ↗
安装
git clone https://github.com/netopsengineer/axiom ~/.claude/skills/axiom5GitHub 星标稳定 - 3活跃
Claude Code plugin marketplace — agentic-engineering (spec-driven shape→decide→execute→measure→eval with adversarial review) + github-keeper (audit/elevate READMEs and make a repo open-source-ready).
github实测增长打开来源 ↗
安装
/plugin marketplace add GiustoPiedimonte/agentic-engineering-marketplace13GitHub 星标稳定 - 4活跃
The container platform tailored for Kubernetes multi-cloud, datacenter, and edge management ⎈ 🖥 ☁️
github实测增长打开来源 ↗
安装
git clone https://github.com/kubesphere/kubesphere ~/.claude/skills/kubesphere17 035GitHub 星标+2 (+0.01 %) - 5休眠
kiboserve
MCPOpen-source Python framework to deploy AI agents via HTTP, A2A, and MCP with built-in observability
github实测增长打开来源 ↗
5GitHub 星标稳定 - 6休眠
A Claude Code skill that adds a rubric-based eval layer to any agent project. Framework-agnostic — generates rubric, test cases, judge prompt, and harness. Returns a weighted score plus a judge-leniency signal.
github实测增长打开来源 ↗
13GitHub 星标稳定 - 7活跃
Amazon Bedrock AgentCore enterprise platform accelerator with AWS CDK, Terraform organization guardrails, MCP/A2A agents, security, memory, and observability.
github实测增长打开来源 ↗
6GitHub 星标+2 (+50.0 %) - 8活跃
AgentStack
智能体Provide clear documentation for AgentStack’s MCP protocol, plugins, and ecosystem API with usage examples and tool references.
github实测增长打开来源 ↗
0GitHub 星标稳定 - 9活跃
memroos
智能体MemroOS / memroos: memory OS and governance layer for AI agents, agent workflows, dispatch, proof, and context continuity.
github实测增长打开来源 ↗
7GitHub 星标稳定 - 10活跃
Production-grade AI coding rules for Cursor and Claude Code. 15 rules + 9 doc templates + skills + agents + MCP setup. Drop into any project.
github实测增长打开来源 ↗
安装
git clone https://github.com/aiagentwithdhruv/ai-dev-stack ~/.claude/skills/ai-dev-stack10GitHub 星标+1 (+11.1 %) - 11活跃
Multi-tier framework for evaluating AI agent skills with quality gates, semantic overlap detection, synthetic evaluation dataset generation, and live agent evaluation that measures how skills affect agent behavior.
github实测增长打开来源 ↗
安装
git clone https://github.com/NVIDIA/SkillEvaluator ~/.claude/skills/SkillEvaluator399GitHub 星标+51 (+14.7 %) - 12休眠
A self-improving harness router for Claude Code.
github实测增长打开来源 ↗
安装
/plugin marketplace add SeongwoongCho/adaptive-harness8GitHub 星标稳定 - 13活跃
Audit which Claude Code skills you actually use — surface dead installs and hallucinated invocations from your session logs.
github实测增长打开来源 ↗
安装
/plugin marketplace add sfrangulov/skill-graveyard10GitHub 星标+1 (+11.1 %) - 14活跃
anchor
技能Anchor — the production-grade AGENTS.md template for AI coding agents. 51 sections of battle-tested rules. Keep your agents grounded.
github实测增长打开来源 ↗
安装
git clone https://github.com/peva3/anchor ~/.claude/skills/anchor8GitHub 星标稳定 - 15活跃
bakeoff
技能Turn one decision into a judged tournament of solutions, then pick the best — a Claude Code skill that generates candidates, auto-derives the rubric, judges independently, and returns a defensible winner.
github实测增长打开来源 ↗
安装
/plugin marketplace add CoriChui/bakeoff10GitHub 星标稳定 - 16休眠
Local open-source dev tool to debug, secure, and evaluate LLM agents. Provides static analysis, dynamic security checks, and runtime monitoring - integrates with Cursor and Claude Code.
github实测增长打开来源 ↗
安装
/plugin marketplace add cylestio/agent-inspector9GitHub 星标稳定 - 17活跃
galdor
其他A Go-native framework for LLM agents, with OpenTelemetry observability built in.
github实测增长打开来源 ↗
10GitHub 星标稳定 - 18活跃
Architecture-first Python scaffold for an auditable multi-agent corporate credit desk using A2A, MCP, deterministic credit policies, model routing, and OpenTelemetry.
github实测增长打开来源 ↗
0GitHub 星标稳定 - 19活跃
Dashboard for monitoring claude code sessions.
github实测增长打开来源 ↗
安装
/plugin marketplace add JayantDevkar/claude-code-karma323GitHub 星标+2 (+0.62 %) - 20休眠
agentanvil
智能体Contract-based testing for LLM agents: hybrid evaluation, multi-agent + A2A, deterministic replay.
github实测增长打开来源 ↗
0GitHub 星标稳定 - 21活跃
aura
MCPAURA is a production-tested SRE agent platform you can deploy in minutes. AURA handles the guardrails, APIs, state management, streaming, and failure handling required to put AI to work safely on production infrastructure.
github实测增长打开来源 ↗
291GitHub 星标+33 (+12.8 %) - 22活跃
Marketing Mix Model expert agent plugin for Claude Code - pymc-marketing v0.18.2+
github实测增长打开来源 ↗
安装
/plugin marketplace add Yakoub-ai/agent-mmm4GitHub 星标稳定 - 23休眠
A Python proof-of-concept for tracing multi-turn Agent-to-Agent (A2A) conversations as a single unified MLflow trace for LLM observability and evaluation.
github实测增长打开来源 ↗
0GitHub 星标稳定 - 24活跃
🪢 Langfuse documentation -- Langfuse is the open source LLM Engineering Platform. Observability, evals, prompt management, playground and metrics to debug and improve LLM apps
github实测增长打开来源 ↗
安装
git clone https://github.com/langfuse/langfuse-docs ~/.claude/skills/langfuse-docs239GitHub 星标+3 (+1.3 %) - 25活跃
green-agent
智能体A2A green-agent orchestrator for evaluating agents on the AppWorld benchmark, built on the AgentBeats SDK
github实测增长打开来源 ↗
0GitHub 星标稳定