ツールの説明は英語です。
LLM可観測性
LLMアプリの監視、トレース、評価、品質管理。
用途
アクティビティ
並び順
119
ツールランキング
情報源間で正規化した測定済み成長順です。 各値は情報源固有の期間を使用し、変化の算出には7日間で2回以上の測定が必要です。
- 1休止中
Two Claude Code skills for adversarial code review (single-model PAR and multi-model MMAR with cross-critique to catch hallucinations), plus a fixture-based eval suite.
github測定済み成長情報源を開く ↗
インストール
/plugin marketplace add prime-radiant-inc/parallel-adversarial-review17GitHubスター安定 - 2休止中
agentops
エージェントAgentOps: Multi-agent infrastructure remediation platform. A2A protocol, agent coordination, HITL approval, auto-rollback.
github測定済み成長情報源を開く ↗
0GitHubスター安定 - 3活動中
agent-stack
スキルProduction patterns for agent orchestrators, harnesses, evals, MCP/A2A interoperability, memory, provider routing, and LLM usage metering.
github測定済み成長情報源を開く ↗
インストール
git clone https://github.com/ssheleg/agent-stack ~/.claude/skills/agent-stack2GitHubスター安定 - 4休止中
agenttap
エージェントReal-time debugging proxy for Agent2Agent (A2A) multi-agent systems
github測定済み成長情報源を開く ↗
3GitHubスター安定 - 5活動中
Multi-tier framework for evaluating AI agent skills with quality gates, semantic overlap detection, synthetic evaluation dataset generation, and live agent evaluation that measures how skills affect agent behavior.
github測定済み成長情報源を開く ↗
インストール
git clone https://github.com/NVIDIA/SkillEvaluator ~/.claude/skills/SkillEvaluator405GitHubスター+42 (+11.6 %) - 6活動中
Architecture-first Python scaffold for an auditable multi-agent corporate credit desk using A2A, MCP, deterministic credit policies, model routing, and OpenTelemetry.
github測定済み成長情報源を開く ↗
0GitHubスター安定 - 7活動中
AgentX python SDK. Build multi-agent AI workforce. Run evaluation. Trace your agent. Full Observability.
github測定済み成長情報源を開く ↗
68GitHubスター安定 - 8活動中
AI Guardian
MCPGoverned local-LLM observability: model policy, prompt scanner, capture proxy, 21 tools.
mcp測定済み成長情報源を開く ↗
インストール
claude mcp add ai-guardian -- uvx ai-guardian-aiops0GitHubスター安定 - 9休止中
A self-improving harness router for Claude Code.
github測定済み成長情報源を開く ↗
インストール
/plugin marketplace add SeongwoongCho/adaptive-harness8GitHubスター安定 - 10活動中
SkillForge
スキルA skill creator that proves its skills work. Evidence-driven skill creation for Claude Code and Codex: baseline-tested generation, per-skill regression evals, ecosystem doctor, cross-runtime compile, and an opt-in proactive advisor.
github測定済み成長情報源を開く ↗
インストール
git clone https://github.com/tripleyak/SkillForge ~/.claude/skills/SkillForge885GitHubスター+1 (+0.11 %) - 11活動中
Open benchmark for Claude Code SEO skills — real headless execution against fixture sites with planted-defect answer keys. Deterministic scoring, pre-registered rubric.
github測定済み成長情報源を開く ↗
インストール
git clone https://github.com/aleclindz/seo-skill-bench ~/.claude/skills/seo-skill-bench51GitHubスター+1 (+2.0 %) - 12休止中
CustoFlow
その他Multi-agent customer support system with Google ADK & Gemini 2.5 Flash Lite. Kaggle capstone demonstrating 11+ concepts. Automates 80%+ queries, <10s response time.
github測定済み成長情報源を開く ↗
2GitHubスター安定 - 13活動中
74 open-source Agent Skills for Claude Code and Codex: AI SEO, AEO and GEO, code review with an A-F ship grade, CI gates, AI evals, design systems, conversion copy, Instagram growth, iOS and Android app shipping, creator rights, and…
github測定済み成長情報源を開く ↗
130GitHubスター+1 (+0.78 %) - 14活動中
nora
MCPOpen-source, self-hosted control plane for OpenClaw and Hermes AI-agent fleets on Docker/Kubernetes — REST, CLI, and MCP.
github測定済み成長情報源を開く ↗
50GitHubスター+1 (+2.0 %) - 15活動中
memroos
エージェントMemroOS / memroos: memory OS and governance layer for AI agents, agent workflows, dispatch, proof, and context continuity.
github測定済み成長情報源を開く ↗
7GitHubスター安定 - 16活動中
a2a-otel-kit
MCPVendor-neutral OpenTelemetry tracing for A2A agents and MCP services, with W3C context propagation and privacy-safe telemetry.
github測定済み成長情報源を開く ↗
1GitHubスター安定 - 17活動中
tulip-agents
エージェントThe agent framework where the model never holds the trigger — every consequential action clears your policy first, waits for a human when it matters, and lands on a record you can verify. Build on it, or put it around the agent you already…
github測定済み成長情報源を開く ↗
1GitHubスター安定 - 18休止中
agentanvil
エージェントContract-based testing for LLM agents: hybrid evaluation, multi-agent + A2A, deterministic replay.
github測定済み成長情報源を開く ↗
0GitHubスター安定 - 19活動中
Make Claude Opus 4.8 behave like Claude Fable 5 — doctrine output style, drift-catching hooks, and an eval loop against golden Fable transcripts. Claude Code plugin.
github測定済み成長情報源を開く ↗
インストール
/plugin marketplace add rennf93/opus-fable-playbook34GitHubスター+1 (+3.0 %) - 20活動中
kubesphere
スキルThe container platform tailored for Kubernetes multi-cloud, datacenter, and edge management ⎈ 🖥 ☁️
github測定済み成長情報源を開く ↗
インストール
git clone https://github.com/kubesphere/kubesphere ~/.claude/skills/kubesphere17 035GitHubスター安定 - 21活動中
🪢 Langfuse documentation -- Langfuse is the open source LLM Engineering Platform. Observability, evals, prompt management, playground and metrics to debug and improve LLM apps
github測定済み成長情報源を開く ↗
インストール
git clone https://github.com/langfuse/langfuse-docs ~/.claude/skills/langfuse-docs240GitHubスター+3 (+1.3 %) - 22活動中
agent-kernel
エージェントThe Operating System for Scalable Enterprise AI Agents - Run, orchestrate, and deploy Compliant Enterprise AI Agents at scale across frameworks, without lock-in, rewrites or fragile glue code. Native support for MCP, A2A. Interface with all…
github測定済み成長情報源を開く ↗
166GitHubスター+20 (+13.7 %) - 23活動中
claudestat
スキルReal-time execution trace and cost intelligence for Claude Code
github測定済み成長情報源を開く ↗
インストール
/plugin marketplace add DeibyGS/claudestat34GitHubスター安定 - 24活動中
pyxen
エージェントA lightweight Python library that decouples agentic runtime from applications it builds
github測定済み成長情報源を開く ↗
1GitHubスター安定 - 25活動中
craft-skills
スキルResearch-backed, eval-driven skills for AI agents
github測定済み成長情報源を開く ↗
インストール
git clone https://github.com/ZSeven-W/craft-skills ~/.claude/skills/craft-skills148GitHubスター+10 (+7.2 %)