ツールの説明は英語です。
LLM可観測性
LLMアプリの監視、トレース、評価、品質管理。
用途
アクティビティ
並び順
122
ツールランキング
Ranked by known GitHub stars, highest first; archived repositories come after maintained repositories. 各値は情報源固有の期間を使用し、変化の算出には7日間で2回以上の測定が必要です。
- 1活動中
cap-evolve
MCPOptimize any AI agent’s skills, tools/MCP, and prompts against your own evals.
github推定モメンタム情報源を開く ↗
47GitHubスター— - 2活動中
Mini Program Engineering Skill Suite is an Agent Skill suite for evidence‑first mini‑program development. It helps agents bring WeChat or other mini‑program projects from vague intent to reliable engineering work: project intake, product…
github測定済み成長情報源を開く ↗
インストール
git clone https://github.com/NocodeMrLi/mini-program-engineering-skill-suite ~/.claude/skills/mini-program-engineering-skill-suite43GitHubスター+4 (+10.3 %) - 3活動中
claudestat
スキルReal-time execution trace and cost intelligence for Claude Code
github測定済み成長情報源を開く ↗
インストール
/plugin marketplace add DeibyGS/claudestat34GitHubスター安定 - 4活動中
Make Claude Opus 4.8 behave like Claude Fable 5 — doctrine output style, drift-catching hooks, and an eval loop against golden Fable transcripts. Claude Code plugin.
github測定済み成長情報源を開く ↗
インストール
/plugin marketplace add rennf93/opus-fable-playbook34GitHubスター+1 (+3.0 %) - 5休止中
🔍 AI observability skill for Claude Code. Debug LangChain/LangGraph agents by fetching execution traces from LangSmith Studio directly in your terminal.
github測定済み成長情報源を開く ↗
インストール
git clone https://github.com/OthmanAdi/langsmith-fetch-skill ~/.claude/skills/langsmith-fetch-skill27GitHubスター安定 - 6休止中
astragraph
エージェントPolicy-enforced observability and fail-closed guardrails for MCP/A2A multi-agent systems.
github測定済み成長情報源を開く ↗
26GitHubスター安定 - 7活動中
Sentry instrumentation skill for system-behavior tracking
github測定済み成長情報源を開く ↗
インストール
git clone https://github.com/tortastudios/sentry-instrumentation ~/.claude/skills/sentry-instrumentation24GitHubスター安定 - 8活動中
Deterministic, local-only audit reports for Claude Code AI agent sessions. Rust + SQLite. Zero network calls. MCP server for agents.
github測定済み成長情報源を開く ↗
22GitHubスター+2 (+10.0 %) - 9活動中
Log Claude Code sessions to Opik, the open-source LLM observability and evaluation platform, built by Comet. Tracing, evaluation, and skills for observable AI applications.
github測定済み成長情報源を開く ↗
22GitHubスター+1 (+4.8 %) - 10活動中
rashomon
スキルMeasure prompt and skill improvements with blind A/B comparison.
github測定済み成長情報源を開く ↗
インストール
/plugin marketplace add shinpr/rashomon18GitHubスター安定 - 11活動中
untell
その他AI-detector auditing toolkit: measures how often a detector flags genuine human writing, how stable its verdict is across seeds, and whether that verdict survives meaning-preserving edits. Detector-in-the-loop measurement harness. Claude…
github測定済み成長情報源を開く ↗
18GitHubスター安定 - 12休止中
MCP as a Judge: a behavioral MCP that strengthens AI coding assistants via explicit LLM evaluations
mcp測定済み成長情報源を開く ↗
インストール
claude mcp add mcp-as-a-judge -- uvx mcp-as-a-judge17GitHubスター安定 - 13活動中
Two Claude Code skills for adversarial code review (single-model PAR and multi-model MMAR with cross-critique to catch hallucinations), plus a fixture-based eval suite.
github測定済み成長情報源を開く ↗
インストール
/plugin marketplace add prime-radiant-inc/parallel-adversarial-review17GitHubスター安定 - 14活動中
Make Claude write clearly, for everyone.
github測定済み成長情報源を開く ↗
インストール
/plugin marketplace add stefanobaghino/simple-output-styles16GitHubスター安定 - 15活動中
adl-cli
その他A command-line tool to scaffold and manage enterprise-ready AI Agents powered by the A2A (Agent-to-Agent) protocol
github測定済み成長情報源を開く ↗
14GitHubスター+1 (+7.7 %) - 16活動中
Claude Code plugin marketplace — agentic-engineering (spec-driven shape→decide→execute→measure→eval with adversarial review) + github-keeper (audit/elevate READMEs and make a repo open-source-ready).
github測定済み成長情報源を開く ↗
インストール
/plugin marketplace add GiustoPiedimonte/agentic-engineering-marketplace13GitHubスター安定 - 17休止中
eval-layer
その他A Claude Code skill that adds a rubric-based eval layer to any agent project. Framework-agnostic — generates rubric, test cases, judge prompt, and harness. Returns a weighted score plus a judge-leniency signal.
github測定済み成長情報源を開く ↗
13GitHubスター安定 - 18活動中
Query Spanlens LLM observability from Cursor, Claude Desktop, or Continue via MCP.
mcp測定済み成長情報源を開く ↗
インストール
claude mcp add mcp-server -- npx @spanlens/mcp-server12GitHubスター安定 - 19活動中
Audit which Claude Code skills you actually use — surface dead installs and hallucinated invocations from your session logs.
github測定済み成長情報源を開く ↗
インストール
/plugin marketplace add sfrangulov/skill-graveyard10GitHubスター+1 (+11.1 %) - 20活動中
bakeoff
スキルTurn one decision into a judged tournament of solutions, then pick the best — a Claude Code skill that generates candidates, auto-derives the rubric, judges independently, and returns a defensible winner.
github測定済み成長情報源を開く ↗
インストール
/plugin marketplace add CoriChui/bakeoff10GitHubスター安定 - 21活動中
galdor
その他A Go-native framework for LLM agents, with OpenTelemetry observability built in.
github測定済み成長情報源を開く ↗
10GitHubスター安定 - 22活動中
ai-dev-stack
スキルProduction-grade AI coding rules for Cursor and Claude Code. 15 rules + 9 doc templates + skills + agents + MCP setup. Drop into any project.
github測定済み成長情報源を開く ↗
インストール
git clone https://github.com/aiagentwithdhruv/ai-dev-stack ~/.claude/skills/ai-dev-stack10GitHubスター+1 (+11.1 %)
学習・参考リソース
Ranked by known GitHub stars, highest first; archived repositories come after maintained repositories. これらは個別に参照でき、主要ランキングには含まれません。
- 1活動中情報源を開く ↗
50+ curated LLM observability tools PLUS 26 Agent Skills (several with runnable, unit-tested scripts) to build, evaluate, debug, secure & monitor reliable LLM apps. Tracing, evals, guardrails, LLMOps.
githubインストール
リソース測定済み成長git clone https://github.com/ContextJet-ai/awesome-llm-observability ~/.claude/skills/awesome-llm-observability33GitHubスター+3 (+10.0 %) - 2活動中情報源を開く ↗
Agentic_AI_Engineer
エージェントMy complete journey to becoming an Agentic AI Engineer through structured learning, projects, experiments, and production-ready implementations of modern AI systems.
githubリソース測定済み成長18GitHubスター安定 - 3活動中情報源を開く ↗
trigger_tree
スキルDocumentation-discovery telemetry for Claude Code — heat/cold maps, health grade, evidence-backed router fixes. 100% local, zero tokens. /tt
githubインストール
リソース測定済み成長/plugin marketplace add Hedde/trigger_tree14GitHubスター+1 (+7.7 %)