ツールの説明は英語です。

LLM可観測性

LLMアプリの監視、トレース、評価、品質管理。

用途

アクティビティ

並び順

122

ツールランキング

混合順位:測定済み成長を優先します。 各値は情報源固有の期間を使用し、変化の算出には7日間で2回以上の測定が必要です。

  1. 1

    スキル
    活動中

    Claude Code skills where every entry ships receipts — accuracy-gated benchmarks against baseline and placebo, rejects published

    github測定済み成長情報源を開く ↗

    インストール git clone https://github.com/sjh9714/skill-receipts ~/.claude/skills/skill-receipts

    2GitHubスター安定
  2. 2
    活動中

    Sentry instrumentation skill for system-behavior tracking

    github測定済み成長情報源を開く ↗

    インストール git clone https://github.com/tortastudios/sentry-instrumentation ~/.claude/skills/sentry-instrumentation

    24GitHubスター安定
  3. 3

    その他
    休止中

    Multi-agent customer support system with Google ADK & Gemini 2.5 Flash Lite. Kaggle capstone demonstrating 11+ concepts. Automates 80%+ queries, <10s response time.

    github測定済み成長情報源を開く ↗

    2GitHubスター安定
  4. 4
    活動中

    Deterministic, local-only audit reports for Claude Code AI agent sessions. Rust + SQLite. Zero network calls. MCP server for agents.

    github測定済み成長情報源を開く ↗

    22GitHubスター+2 (+10.0 %)
  5. 5

    MCP
    活動中

    A typed, replayable, budget-aware programming language for LLM agents. Compiles to Python & TypeScript.

    github測定済み成長情報源を開く ↗

    2GitHubスター安定
  6. 6
    活動中

    Log Claude Code sessions to Opik, the open-source LLM observability and evaluation platform, built by Comet. Tracing, evaluation, and skills for observable AI applications.

    github測定済み成長情報源を開く ↗

    22GitHubスター+1 (+4.8 %)
  7. 7

    その他
    活動中

    AI-detector auditing toolkit: measures how often a detector flags genuine human writing, how stable its verdict is across seeds, and whether that verdict survives meaning-preserving edits. Detector-in-the-loop measurement harness. Claude…

    github測定済み成長情報源を開く ↗

    18GitHubスター安定
  8. 8

    エージェント
    休止中

    OpenTelemetry semantic conventions and instrumentation for agent provenance, derivation lineage, and acceptance criteria evaluation. Fills the Microsoft AI stack observability gap.

    github測定済み成長情報源を開く ↗

    0GitHubスター安定
  9. 9

    エージェント
    休止中

    A Multi-Agent System (MAS) evaluation framework using PydanticAI that generates and evaluates scientific paper reviews through a three-tiered assessment approach: traditional metrics, LLM-as-a-Judge, and graph-based complexity analysis.

    github測定済み成長情報源を開く ↗

    2GitHubスター安定
  10. 10
    活動中

    Multi-agent quality gate skill for Claude Code that researches, reviews, tests and challenges AI-generated work before the final answer.

    github測定済み成長情報源を開く ↗

    インストール git clone https://github.com/ma-nucho-pro/supervisor-skill-claude ~/.claude/skills/supervisor-skill-claude

    2GitHubスター安定
  11. 11

    スキル
    活動中

    Measure prompt and skill improvements with blind A/B comparison.

    github測定済み成長情報源を開く ↗

    インストール /plugin marketplace add shinpr/rashomon

    18GitHubスター安定
  12. 12

    スキル
    活動中

    Production patterns for agent orchestrators, harnesses, evals, MCP/A2A interoperability, memory, provider routing, and LLM usage metering.

    github測定済み成長情報源を開く ↗

    インストール git clone https://github.com/ssheleg/agent-stack ~/.claude/skills/agent-stack

    2GitHubスター+1 (+100.0 %)
  13. 13
    休止中

    MCP as a Judge: a behavioral MCP that strengthens AI coding assistants via explicit LLM evaluations

    mcp測定済み成長情報源を開く ↗

    インストール claude mcp add mcp-as-a-judge -- uvx mcp-as-a-judge

    17GitHubスター安定
  14. 14

    エージェント
    活動中

    Multi-agent SRE on-call investigator that auto-triages Slack/Discord infrastructure alerts via AWS Bedrock AgentCore, fanning out to specialized agents (CloudWatch, EKS, Slack/Discord scanners) for parallel investigation.

    github測定済み成長情報源を開く ↗

    3GitHubスター安定
  15. 15

    スキル
    活動中

    Deletion-first Agent Skill for removing test bloat, verification theater, and speculative fallbacks while preserving behavior.

    github測定済み成長情報源を開く ↗

    インストール git clone https://github.com/MrZoyo/deslop-GPT ~/.claude/skills/deslop-GPT

    76GitHubスター+48 (+171.4 %)
  16. 16
    活動中

    Mini Program Engineering Skill Suite is an Agent Skill suite for evidence‑first mini‑program development. It helps agents bring WeChat or other mini‑program projects from vague intent to reliable engineering work: project intake, product…

    github測定済み成長情報源を開く ↗

    インストール git clone https://github.com/NocodeMrLi/mini-program-engineering-skill-suite ~/.claude/skills/mini-program-engineering-skill-suite

    43GitHubスター+4 (+10.3 %)
  17. 17
    活動中

    Build tested agent skills and govern their lifecycle through a user-defined marketplace: evidence, discovery, updates, rollback, quarantine, and 17-platform distribution.

    github推定モメンタム情報源を開く ↗

    インストール git clone https://github.com/FrancyJGLisboa/agent-skills-platform ~/.claude/skills/agent-skills-platform

    2 376GitHubスター

学習・参考リソース

情報源間で正規化した測定済み成長順です。 これらは個別に参照でき、主要ランキングには含まれません。

  1. 1

    trigger_tree

    スキル
    活動中情報源を開く ↗

    Documentation-discovery telemetry for Claude Code — heat/cold maps, health grade, evidence-backed router fixes. 100% local, zero tokens. /tt

    github

    インストール /plugin marketplace add Hedde/trigger_tree

    リソース測定済み成長
    14GitHubスター+1 (+7.7 %)
  2. 2

    production-ai-engineering

    エージェント
    活動中情報源を開く ↗

    Handbook técnico aberto sobre engenharia de IA em produção: ML tradicional, LLMs, RAG, agentes, segurança, observabilidade, FinOps e deployment.

    githubリソース測定済み成長
    1GitHubスター安定
  3. 3
    活動中情報源を開く ↗

    50+ curated LLM observability tools PLUS 26 Agent Skills (several with runnable, unit-tested scripts) to build, evaluate, debug, secure & monitor reliable LLM apps. Tracing, evals, guardrails, LLMOps.

    github

    インストール git clone https://github.com/ContextJet-ai/awesome-llm-observability ~/.claude/skills/awesome-llm-observability

    リソース測定済み成長
    33GitHubスター+3 (+10.0 %)
  4. 4

    Agentic_AI_Engineer

    エージェント
    活動中情報源を開く ↗

    My complete journey to becoming an Agentic AI Engineer through structured learning, projects, experiments, and production-ready implementations of modern AI systems.

    githubリソース測定済み成長
    18GitHubスター安定
  5. 5

    tunelab

    スキル
    活動中情報源を開く ↗

    Claude Code plugin for LLM fine-tuning, distillation, and evaluation — decide whether you need fine-tuning at all, distill your LLM logs into small local models (MLX/LoRA), evaluate with held-out discipline, and learn the why at every step.

    github

    インストール /plugin marketplace add rchaz/tunelab

    リソース測定済み成長
    6GitHubスター安定