ツールの説明は英語です。

LLM可観測性

LLMアプリの監視、トレース、評価、品質管理。

用途

アクティビティ

並び順

122

ツールランキング

情報源間で正規化した測定済み成長順です。 各値は情報源固有の期間を使用し、変化の算出には7日間で2回以上の測定が必要です。

  1. 1
    活動中

    MCP server for Langfuse LLM observability — trace and observation analysis.

    mcp測定済み成長情報源を開く ↗

    インストール claude mcp add langfuse -- npx langfuse-observability-mcp-server

    78GitHubスター安定
  2. 2

    その他
    活動中

    🔑 HomeStream · 家园·流 — 零成本自托管多Agent协作框架,通往AI世界的那把钥匙 | Zero-cost self-hosted multi-agent framework — The key to AI world

    github測定済み成長情報源を開く ↗

    0GitHubスター安定
  3. 3

    その他
    活動中

    AgentX python SDK. Build multi-agent AI workforce. Run evaluation. Trace your agent. Full Observability.

    github測定済み成長情報源を開く ↗

    68GitHubスター安定
  4. 4

    エージェント
    休止中

    A Python proof-of-concept for tracing multi-turn Agent-to-Agent (A2A) conversations as a single unified MLflow trace for LLM observability and evaluation.

    github測定済み成長情報源を開く ↗

    0GitHubスター安定
  5. 5
    活動中

    Make Claude Opus 4.8 behave like Claude Fable 5 — doctrine output style, drift-catching hooks, and an eval loop against golden Fable transcripts. Claude Code plugin.

    github測定済み成長情報源を開く ↗

    インストール /plugin marketplace add rennf93/opus-fable-playbook

    33GitHubスター安定
  6. 6
    活動中

    AgentGateway — independent third-party profile of a public API surface, by API Evangelist. AgentGateway is an open-source, AI-native proxy and gateway for routing, observing, and governing traffic to and from AI agents, LLM providers, and…

    github測定済み成長情報源を開く ↗

    0GitHubスター安定
  7. 7
    休止中

    🔍 AI observability skill for Claude Code. Debug LangChain/LangGraph agents by fetching execution traces from LangSmith Studio directly in your terminal.

    github測定済み成長情報源を開く ↗

    インストール git clone https://github.com/OthmanAdi/langsmith-fetch-skill ~/.claude/skills/langsmith-fetch-skill

    27GitHubスター安定
  8. 8

    エージェント
    休止中

    Policy-enforced observability and fail-closed guardrails for MCP/A2A multi-agent systems.

    github測定済み成長情報源を開く ↗

    26GitHubスター安定
  9. 9

    エージェント
    休止中

    OpenTelemetry semantic conventions and instrumentation for agent provenance, derivation lineage, and acceptance criteria evaluation. Fills the Microsoft AI stack observability gap.

    github測定済み成長情報源を開く ↗

    0GitHubスター安定
  10. 10
    活動中

    Sentry instrumentation skill for system-behavior tracking

    github測定済み成長情報源を開く ↗

    インストール git clone https://github.com/tortastudios/sentry-instrumentation ~/.claude/skills/sentry-instrumentation

    24GitHubスター安定
  11. 11
    活動中

    Governed local-LLM observability: model policy, prompt scanner, capture proxy, 21 tools.

    mcp測定済み成長情報源を開く ↗

    インストール claude mcp add ai-guardian -- uvx ai-guardian-aiops

    0GitHubスター安定
  12. 12

    その他
    活動中

    AI-detector auditing toolkit: measures how often a detector flags genuine human writing, how stable its verdict is across seeds, and whether that verdict survives meaning-preserving edits. Detector-in-the-loop measurement harness. Claude…

    github測定済み成長情報源を開く ↗

    18GitHubスター安定
  13. 13

    スキル
    活動中

    Measure prompt and skill improvements with blind A/B comparison.

    github測定済み成長情報源を開く ↗

    インストール /plugin marketplace add shinpr/rashomon

    18GitHubスター安定
  14. 14
    休止中

    An AI-powered multi-agent system that demonstrates clinical triage, OTC medication recommendations, and e-pharmacy integration for respiratory conditions. Built with modular agents that collaborate to provide safe, intelligent healthcare…

    github測定済み成長情報源を開く ↗

    0GitHubスター安定
  15. 15
    活動中

    Two Claude Code skills for adversarial code review (single-model PAR and multi-model MMAR with cross-critique to catch hallucinations), plus a fixture-based eval suite.

    github測定済み成長情報源を開く ↗

    インストール /plugin marketplace add prime-radiant-inc/parallel-adversarial-review

    17GitHubスター安定
  16. 16
    休止中

    MCP as a Judge: a behavioral MCP that strengthens AI coding assistants via explicit LLM evaluations

    mcp測定済み成長情報源を開く ↗

    インストール claude mcp add mcp-as-a-judge -- uvx mcp-as-a-judge

    17GitHubスター安定
  17. 17
    活動中

    Make Claude write clearly, for everyone.

    github測定済み成長情報源を開く ↗

    インストール /plugin marketplace add stefanobaghino/simple-output-styles

    16GitHubスター安定
  18. 18

    エージェント
    休止中

    AgentOps: Multi-agent infrastructure remediation platform. A2A protocol, agent coordination, HITL approval, auto-rollback.

    github測定済み成長情報源を開く ↗

    0GitHubスター安定
  19. 19
    活動中

    Claude Code plugin marketplace — agentic-engineering (spec-driven shape→decide→execute→measure→eval with adversarial review) + github-keeper (audit/elevate READMEs and make a repo open-source-ready).

    github測定済み成長情報源を開く ↗

    インストール /plugin marketplace add GiustoPiedimonte/agentic-engineering-marketplace

    13GitHubスター安定
  20. 20

    スキル
    活動中

    Glanceable Claude Code and Codex state in your terminal tabs: white=idle, blue=working, orange=waiting. Multi-terminal (iTerm2, WezTerm, AI Power Term), a live status bar with Anthropic usage limits, /sfl and /nil window save-and-restore,…

    github測定済み成長情報源を開く ↗

    インストール git clone https://github.com/wasulajr/headsup ~/.claude/skills/headsup

    1GitHubスター安定
  21. 21

    その他
    休止中

    A Claude Code skill that adds a rubric-based eval layer to any agent project. Framework-agnostic — generates rubric, test cases, judge prompt, and harness. Returns a weighted score plus a judge-leniency signal.

    github測定済み成長情報源を開く ↗

    13GitHubスター安定
  22. 22
    活動中

    Query Spanlens LLM observability from Cursor, Claude Desktop, or Continue via MCP.

    mcp測定済み成長情報源を開く ↗

    インストール claude mcp add mcp-server -- npx @spanlens/mcp-server

    12GitHubスター安定
  23. 23

    エージェント
    活動中

    Provide clear documentation for AgentStack’s MCP protocol, plugins, and ecosystem API with usage examples and tool references.

    github測定済み成長情報源を開く ↗

    0GitHubスター安定
  24. 24

    その他
    活動中

    Terminal-first validation harness for reviewing PRDs, RFCs, strategy docs, and experiment plans.

    github測定済み成長情報源を開く ↗

    1GitHubスター安定
  25. 25

    スキル
    活動中

    Turn one decision into a judged tournament of solutions, then pick the best — a Claude Code skill that generates candidates, auto-derives the rubric, judges independently, and returns a defensible winner.

    github測定済み成長情報源を開く ↗

    インストール /plugin marketplace add CoriChui/bakeoff

    10GitHubスター安定