Le descrizioni degli strumenti sono in inglese.

Osservabilità LLM

Monitoraggio, tracce, valutazione e qualità delle applicazioni LLM.

Utilizzo

Attività

Ordina per

121

Classifica degli strumenti

Classifica per crescita misurata e normalizzata tra le fonti. I valori grezzi mantengono la propria finestra; una variazione richiede almeno due rilevazioni in 7 giorni.

  1. 1

    Skill
    Dormiente

    Don't make LLMs honest. Make every factual claim auditable. — An LLM Claim Auditing Layer with T1-T7 truth gradients. 98.1% business effectiveness on LiarBench v0.2.

    githubcrescita misurataApri fonte ↗

    Installa git clone https://github.com/lc198707/anti-lie ~/.claude/skills/anti-lie

    89stelle GitHubstabile
  2. 2

    Skill
    Attivo

    Production patterns for agent orchestrators, harnesses, evals, MCP/A2A interoperability, memory, provider routing, and LLM usage metering.

    githubcrescita misurataApri fonte ↗

    Installa git clone https://github.com/ssheleg/agent-stack ~/.claude/skills/agent-stack

    2stelle GitHub+1 (+100.0 %)
  3. 3

    Altro
    Attivo

    AI-detector auditing toolkit: measures how often a detector flags genuine human writing, how stable its verdict is across seeds, and whether that verdict survives meaning-preserving edits. Detector-in-the-loop measurement harness. Claude…

    githubcrescita misurataApri fonte ↗

    18stelle GitHubstabile
  4. 4
    Attivo

    Log Claude Code sessions to Opik, the open-source LLM observability and evaluation platform, built by Comet. Tracing, evaluation, and skills for observable AI applications.

    githubcrescita misurataApri fonte ↗

    22stelle GitHub+1 (+4.8 %)
  5. 5
    Attivo

    Deterministic, local-only audit reports for Claude Code AI agent sessions. Rust + SQLite. Zero network calls. MCP server for agents.

    githubcrescita misurataApri fonte ↗

    22stelle GitHub+2 (+10.0 %)
  6. 6

    Skill
    Attivo

    Deletion-first Agent Skill for removing test bloat, verification theater, and speculative fallbacks while preserving behavior.

    githubcrescita misurataApri fonte ↗

    Installa git clone https://github.com/MrZoyo/deslop-GPT ~/.claude/skills/deslop-GPT

    87stelle GitHub+59 (+210.7 %)
  7. 7

    Agente
    Dormiente

    Real-time debugging proxy for Agent2Agent (A2A) multi-agent systems

    githubcrescita misurataApri fonte ↗

    3stelle GitHubstabile
  8. 8

    Agente
    Dormiente

    Contract-based testing for LLM agents: hybrid evaluation, multi-agent + A2A, deterministic replay.

    githubcrescita misurataApri fonte ↗

    0stelle GitHubstabile
  9. 9
    Attivo

    Sentry instrumentation skill for system-behavior tracking

    githubcrescita misurataApri fonte ↗

    Installa git clone https://github.com/tortastudios/sentry-instrumentation ~/.claude/skills/sentry-instrumentation

    24stelle GitHubstabile
  10. 10
    Dormiente

    A self-improving harness router for Claude Code.

    githubcrescita misurataApri fonte ↗

    Installa /plugin marketplace add SeongwoongCho/adaptive-harness

    8stelle GitHubstabile
  11. 11
    Attivo

    MCP server for Langfuse LLM observability — trace and observation analysis.

    mcpcrescita misurataApri fonte ↗

    Installa claude mcp add langfuse -- npx langfuse-observability-mcp-server

    80stelle GitHub+2 (+2.6 %)
  12. 12

    Skill
    Attivo

    Axiom is a curated marketplace of shared plugins for Claude Code and Codex.

    githubcrescita misurataApri fonte ↗

    Installa git clone https://github.com/netopsengineer/axiom ~/.claude/skills/axiom

    5stelle GitHubstabile
  13. 13

    Agente
    Dormiente

    Policy-enforced observability and fail-closed guardrails for MCP/A2A multi-agent systems.

    githubcrescita misurataApri fonte ↗

    26stelle GitHubstabile
  14. 14
    Attivo

    A Go framework for building production agent systems with graph workflows, tools, memory, A2A, AG-UI, MCP, evaluation, and observability.

    githubcrescita misurataApri fonte ↗

    1 763stelle GitHub+15 (+0.86 %)
  15. 15

    Agente
    Attivo

    面向长程研发任务的 Java Agent Harness,支持 A2A 跨语言协作、可恢复执行、上下文工程与 Eval 驱动开发。

    githubcrescita misurataApri fonte ↗

    0stelle GitHubstabile
  16. 16
    Dormiente

    A safety-first multi-agent mental health companion with real-time distress tracking, triple-layer guardrails, and evidence-based grounding techniques. Built for Kaggle × Google Agents Intensive 2025 Capstone (Agents for Good Track)

    githubcrescita misurataApri fonte ↗

    1stelle GitHubstabile
  17. 17

    Skill
    Attivo

    OpenJudge: A Unified Framework for Holistic Evaluation and Quality Rewards

    githubcrescita misurataApri fonte ↗

    Installa git clone https://github.com/agentscope-ai/OpenJudge ~/.claude/skills/OpenJudge

    816stelle GitHub+11 (+1.4 %)
  18. 18

    Altro
    Dormiente

    Official Python SDK for GT8004 — AI agent observability with MCP, A2A, x402 payment tracking. FastAPI, Flask, FastMCP middleware included.

    githubcrescita misurataApri fonte ↗

    1stelle GitHubstabile
  19. 19
    Dormiente

    🔍 AI observability skill for Claude Code. Debug LangChain/LangGraph agents by fetching execution traces from LangSmith Studio directly in your terminal.

    githubcrescita misurataApri fonte ↗

    Installa git clone https://github.com/OthmanAdi/langsmith-fetch-skill ~/.claude/skills/langsmith-fetch-skill

    27stelle GitHubstabile
  20. 20

    Skill
    Attivo

    local-first analytics for AI agent skills

    githubcrescita misurataApri fonte ↗

    Installa git clone https://github.com/crafter-station/skill-kit ~/.claude/skills/skill-kit

    77stelle GitHub+1 (+1.3 %)
  21. 21
    Attivo

    8 MCP SMB products — standalone AI servers for SMBs: Guardrails, FinOps, Observability, Router, Trust Score, Memory, ThinkSecure, A2A Lite. 37 tools, TNC credits billing.

    githubcrescita misurataApri fonte ↗

    1stelle GitHubstabile
  22. 22

    Agente
    Dormiente

    AgentOps: Multi-agent infrastructure remediation platform. A2A protocol, agent coordination, HITL approval, auto-rollback.

    githubcrescita misurataApri fonte ↗

    0stelle GitHubstabile
  23. 23

    MCP
    Attivo

    AURA is a production-tested SRE agent platform you can deploy in minutes. AURA handles the guardrails, APIs, state management, streaming, and failure handling required to put AI to work safely on production infrastructure.

    githubcrescita misurataApri fonte ↗

    339stelle GitHub+81 (+31.4 %)
  24. 24
    Attivo

    Agent Skill evaluation harness for paired variants, trace artifacts, and runner adapters

    githubcrescita misurataApri fonte ↗

    Installa git clone https://github.com/adewale/skill-eval-harness ~/.claude/skills/skill-eval-harness

    73stelle GitHub+2 (+2.8 %)
  25. 25
    Dormiente

    Local open-source dev tool to debug, secure, and evaluate LLM agents. Provides static analysis, dynamic security checks, and runtime monitoring - integrates with Cursor and Claude Code.

    githubcrescita misurataApri fonte ↗

    Installa /plugin marketplace add cylestio/agent-inspector

    9stelle GitHubstabile