Le descrizioni degli strumenti sono in inglese.
Osservabilità LLM
Monitoraggio, tracce, valutazione e qualità delle applicazioni LLM.
Utilizzo
Attività
Ordina per
121
Classifica degli strumenti
Classifica per crescita misurata e normalizzata tra le fonti. I valori grezzi mantengono la propria finestra; una variazione richiede almeno due rilevazioni in 7 giorni.
- 1Dormiente
anti-lie
SkillDon't make LLMs honest. Make every factual claim auditable. — An LLM Claim Auditing Layer with T1-T7 truth gradients. 98.1% business effectiveness on LiarBench v0.2.
githubcrescita misurataApri fonte ↗
Installa
git clone https://github.com/lc198707/anti-lie ~/.claude/skills/anti-lie89stelle GitHubstabile - 2Attivo
agent-stack
SkillProduction patterns for agent orchestrators, harnesses, evals, MCP/A2A interoperability, memory, provider routing, and LLM usage metering.
githubcrescita misurataApri fonte ↗
Installa
git clone https://github.com/ssheleg/agent-stack ~/.claude/skills/agent-stack2stelle GitHub+1 (+100.0 %) - 3Attivo
untell
AltroAI-detector auditing toolkit: measures how often a detector flags genuine human writing, how stable its verdict is across seeds, and whether that verdict survives meaning-preserving edits. Detector-in-the-loop measurement harness. Claude…
githubcrescita misurataApri fonte ↗
18stelle GitHubstabile - 4Attivo
Log Claude Code sessions to Opik, the open-source LLM observability and evaluation platform, built by Comet. Tracing, evaluation, and skills for observable AI applications.
githubcrescita misurataApri fonte ↗
22stelle GitHub+1 (+4.8 %) - 5Attivo
Deterministic, local-only audit reports for Claude Code AI agent sessions. Rust + SQLite. Zero network calls. MCP server for agents.
githubcrescita misurataApri fonte ↗
22stelle GitHub+2 (+10.0 %) - 6Attivo
deslop-GPT
SkillDeletion-first Agent Skill for removing test bloat, verification theater, and speculative fallbacks while preserving behavior.
githubcrescita misurataApri fonte ↗
Installa
git clone https://github.com/MrZoyo/deslop-GPT ~/.claude/skills/deslop-GPT87stelle GitHub+59 (+210.7 %) - 7Dormiente
agenttap
AgenteReal-time debugging proxy for Agent2Agent (A2A) multi-agent systems
githubcrescita misurataApri fonte ↗
3stelle GitHubstabile - 8Dormiente
agentanvil
AgenteContract-based testing for LLM agents: hybrid evaluation, multi-agent + A2A, deterministic replay.
githubcrescita misurataApri fonte ↗
0stelle GitHubstabile - 9Attivo
Sentry instrumentation skill for system-behavior tracking
githubcrescita misurataApri fonte ↗
Installa
git clone https://github.com/tortastudios/sentry-instrumentation ~/.claude/skills/sentry-instrumentation24stelle GitHubstabile - 10Dormiente
adaptive-harness
SkillA self-improving harness router for Claude Code.
githubcrescita misurataApri fonte ↗
Installa
/plugin marketplace add SeongwoongCho/adaptive-harness8stelle GitHubstabile - 11Attivo
MCP server for Langfuse LLM observability — trace and observation analysis.
mcpcrescita misurataApri fonte ↗
Installa
claude mcp add langfuse -- npx langfuse-observability-mcp-server80stelle GitHub+2 (+2.6 %) - 12Attivo
axiom
SkillAxiom is a curated marketplace of shared plugins for Claude Code and Codex.
githubcrescita misurataApri fonte ↗
Installa
git clone https://github.com/netopsengineer/axiom ~/.claude/skills/axiom5stelle GitHubstabile - 13Dormiente
astragraph
AgentePolicy-enforced observability and fail-closed guardrails for MCP/A2A multi-agent systems.
githubcrescita misurataApri fonte ↗
26stelle GitHubstabile - 14Attivo
A Go framework for building production agent systems with graph workflows, tools, memory, A2A, AG-UI, MCP, evaluation, and observability.
githubcrescita misurataApri fonte ↗
1 763stelle GitHub+15 (+0.86 %) - 15Attivo
CodeFlow
Agente面向长程研发任务的 Java Agent Harness,支持 A2A 跨语言协作、可恢复执行、上下文工程与 Eval 驱动开发。
githubcrescita misurataApri fonte ↗
0stelle GitHubstabile - 16Dormiente
kaggle-capstone-ai-agent
AgenteA safety-first multi-agent mental health companion with real-time distress tracking, triple-layer guardrails, and evidence-based grounding techniques. Built for Kaggle × Google Agents Intensive 2025 Capstone (Agents for Good Track)
githubcrescita misurataApri fonte ↗
1stelle GitHubstabile - 17Attivo
OpenJudge
SkillOpenJudge: A Unified Framework for Holistic Evaluation and Quality Rewards
githubcrescita misurataApri fonte ↗
Installa
git clone https://github.com/agentscope-ai/OpenJudge ~/.claude/skills/OpenJudge816stelle GitHub+11 (+1.4 %) - 18Dormiente
gt8004-sdk
AltroOfficial Python SDK for GT8004 — AI agent observability with MCP, A2A, x402 payment tracking. FastAPI, Flask, FastMCP middleware included.
githubcrescita misurataApri fonte ↗
1stelle GitHubstabile - 19Dormiente
🔍 AI observability skill for Claude Code. Debug LangChain/LangGraph agents by fetching execution traces from LangSmith Studio directly in your terminal.
githubcrescita misurataApri fonte ↗
Installa
git clone https://github.com/OthmanAdi/langsmith-fetch-skill ~/.claude/skills/langsmith-fetch-skill27stelle GitHubstabile - 20Attivo
skill-kit
Skilllocal-first analytics for AI agent skills
githubcrescita misurataApri fonte ↗
Installa
git clone https://github.com/crafter-station/skill-kit ~/.claude/skills/skill-kit77stelle GitHub+1 (+1.3 %) - 21Attivo
8 MCP SMB products — standalone AI servers for SMBs: Guardrails, FinOps, Observability, Router, Trust Score, Memory, ThinkSecure, A2A Lite. 37 tools, TNC credits billing.
githubcrescita misurataApri fonte ↗
1stelle GitHubstabile - 22Dormiente
agentops
AgenteAgentOps: Multi-agent infrastructure remediation platform. A2A protocol, agent coordination, HITL approval, auto-rollback.
githubcrescita misurataApri fonte ↗
0stelle GitHubstabile - 23Attivo
aura
MCPAURA is a production-tested SRE agent platform you can deploy in minutes. AURA handles the guardrails, APIs, state management, streaming, and failure handling required to put AI to work safely on production infrastructure.
githubcrescita misurataApri fonte ↗
339stelle GitHub+81 (+31.4 %) - 24Attivo
skill-eval-harness
SkillAgent Skill evaluation harness for paired variants, trace artifacts, and runner adapters
githubcrescita misurataApri fonte ↗
Installa
git clone https://github.com/adewale/skill-eval-harness ~/.claude/skills/skill-eval-harness73stelle GitHub+2 (+2.8 %) - 25Dormiente
agent-inspector
SkillLocal open-source dev tool to debug, secure, and evaluate LLM agents. Provides static analysis, dynamic security checks, and runtime monitoring - integrates with Cursor and Claude Code.
githubcrescita misurataApri fonte ↗
Installa
/plugin marketplace add cylestio/agent-inspector9stelle GitHubstabile