Le descrizioni degli strumenti sono in inglese.
Osservabilità LLM
Monitoraggio, tracce, valutazione e qualità delle applicazioni LLM.
Utilizzo
Attività
Ordina per
121
Classifica degli strumenti
Classifica per crescita misurata e normalizzata tra le fonti. I valori grezzi mantengono la propria finestra; una variazione richiede almeno due rilevazioni in 7 giorni.
- 1Attivo
ai-dev-stack
SkillProduction-grade AI coding rules for Cursor and Claude Code. 15 rules + 9 doc templates + skills + agents + MCP setup. Drop into any project.
githubcrescita misurataApri fonte ↗
Installa
git clone https://github.com/aiagentwithdhruv/ai-dev-stack ~/.claude/skills/ai-dev-stack10stelle GitHubstabile - 2Attivo
Claude Code plugin marketplace — agentic-engineering (spec-driven shape→decide→execute→measure→eval with adversarial review) + github-keeper (audit/elevate READMEs and make a repo open-source-ready).
githubcrescita misurataApri fonte ↗
Installa
/plugin marketplace add GiustoPiedimonte/agentic-engineering-marketplace13stelle GitHubstabile - 3Attivo
nudge
MCPA typed, replayable, budget-aware programming language for LLM agents. Compiles to Python & TypeScript.
githubcrescita misurataApri fonte ↗
2stelle GitHubstabile - 4Attivo
agent-mmm
SkillMarketing Mix Model expert agent plugin for Claude Code - pymc-marketing v0.18.2+
githubcrescita misurataApri fonte ↗
Installa
/plugin marketplace add Yakoub-ai/agent-mmm4stelle GitHubstabile - 5Attivo
rashomon
SkillMeasure prompt and skill improvements with blind A/B comparison.
githubcrescita misurataApri fonte ↗
Installa
/plugin marketplace add shinpr/rashomon18stelle GitHubstabile - 6Attivo
agentgateway
MCPAgentGateway — independent third-party profile of a public API surface, by API Evangelist. AgentGateway is an open-source, AI-native proxy and gateway for routing, observing, and governing traffic to and from AI agents, LLM providers, and…
githubcrescita misurataApri fonte ↗
0stelle GitHubstabile - 7Attivo
evoagent-os
AgenteLocal-first control plane for durable, governed agent teams: DAG orchestration, approvals, memory, signed skills, trace contracts, and realtime.
githubcrescita misurataApri fonte ↗
0stelle GitHubstabile - 8Attivo
Production reliability, benchmarking, and evaluation layer for MCP agent runtimes — protocol negotiation, A2A interop, circuit breakers, connection pooling, adaptive streaming, and autoscaling, with measured throughput/latency and a…
githubcrescita misurataApri fonte ↗
3stelle GitHubstabile - 9Attivo
dsh-plugins
AgenteGeneric DeepSeek Harness (dsh) plugins: A2A protocol server, session storage mirror, and Langfuse observability.
githubcrescita misurataApri fonte ↗
1stelle GitHubstabile - 10Dormiente
anti-lie
SkillDon't make LLMs honest. Make every factual claim auditable. — An LLM Claim Auditing Layer with T1-T7 truth gradients. 98.1% business effectiveness on LiarBench v0.2.
githubcrescita misurataApri fonte ↗
Installa
git clone https://github.com/lc198707/anti-lie ~/.claude/skills/anti-lie89stelle GitHubstabile - 11Attivo
skill-graveyard
SkillAudit which Claude Code skills you actually use — surface dead installs and hallucinated invocations from your session logs.
githubcrescita misurataApri fonte ↗
Installa
/plugin marketplace add sfrangulov/skill-graveyard10stelle GitHubstabile - 12Attivo
adherencia-reglas
SkillMide si tu agente de IA cumple las reglas que le escribiste. Lee el historial local de Claude Code y devuelve un porcentaje por regla. Sin instalar nada, sin red, biblioteca estándar.
githubcrescita misurataApri fonte ↗
Installa
git clone https://github.com/jleonceo/adherencia-reglas ~/.claude/skills/adherencia-reglas1stelle GitHubstabile - 13Attivo
DeepSeek-Infra
AltroLocal-first Agentic AI Infrastructure Platform with LLM Gateway, Agent DAG Runtime, MCP Tool Hub, A2A Agent Mesh, Local RAG, Tool Sandbox and Observability.
githubcrescita misurataApri fonte ↗
1stelle GitHubstabile - 14Dormiente
primal-core
AgenteThe reliability and interoperability layer for AI agents. A2A-native. MCP-bridged.
githubcrescita misurataApri fonte ↗
1stelle GitHubstabile - 15Dormiente
An AI-powered multi-agent system that demonstrates clinical triage, OTC medication recommendations, and e-pharmacy integration for respiratory conditions. Built with modular agents that collaborate to provide safe, intelligent healthcare…
githubcrescita misurataApri fonte ↗
0stelle GitHubstabile - 16Dormiente
agent-inspector
SkillLocal open-source dev tool to debug, secure, and evaluate LLM agents. Provides static analysis, dynamic security checks, and runtime monitoring - integrates with Cursor and Claude Code.
githubcrescita misurataApri fonte ↗
Installa
/plugin marketplace add cylestio/agent-inspector9stelle GitHubstabile - 17Dormiente
MCP as a Judge: a behavioral MCP that strengthens AI coding assistants via explicit LLM evaluations
mcpcrescita misurataApri fonte ↗
Installa
claude mcp add mcp-as-a-judge -- uvx mcp-as-a-judge17stelle GitHubstabile - 18Attivo
8 MCP SMB products — standalone AI servers for SMBs: Guardrails, FinOps, Observability, Router, Trust Score, Memory, ThinkSecure, A2A Lite. 37 tools, TNC credits billing.
githubcrescita misurataApri fonte ↗
1stelle GitHubstabile - 19Attivo
skill-receipts
SkillClaude Code skills where every entry ships receipts — accuracy-gated benchmarks against baseline and placebo, rejects published
githubcrescita misurataApri fonte ↗
Installa
git clone https://github.com/sjh9714/skill-receipts ~/.claude/skills/skill-receipts2stelle GitHubstabile - 20Dormiente
eval-layer
AltroA Claude Code skill that adds a rubric-based eval layer to any agent project. Framework-agnostic — generates rubric, test cases, judge prompt, and harness. Returns a weighted score plus a judge-leniency signal.
githubcrescita misurataApri fonte ↗
13stelle GitHubstabile - 21Dormiente
kaggle-capstone-ai-agent
AgenteA safety-first multi-agent mental health companion with real-time distress tracking, triple-layer guardrails, and evidence-based grounding techniques. Built for Kaggle × Google Agents Intensive 2025 Capstone (Agents for Good Track)
githubcrescita misurataApri fonte ↗
1stelle GitHubstabile - 22Attivo
anchor
SkillAnchor — the production-grade AGENTS.md template for AI coding agents. 51 sections of battle-tested rules. Keep your agents grounded.
githubcrescita misurataApri fonte ↗
Installa
git clone https://github.com/peva3/anchor ~/.claude/skills/anchor8stelle GitHubstabile - 23Dormiente
Two Claude Code skills for adversarial code review (single-model PAR and multi-model MMAR with cross-critique to catch hallucinations), plus a fixture-based eval suite.
githubcrescita misurataApri fonte ↗
Installa
/plugin marketplace add prime-radiant-inc/parallel-adversarial-review17stelle GitHubstabile - 24Dormiente
agentops
AgenteAgentOps: Multi-agent infrastructure remediation platform. A2A protocol, agent coordination, HITL approval, auto-rollback.
githubcrescita misurataApri fonte ↗
0stelle GitHubstabile - 25Attivo
agent-stack
SkillProduction patterns for agent orchestrators, harnesses, evals, MCP/A2A interoperability, memory, provider routing, and LLM usage metering.
githubcrescita misurataApri fonte ↗
Installa
git clone https://github.com/ssheleg/agent-stack ~/.claude/skills/agent-stack2stelle GitHubstabile