Le descrizioni degli strumenti sono in inglese.

Osservabilità LLM

Monitoraggio, tracce, valutazione e qualità delle applicazioni LLM.

Utilizzo

Attività

Ordina per

121

Classifica degli strumenti

Classifica per crescita misurata e normalizzata tra le fonti. I valori grezzi mantengono la propria finestra; una variazione richiede almeno due rilevazioni in 7 giorni.

  1. 1

    Skill
    Attivo

    Production-grade AI coding rules for Cursor and Claude Code. 15 rules + 9 doc templates + skills + agents + MCP setup. Drop into any project.

    githubcrescita misurataApri fonte ↗

    Installa git clone https://github.com/aiagentwithdhruv/ai-dev-stack ~/.claude/skills/ai-dev-stack

    10stelle GitHubstabile
  2. 2
    Attivo

    Claude Code plugin marketplace — agentic-engineering (spec-driven shape→decide→execute→measure→eval with adversarial review) + github-keeper (audit/elevate READMEs and make a repo open-source-ready).

    githubcrescita misurataApri fonte ↗

    Installa /plugin marketplace add GiustoPiedimonte/agentic-engineering-marketplace

    13stelle GitHubstabile
  3. 3

    MCP
    Attivo

    A typed, replayable, budget-aware programming language for LLM agents. Compiles to Python & TypeScript.

    githubcrescita misurataApri fonte ↗

    2stelle GitHubstabile
  4. 4

    Skill
    Attivo

    Marketing Mix Model expert agent plugin for Claude Code - pymc-marketing v0.18.2+

    githubcrescita misurataApri fonte ↗

    Installa /plugin marketplace add Yakoub-ai/agent-mmm

    4stelle GitHubstabile
  5. 5

    Skill
    Attivo

    Measure prompt and skill improvements with blind A/B comparison.

    githubcrescita misurataApri fonte ↗

    Installa /plugin marketplace add shinpr/rashomon

    18stelle GitHubstabile
  6. 6
    Attivo

    AgentGateway — independent third-party profile of a public API surface, by API Evangelist. AgentGateway is an open-source, AI-native proxy and gateway for routing, observing, and governing traffic to and from AI agents, LLM providers, and…

    githubcrescita misurataApri fonte ↗

    0stelle GitHubstabile
  7. 7

    Agente
    Attivo

    Local-first control plane for durable, governed agent teams: DAG orchestration, approvals, memory, signed skills, trace contracts, and realtime.

    githubcrescita misurataApri fonte ↗

    0stelle GitHubstabile
  8. 8
    Attivo

    Production reliability, benchmarking, and evaluation layer for MCP agent runtimes — protocol negotiation, A2A interop, circuit breakers, connection pooling, adaptive streaming, and autoscaling, with measured throughput/latency and a…

    githubcrescita misurataApri fonte ↗

    3stelle GitHubstabile
  9. 9

    Agente
    Attivo

    Generic DeepSeek Harness (dsh) plugins: A2A protocol server, session storage mirror, and Langfuse observability.

    githubcrescita misurataApri fonte ↗

    1stelle GitHubstabile
  10. 10

    Skill
    Dormiente

    Don't make LLMs honest. Make every factual claim auditable. — An LLM Claim Auditing Layer with T1-T7 truth gradients. 98.1% business effectiveness on LiarBench v0.2.

    githubcrescita misurataApri fonte ↗

    Installa git clone https://github.com/lc198707/anti-lie ~/.claude/skills/anti-lie

    89stelle GitHubstabile
  11. 11
    Attivo

    Audit which Claude Code skills you actually use — surface dead installs and hallucinated invocations from your session logs.

    githubcrescita misurataApri fonte ↗

    Installa /plugin marketplace add sfrangulov/skill-graveyard

    10stelle GitHubstabile
  12. 12
    Attivo

    Mide si tu agente de IA cumple las reglas que le escribiste. Lee el historial local de Claude Code y devuelve un porcentaje por regla. Sin instalar nada, sin red, biblioteca estándar.

    githubcrescita misurataApri fonte ↗

    Installa git clone https://github.com/jleonceo/adherencia-reglas ~/.claude/skills/adherencia-reglas

    1stelle GitHubstabile
  13. 13
    Attivo

    Local-first Agentic AI Infrastructure Platform with LLM Gateway, Agent DAG Runtime, MCP Tool Hub, A2A Agent Mesh, Local RAG, Tool Sandbox and Observability.

    githubcrescita misurataApri fonte ↗

    1stelle GitHubstabile
  14. 14

    Agente
    Dormiente

    The reliability and interoperability layer for AI agents. A2A-native. MCP-bridged.

    githubcrescita misurataApri fonte ↗

    1stelle GitHubstabile
  15. 15
    Dormiente

    An AI-powered multi-agent system that demonstrates clinical triage, OTC medication recommendations, and e-pharmacy integration for respiratory conditions. Built with modular agents that collaborate to provide safe, intelligent healthcare…

    githubcrescita misurataApri fonte ↗

    0stelle GitHubstabile
  16. 16
    Dormiente

    Local open-source dev tool to debug, secure, and evaluate LLM agents. Provides static analysis, dynamic security checks, and runtime monitoring - integrates with Cursor and Claude Code.

    githubcrescita misurataApri fonte ↗

    Installa /plugin marketplace add cylestio/agent-inspector

    9stelle GitHubstabile
  17. 17
    Dormiente

    MCP as a Judge: a behavioral MCP that strengthens AI coding assistants via explicit LLM evaluations

    mcpcrescita misurataApri fonte ↗

    Installa claude mcp add mcp-as-a-judge -- uvx mcp-as-a-judge

    17stelle GitHubstabile
  18. 18
    Attivo

    8 MCP SMB products — standalone AI servers for SMBs: Guardrails, FinOps, Observability, Router, Trust Score, Memory, ThinkSecure, A2A Lite. 37 tools, TNC credits billing.

    githubcrescita misurataApri fonte ↗

    1stelle GitHubstabile
  19. 19
    Attivo

    Claude Code skills where every entry ships receipts — accuracy-gated benchmarks against baseline and placebo, rejects published

    githubcrescita misurataApri fonte ↗

    Installa git clone https://github.com/sjh9714/skill-receipts ~/.claude/skills/skill-receipts

    2stelle GitHubstabile
  20. 20

    Altro
    Dormiente

    A Claude Code skill that adds a rubric-based eval layer to any agent project. Framework-agnostic — generates rubric, test cases, judge prompt, and harness. Returns a weighted score plus a judge-leniency signal.

    githubcrescita misurataApri fonte ↗

    13stelle GitHubstabile
  21. 21
    Dormiente

    A safety-first multi-agent mental health companion with real-time distress tracking, triple-layer guardrails, and evidence-based grounding techniques. Built for Kaggle × Google Agents Intensive 2025 Capstone (Agents for Good Track)

    githubcrescita misurataApri fonte ↗

    1stelle GitHubstabile
  22. 22

    Skill
    Attivo

    Anchor — the production-grade AGENTS.md template for AI coding agents. 51 sections of battle-tested rules. Keep your agents grounded.

    githubcrescita misurataApri fonte ↗

    Installa git clone https://github.com/peva3/anchor ~/.claude/skills/anchor

    8stelle GitHubstabile
  23. 23
    Dormiente

    Two Claude Code skills for adversarial code review (single-model PAR and multi-model MMAR with cross-critique to catch hallucinations), plus a fixture-based eval suite.

    githubcrescita misurataApri fonte ↗

    Installa /plugin marketplace add prime-radiant-inc/parallel-adversarial-review

    17stelle GitHubstabile
  24. 24

    Agente
    Dormiente

    AgentOps: Multi-agent infrastructure remediation platform. A2A protocol, agent coordination, HITL approval, auto-rollback.

    githubcrescita misurataApri fonte ↗

    0stelle GitHubstabile
  25. 25

    Skill
    Attivo

    Production patterns for agent orchestrators, harnesses, evals, MCP/A2A interoperability, memory, provider routing, and LLM usage metering.

    githubcrescita misurataApri fonte ↗

    Installa git clone https://github.com/ssheleg/agent-stack ~/.claude/skills/agent-stack

    2stelle GitHubstabile