Le descrizioni degli strumenti sono in inglese.

Osservabilità LLM

Monitoraggio, tracce, valutazione e qualità delle applicazioni LLM.

Utilizzo

Attività

Ordina per

122

Classifica degli strumenti

Classifica mista: la crescita misurata ha la precedenza. I valori grezzi mantengono la propria finestra; una variazione richiede almeno due rilevazioni in 7 giorni.

  1. 1
    Attivo

    Claude Code skills where every entry ships receipts — accuracy-gated benchmarks against baseline and placebo, rejects published

    githubcrescita misurataApri fonte ↗

    Installa git clone https://github.com/sjh9714/skill-receipts ~/.claude/skills/skill-receipts

    2stelle GitHubstabile
  2. 2
    Attivo

    Sentry instrumentation skill for system-behavior tracking

    githubcrescita misurataApri fonte ↗

    Installa git clone https://github.com/tortastudios/sentry-instrumentation ~/.claude/skills/sentry-instrumentation

    24stelle GitHubstabile
  3. 3

    Altro
    Dormiente

    Multi-agent customer support system with Google ADK & Gemini 2.5 Flash Lite. Kaggle capstone demonstrating 11+ concepts. Automates 80%+ queries, <10s response time.

    githubcrescita misurataApri fonte ↗

    2stelle GitHubstabile
  4. 4
    Attivo

    Deterministic, local-only audit reports for Claude Code AI agent sessions. Rust + SQLite. Zero network calls. MCP server for agents.

    githubcrescita misurataApri fonte ↗

    22stelle GitHub+2 (+10.0 %)
  5. 5

    MCP
    Attivo

    A typed, replayable, budget-aware programming language for LLM agents. Compiles to Python & TypeScript.

    githubcrescita misurataApri fonte ↗

    2stelle GitHubstabile
  6. 6
    Attivo

    Log Claude Code sessions to Opik, the open-source LLM observability and evaluation platform, built by Comet. Tracing, evaluation, and skills for observable AI applications.

    githubcrescita misurataApri fonte ↗

    22stelle GitHub+1 (+4.8 %)
  7. 7

    Altro
    Attivo

    AI-detector auditing toolkit: measures how often a detector flags genuine human writing, how stable its verdict is across seeds, and whether that verdict survives meaning-preserving edits. Detector-in-the-loop measurement harness. Claude…

    githubcrescita misurataApri fonte ↗

    18stelle GitHubstabile
  8. 8
    Dormiente

    OpenTelemetry semantic conventions and instrumentation for agent provenance, derivation lineage, and acceptance criteria evaluation. Fills the Microsoft AI stack observability gap.

    githubcrescita misurataApri fonte ↗

    0stelle GitHubstabile
  9. 9

    Agente
    Dormiente

    A Multi-Agent System (MAS) evaluation framework using PydanticAI that generates and evaluates scientific paper reviews through a three-tiered assessment approach: traditional metrics, LLM-as-a-Judge, and graph-based complexity analysis.

    githubcrescita misurataApri fonte ↗

    2stelle GitHubstabile
  10. 10
    Attivo

    Multi-agent quality gate skill for Claude Code that researches, reviews, tests and challenges AI-generated work before the final answer.

    githubcrescita misurataApri fonte ↗

    Installa git clone https://github.com/ma-nucho-pro/supervisor-skill-claude ~/.claude/skills/supervisor-skill-claude

    2stelle GitHubstabile
  11. 11

    Skill
    Attivo

    Measure prompt and skill improvements with blind A/B comparison.

    githubcrescita misurataApri fonte ↗

    Installa /plugin marketplace add shinpr/rashomon

    18stelle GitHubstabile
  12. 12

    Skill
    Attivo

    Production patterns for agent orchestrators, harnesses, evals, MCP/A2A interoperability, memory, provider routing, and LLM usage metering.

    githubcrescita misurataApri fonte ↗

    Installa git clone https://github.com/ssheleg/agent-stack ~/.claude/skills/agent-stack

    2stelle GitHub+1 (+100.0 %)
  13. 13
    Dormiente

    MCP as a Judge: a behavioral MCP that strengthens AI coding assistants via explicit LLM evaluations

    mcpcrescita misurataApri fonte ↗

    Installa claude mcp add mcp-as-a-judge -- uvx mcp-as-a-judge

    17stelle GitHubstabile
  14. 14

    Agente
    Attivo

    Multi-agent SRE on-call investigator that auto-triages Slack/Discord infrastructure alerts via AWS Bedrock AgentCore, fanning out to specialized agents (CloudWatch, EKS, Slack/Discord scanners) for parallel investigation.

    githubcrescita misurataApri fonte ↗

    3stelle GitHubstabile
  15. 15

    Skill
    Attivo

    Deletion-first Agent Skill for removing test bloat, verification theater, and speculative fallbacks while preserving behavior.

    githubcrescita misurataApri fonte ↗

    Installa git clone https://github.com/MrZoyo/deslop-GPT ~/.claude/skills/deslop-GPT

    76stelle GitHub+48 (+171.4 %)
  16. 16
    Attivo

    Mini Program Engineering Skill Suite is an Agent Skill suite for evidence‑first mini‑program development. It helps agents bring WeChat or other mini‑program projects from vague intent to reliable engineering work: project intake, product…

    githubcrescita misurataApri fonte ↗

    Installa git clone https://github.com/NocodeMrLi/mini-program-engineering-skill-suite ~/.claude/skills/mini-program-engineering-skill-suite

    43stelle GitHub+4 (+10.3 %)
  17. 17
    Attivo

    Build tested agent skills and govern their lifecycle through a user-defined marketplace: evidence, discovery, updates, rollback, quarantine, and 17-platform distribution.

    githubmomentum stimatoApri fonte ↗

    Installa git clone https://github.com/FrancyJGLisboa/agent-skills-platform ~/.claude/skills/agent-skills-platform

    2 376stelle GitHub

Risorse didattiche e di riferimento

Classifica per crescita misurata e normalizzata tra le fonti. Queste risorse restano accessibili separatamente e non partecipano alla classifica principale.

  1. 1
    AttivoApri fonte ↗

    Documentation-discovery telemetry for Claude Code — heat/cold maps, health grade, evidence-backed router fixes. 100% local, zero tokens. /tt

    github

    Installa /plugin marketplace add Hedde/trigger_tree

    Risorsacrescita misurata
    14stelle GitHub+1 (+7.7 %)
  2. 2
    AttivoApri fonte ↗

    Handbook técnico aberto sobre engenharia de IA em produção: ML tradicional, LLMs, RAG, agentes, segurança, observabilidade, FinOps e deployment.

    githubRisorsacrescita misurata
    1stelle GitHubstabile
  3. 3
    AttivoApri fonte ↗

    50+ curated LLM observability tools PLUS 26 Agent Skills (several with runnable, unit-tested scripts) to build, evaluate, debug, secure & monitor reliable LLM apps. Tracing, evals, guardrails, LLMOps.

    github

    Installa git clone https://github.com/ContextJet-ai/awesome-llm-observability ~/.claude/skills/awesome-llm-observability

    Risorsacrescita misurata
    33stelle GitHub+3 (+10.0 %)
  4. 4
    AttivoApri fonte ↗

    My complete journey to becoming an Agentic AI Engineer through structured learning, projects, experiments, and production-ready implementations of modern AI systems.

    githubRisorsacrescita misurata
    18stelle GitHubstabile
  5. 5

    tunelab

    Skill
    AttivoApri fonte ↗

    Claude Code plugin for LLM fine-tuning, distillation, and evaluation — decide whether you need fine-tuning at all, distill your LLM logs into small local models (MLX/LoRA), evaluate with held-out discipline, and learn the why at every step.

    github

    Installa /plugin marketplace add rchaz/tunelab

    Risorsacrescita misurata
    6stelle GitHubstabile