Le descrizioni degli strumenti sono in inglese.

Osservabilità LLM

Monitoraggio, tracce, valutazione e qualità delle applicazioni LLM.

Utilizzo

Attività

Ordina per

122

Classifica degli strumenti

Ranked by known GitHub stars, highest first; archived repositories come after maintained repositories. I valori grezzi mantengono la propria finestra; una variazione richiede almeno due rilevazioni in 7 giorni.

  1. 1
    Attivo

    Mini Program Engineering Skill Suite is an Agent Skill suite for evidence‑first mini‑program development. It helps agents bring WeChat or other mini‑program projects from vague intent to reliable engineering work: project intake, product…

    githubmomentum stimatoApri fonte ↗

    Installa git clone https://github.com/NocodeMrLi/mini-program-engineering-skill-suite ~/.claude/skills/mini-program-engineering-skill-suite

    39stelle GitHub
  2. 2

    Skill
    Attivo

    Real-time execution trace and cost intelligence for Claude Code

    githubcrescita misurataApri fonte ↗

    Installa /plugin marketplace add DeibyGS/claudestat

    34stelle GitHub+1 (+3.0 %)
  3. 3
    Attivo

    Make Claude Opus 4.8 behave like Claude Fable 5 — doctrine output style, drift-catching hooks, and an eval loop against golden Fable transcripts. Claude Code plugin.

    githubcrescita misurataApri fonte ↗

    Installa /plugin marketplace add rennf93/opus-fable-playbook

    33stelle GitHubstabile
  4. 4
    Dormiente

    🔍 AI observability skill for Claude Code. Debug LangChain/LangGraph agents by fetching execution traces from LangSmith Studio directly in your terminal.

    githubcrescita misurataApri fonte ↗

    Installa git clone https://github.com/OthmanAdi/langsmith-fetch-skill ~/.claude/skills/langsmith-fetch-skill

    27stelle GitHubstabile
  5. 5

    Agente
    Dormiente

    Policy-enforced observability and fail-closed guardrails for MCP/A2A multi-agent systems.

    githubcrescita misurataApri fonte ↗

    26stelle GitHubstabile
  6. 6
    Attivo

    Sentry instrumentation skill for system-behavior tracking

    githubcrescita misurataApri fonte ↗

    Installa git clone https://github.com/tortastudios/sentry-instrumentation ~/.claude/skills/sentry-instrumentation

    24stelle GitHubstabile
  7. 7
    Attivo

    Log Claude Code sessions to Opik, the open-source LLM observability and evaluation platform, built by Comet. Tracing, evaluation, and skills for observable AI applications.

    githubcrescita misurataApri fonte ↗

    22stelle GitHub+1 (+4.8 %)
  8. 8
    Attivo

    Deterministic, local-only audit reports for Claude Code AI agent sessions. Rust + SQLite. Zero network calls. MCP server for agents.

    githubcrescita misurataApri fonte ↗

    22stelle GitHub+2 (+10.0 %)
  9. 9

    Skill
    Attivo

    Measure prompt and skill improvements with blind A/B comparison.

    githubcrescita misurataApri fonte ↗

    Installa /plugin marketplace add shinpr/rashomon

    18stelle GitHubstabile
  10. 10

    Altro
    Attivo

    AI-detector auditing toolkit: measures how often a detector flags genuine human writing, how stable its verdict is across seeds, and whether that verdict survives meaning-preserving edits. Detector-in-the-loop measurement harness. Claude…

    githubcrescita misurataApri fonte ↗

    18stelle GitHubstabile
  11. 11
    Attivo

    Two Claude Code skills for adversarial code review (single-model PAR and multi-model MMAR with cross-critique to catch hallucinations), plus a fixture-based eval suite.

    githubcrescita misurataApri fonte ↗

    Installa /plugin marketplace add prime-radiant-inc/parallel-adversarial-review

    17stelle GitHubstabile
  12. 12
    Dormiente

    MCP as a Judge: a behavioral MCP that strengthens AI coding assistants via explicit LLM evaluations

    mcpcrescita misurataApri fonte ↗

    Installa claude mcp add mcp-as-a-judge -- uvx mcp-as-a-judge

    17stelle GitHubstabile
  13. 13
    Attivo

    Make Claude write clearly, for everyone.

    githubcrescita misurataApri fonte ↗

    Installa /plugin marketplace add stefanobaghino/simple-output-styles

    16stelle GitHubstabile
  14. 14

    Altro
    Attivo

    A command-line tool to scaffold and manage enterprise-ready AI Agents powered by the A2A (Agent-to-Agent) protocol

    githubcrescita misurataApri fonte ↗

    14stelle GitHub+1 (+7.7 %)
  15. 15
    Attivo

    Claude Code plugin marketplace — agentic-engineering (spec-driven shape→decide→execute→measure→eval with adversarial review) + github-keeper (audit/elevate READMEs and make a repo open-source-ready).

    githubcrescita misurataApri fonte ↗

    Installa /plugin marketplace add GiustoPiedimonte/agentic-engineering-marketplace

    13stelle GitHubstabile
  16. 16

    Altro
    Dormiente

    A Claude Code skill that adds a rubric-based eval layer to any agent project. Framework-agnostic — generates rubric, test cases, judge prompt, and harness. Returns a weighted score plus a judge-leniency signal.

    githubcrescita misurataApri fonte ↗

    13stelle GitHubstabile
  17. 17
    Attivo

    Query Spanlens LLM observability from Cursor, Claude Desktop, or Continue via MCP.

    mcpcrescita misurataApri fonte ↗

    Installa claude mcp add mcp-server -- npx @spanlens/mcp-server

    12stelle GitHubstabile
  18. 18

    Skill
    Attivo

    Turn one decision into a judged tournament of solutions, then pick the best — a Claude Code skill that generates candidates, auto-derives the rubric, judges independently, and returns a defensible winner.

    githubcrescita misurataApri fonte ↗

    Installa /plugin marketplace add CoriChui/bakeoff

    10stelle GitHubstabile
  19. 19
    Attivo

    Audit which Claude Code skills you actually use — surface dead installs and hallucinated invocations from your session logs.

    githubcrescita misurataApri fonte ↗

    Installa /plugin marketplace add sfrangulov/skill-graveyard

    10stelle GitHub+1 (+11.1 %)
  20. 20

    Altro
    Attivo

    A Go-native framework for LLM agents, with OpenTelemetry observability built in.

    githubcrescita misurataApri fonte ↗

    10stelle GitHubstabile
  21. 21

    Skill
    Attivo

    Production-grade AI coding rules for Cursor and Claude Code. 15 rules + 9 doc templates + skills + agents + MCP setup. Drop into any project.

    githubcrescita misurataApri fonte ↗

    Installa git clone https://github.com/aiagentwithdhruv/ai-dev-stack ~/.claude/skills/ai-dev-stack

    10stelle GitHub+1 (+11.1 %)
  22. 22
    Dormiente

    Local open-source dev tool to debug, secure, and evaluate LLM agents. Provides static analysis, dynamic security checks, and runtime monitoring - integrates with Cursor and Claude Code.

    githubcrescita misurataApri fonte ↗

    Installa /plugin marketplace add cylestio/agent-inspector

    9stelle GitHubstabile

Risorse didattiche e di riferimento

Ranked by known GitHub stars, highest first; archived repositories come after maintained repositories. Queste risorse restano accessibili separatamente e non partecipano alla classifica principale.

  1. 1
    AttivoApri fonte ↗

    50+ curated LLM observability tools PLUS 26 Agent Skills (several with runnable, unit-tested scripts) to build, evaluate, debug, secure & monitor reliable LLM apps. Tracing, evals, guardrails, LLMOps.

    github

    Installa git clone https://github.com/ContextJet-ai/awesome-llm-observability ~/.claude/skills/awesome-llm-observability

    Risorsacrescita misurata
    32stelle GitHub+2 (+6.7 %)
  2. 2
    AttivoApri fonte ↗

    My complete journey to becoming an Agentic AI Engineer through structured learning, projects, experiments, and production-ready implementations of modern AI systems.

    githubRisorsacrescita misurata
    18stelle GitHubstabile
  3. 3
    AttivoApri fonte ↗

    Documentation-discovery telemetry for Claude Code — heat/cold maps, health grade, evidence-backed router fixes. 100% local, zero tokens. /tt

    github

    Installa /plugin marketplace add Hedde/trigger_tree

    Risorsacrescita misurata
    14stelle GitHub+1 (+7.7 %)