Le descrizioni degli strumenti sono in inglese.

Osservabilità LLM

Monitoraggio, tracce, valutazione e qualità delle applicazioni LLM.

Utilizzo

Attività

Ordina per

118

Classifica degli strumenti

Ranked by creation date, newest first; undated entries come last. I valori grezzi mantengono la propria finestra; una variazione richiede almeno due rilevazioni in 7 giorni.

  1. 1
    Attivo

    Governed local-LLM observability: model policy, prompt scanner, capture proxy, 21 tools.

    mcpcrescita misurataApri fonte ↗

    Installa claude mcp add ai-guardian -- uvx ai-guardian-aiops

    0stelle GitHubstabile
  2. 2

    Altro
    Attivo

    AI SRE AgenticOps for Kubernetes and cloud infrastructure.

    githubcrescita misurataApri fonte ↗

    782stelle GitHub+1 (+0.13 %)
  3. 3

    Altro
    Attivo

    🔑 HomeStream · 家园·流 — 零成本自托管多Agent协作框架,通往AI世界的那把钥匙 | Zero-cost self-hosted multi-agent framework — The key to AI world

    githubcrescita misurataApri fonte ↗

    0stelle GitHubstabile
  4. 4

    Skill
    Attivo

    Turn one decision into a judged tournament of solutions, then pick the best — a Claude Code skill that generates candidates, auto-derives the rubric, judges independently, and returns a defensible winner.

    githubcrescita misurataApri fonte ↗

    Installa /plugin marketplace add CoriChui/bakeoff

    10stelle GitHubstabile
  5. 5

    Skill
    Attivo

    The Fable Workflow: how Claude Fable 5 worked, distilled into skills any model can run, with the eval that keeps it honest. Think / act / prove.

    githubcrescita misurataApri fonte ↗

    Installa git clone https://github.com/Sahir619/fable-method ~/.claude/skills/fable-method

    2 272stelle GitHub+16 (+0.71 %)
  6. 6
    Attivo

    Self-hosted AI SRE for Kubernetes — zero-instrumentation eBPF observability plus a copilot that fixes issues through guardrailed, self-verifying actions. BYO-LLM, air-gapped capable.

    githubcrescita misurataApri fonte ↗

    131stelle GitHub-4 (-3.0 %)
  7. 7
    Attivo

    Open benchmark for Claude Code SEO skills — real headless execution against fixture sites with planted-defect answer keys. Deterministic scoring, pre-registered rubric.

    githubcrescita misurataApri fonte ↗

    Installa git clone https://github.com/aleclindz/seo-skill-bench ~/.claude/skills/seo-skill-bench

    51stelle GitHub-5 (-8.9 %)
  8. 8
    Attivo

    Make Claude Opus 4.8 behave like Claude Fable 5 — doctrine output style, drift-catching hooks, and an eval loop against golden Fable transcripts. Claude Code plugin.

    githubcrescita misurataApri fonte ↗

    Installa /plugin marketplace add rennf93/opus-fable-playbook

    33stelle GitHubstabile
  9. 9
    Attivo

    Claude Code plugin marketplace — agentic-engineering (spec-driven shape→decide→execute→measure→eval with adversarial review) + github-keeper (audit/elevate READMEs and make a repo open-source-ready).

    githubcrescita misurataApri fonte ↗

    Installa /plugin marketplace add GiustoPiedimonte/agentic-engineering-marketplace

    13stelle GitHubstabile
  10. 10

    Altro
    Attivo

    AI-detector auditing toolkit: measures how often a detector flags genuine human writing, how stable its verdict is across seeds, and whether that verdict survives meaning-preserving edits. Detector-in-the-loop measurement harness. Claude…

    githubcrescita misurataApri fonte ↗

    18stelle GitHubstabile
  11. 11
    Attivo

    Multi-tier framework for evaluating AI agent skills with quality gates, semantic overlap detection, synthetic evaluation dataset generation, and live agent evaluation that measures how skills affect agent behavior.

    githubcrescita misurataApri fonte ↗

    Installa git clone https://github.com/NVIDIA/SkillEvaluator ~/.claude/skills/SkillEvaluator

    389stelle GitHub+57 (+17.2 %)
  12. 12

    Agente
    Attivo

    Agent 降智检测与自愈公评网络 — an immune system for the AI agent society

    githubcrescita misurataApri fonte ↗

    1stelle GitHubstabile
  13. 13

    Agente
    Attivo

    DataBuff is an AI-native APM built on Opentelemetry,with multi-agent troubleshooting out of the box.

    githubcrescita misurataApri fonte ↗

    642stelle GitHub+32 (+5.2 %)
  14. 14

    Agente
    Attivo

    A local A2A event and continuity service for agent applications, with structured history, provenance-safe transcripts, literal search, and natural-language queries.

    githubcrescita misurataApri fonte ↗

    1stelle GitHub+1
  15. 15
    Attivo

    Terminal-first validation harness for reviewing PRDs, RFCs, strategy docs, and experiment plans.

    githubcrescita misurataApri fonte ↗

    1stelle GitHubstabile
  16. 16
    Attivo

    Optimize any AI agent’s skills, tools/MCP, and prompts against your own evals.

    githubcrescita misurataApri fonte ↗

    47stelle GitHub+1 (+2.2 %)
  17. 17

    Agente
    Attivo

    A lightweight Python library that decouples agentic runtime from applications it builds

    githubcrescita misurataApri fonte ↗

    1stelle GitHubstabile
  18. 18
    Attivo

    Agent Skill evaluation harness for paired variants, trace artifacts, and runner adapters

    githubcrescita misurataApri fonte ↗

    Installa git clone https://github.com/adewale/skill-eval-harness ~/.claude/skills/skill-eval-harness

    73stelle GitHub+4 (+5.8 %)
  19. 19

    Agente
    Attivo

    The agent framework where the model never holds the trigger — every consequential action clears your policy first, waits for a human when it matters, and lands on a record you can verify. Build on it, or put it around the agent you already…

    githubcrescita misurataApri fonte ↗

    1stelle GitHubstabile
  20. 20

    Skill
    Attivo

    Axiom is a curated marketplace of shared plugins for Claude Code and Codex.

    githubcrescita misurataApri fonte ↗

    Installa git clone https://github.com/netopsengineer/axiom ~/.claude/skills/axiom

    5stelle GitHubstabile
  21. 21
    Attivo

    Local-first Agentic AI Infrastructure Platform with LLM Gateway, Agent DAG Runtime, MCP Tool Hub, A2A Agent Mesh, Local RAG, Tool Sandbox and Observability.

    githubcrescita misurataApri fonte ↗

    1stelle GitHubstabile
  22. 22

    Skill
    Attivo

    Anchor — the production-grade AGENTS.md template for AI coding agents. 51 sections of battle-tested rules. Keep your agents grounded.

    githubcrescita misurataApri fonte ↗

    Installa git clone https://github.com/peva3/anchor ~/.claude/skills/anchor

    8stelle GitHubstabile

Risorse didattiche e di riferimento

Ranked by creation date, newest first; undated entries come last. Queste risorse restano accessibili separatamente e non partecipano alla classifica principale.

  1. 1
    AttivoApri fonte ↗

    50+ curated LLM observability tools PLUS 26 Agent Skills (several with runnable, unit-tested scripts) to build, evaluate, debug, secure & monitor reliable LLM apps. Tracing, evals, guardrails, LLMOps.

    github

    Installa git clone https://github.com/ContextJet-ai/awesome-llm-observability ~/.claude/skills/awesome-llm-observability

    Risorsacrescita misurata
    32stelle GitHub+2 (+6.7 %)
  2. 2
    AttivoApri fonte ↗

    Curated + scored map of official MCP servers and agents for DevOps, Cloud, SRE, and Platform Engineering — every entry rated on production access, approval gates, and audit evidence.

    githubRisorsacrescita misurata
    77stelle GitHubstabile
  3. 3

    tunelab

    Skill
    AttivoApri fonte ↗

    Claude Code plugin for LLM fine-tuning, distillation, and evaluation — decide whether you need fine-tuning at all, distill your LLM logs into small local models (MLX/LoRA), evaluate with held-out discipline, and learn the why at every step.

    github

    Installa /plugin marketplace add rchaz/tunelab

    Risorsacrescita misurata
    6stelle GitHubstabile