Le descrizioni degli strumenti sono in inglese.
Osservabilità LLM
Monitoraggio, tracce, valutazione e qualità delle applicazioni LLM.
Utilizzo
Attività
Ordina per
118
Classifica degli strumenti
Ranked by creation date, newest first; undated entries come last. I valori grezzi mantengono la propria finestra; una variazione richiede almeno due rilevazioni in 7 giorni.
- 1Attivo
AI Guardian
MCPGoverned local-LLM observability: model policy, prompt scanner, capture proxy, 21 tools.
mcpcrescita misurataApri fonte ↗
Installa
claude mcp add ai-guardian -- uvx ai-guardian-aiops0stelle GitHubstabile - 2Attivo
Flawless
AltroAI SRE AgenticOps for Kubernetes and cloud infrastructure.
githubcrescita misurataApri fonte ↗
782stelle GitHub+1 (+0.13 %) - 3Attivo
homestream
Altro🔑 HomeStream · 家园·流 — 零成本自托管多Agent协作框架,通往AI世界的那把钥匙 | Zero-cost self-hosted multi-agent framework — The key to AI world
githubcrescita misurataApri fonte ↗
0stelle GitHubstabile - 4Attivo
bakeoff
SkillTurn one decision into a judged tournament of solutions, then pick the best — a Claude Code skill that generates candidates, auto-derives the rubric, judges independently, and returns a defensible winner.
githubcrescita misurataApri fonte ↗
Installa
/plugin marketplace add CoriChui/bakeoff10stelle GitHubstabile - 5Attivo
fable-method
SkillThe Fable Workflow: how Claude Fable 5 worked, distilled into skills any model can run, with the eval that keeps it honest. Think / act / prove.
githubcrescita misurataApri fonte ↗
Installa
git clone https://github.com/Sahir619/fable-method ~/.claude/skills/fable-method2 272stelle GitHub+16 (+0.71 %) - 6Attivo
rocketplaneIO
AltroSelf-hosted AI SRE for Kubernetes — zero-instrumentation eBPF observability plus a copilot that fixes issues through guardrailed, self-verifying actions. BYO-LLM, air-gapped capable.
githubcrescita misurataApri fonte ↗
131stelle GitHub-4 (-3.0 %) - 7Attivo
seo-skill-bench
SkillOpen benchmark for Claude Code SEO skills — real headless execution against fixture sites with planted-defect answer keys. Deterministic scoring, pre-registered rubric.
githubcrescita misurataApri fonte ↗
Installa
git clone https://github.com/aleclindz/seo-skill-bench ~/.claude/skills/seo-skill-bench51stelle GitHub-5 (-8.9 %) - 8Attivo
opus-fable-playbook
SkillMake Claude Opus 4.8 behave like Claude Fable 5 — doctrine output style, drift-catching hooks, and an eval loop against golden Fable transcripts. Claude Code plugin.
githubcrescita misurataApri fonte ↗
Installa
/plugin marketplace add rennf93/opus-fable-playbook33stelle GitHubstabile - 9Attivo
Claude Code plugin marketplace — agentic-engineering (spec-driven shape→decide→execute→measure→eval with adversarial review) + github-keeper (audit/elevate READMEs and make a repo open-source-ready).
githubcrescita misurataApri fonte ↗
Installa
/plugin marketplace add GiustoPiedimonte/agentic-engineering-marketplace13stelle GitHubstabile - 10Attivo
untell
AltroAI-detector auditing toolkit: measures how often a detector flags genuine human writing, how stable its verdict is across seeds, and whether that verdict survives meaning-preserving edits. Detector-in-the-loop measurement harness. Claude…
githubcrescita misurataApri fonte ↗
18stelle GitHubstabile - 11Attivo
SkillEvaluator
SkillMulti-tier framework for evaluating AI agent skills with quality gates, semantic overlap detection, synthetic evaluation dataset generation, and live agent evaluation that measures how skills affect agent behavior.
githubcrescita misurataApri fonte ↗
Installa
git clone https://github.com/NVIDIA/SkillEvaluator ~/.claude/skills/SkillEvaluator389stelle GitHub+57 (+17.2 %) - 12Attivo
DriftSentinel
AgenteAgent 降智检测与自愈公评网络 — an immune system for the AI agent society
githubcrescita misurataApri fonte ↗
1stelle GitHubstabile - 13Attivo
databuff
AgenteDataBuff is an AI-native APM built on Opentelemetry,with multi-agent troubleshooting out of the box.
githubcrescita misurataApri fonte ↗
642stelle GitHub+32 (+5.2 %) - 14Attivo
historian
AgenteA local A2A event and continuity service for agent applications, with structured history, provenance-safe transcripts, literal search, and natural-language queries.
githubcrescita misurataApri fonte ↗
1stelle GitHub+1 - 15Attivo
shokunin-review
AltroTerminal-first validation harness for reviewing PRDs, RFCs, strategy docs, and experiment plans.
githubcrescita misurataApri fonte ↗
1stelle GitHubstabile - 16Attivo
cap-evolve
MCPOptimize any AI agent’s skills, tools/MCP, and prompts against your own evals.
githubcrescita misurataApri fonte ↗
47stelle GitHub+1 (+2.2 %) - 17Attivo
pyxen
AgenteA lightweight Python library that decouples agentic runtime from applications it builds
githubcrescita misurataApri fonte ↗
1stelle GitHubstabile - 18Attivo
skill-eval-harness
SkillAgent Skill evaluation harness for paired variants, trace artifacts, and runner adapters
githubcrescita misurataApri fonte ↗
Installa
git clone https://github.com/adewale/skill-eval-harness ~/.claude/skills/skill-eval-harness73stelle GitHub+4 (+5.8 %) - 19Attivo
tulip-agents
AgenteThe agent framework where the model never holds the trigger — every consequential action clears your policy first, waits for a human when it matters, and lands on a record you can verify. Build on it, or put it around the agent you already…
githubcrescita misurataApri fonte ↗
1stelle GitHubstabile - 20Attivo
axiom
SkillAxiom is a curated marketplace of shared plugins for Claude Code and Codex.
githubcrescita misurataApri fonte ↗
Installa
git clone https://github.com/netopsengineer/axiom ~/.claude/skills/axiom5stelle GitHubstabile - 21Attivo
DeepSeek-Infra
AltroLocal-first Agentic AI Infrastructure Platform with LLM Gateway, Agent DAG Runtime, MCP Tool Hub, A2A Agent Mesh, Local RAG, Tool Sandbox and Observability.
githubcrescita misurataApri fonte ↗
1stelle GitHubstabile - 22Attivo
anchor
SkillAnchor — the production-grade AGENTS.md template for AI coding agents. 51 sections of battle-tested rules. Keep your agents grounded.
githubcrescita misurataApri fonte ↗
Installa
git clone https://github.com/peva3/anchor ~/.claude/skills/anchor8stelle GitHubstabile
Risorse didattiche e di riferimento
Ranked by creation date, newest first; undated entries come last. Queste risorse restano accessibili separatamente e non partecipano alla classifica principale.
- 1AttivoApri fonte ↗
50+ curated LLM observability tools PLUS 26 Agent Skills (several with runnable, unit-tested scripts) to build, evaluate, debug, secure & monitor reliable LLM apps. Tracing, evals, guardrails, LLMOps.
githubInstalla
Risorsacrescita misuratagit clone https://github.com/ContextJet-ai/awesome-llm-observability ~/.claude/skills/awesome-llm-observability32stelle GitHub+2 (+6.7 %) - 2AttivoApri fonte ↗
Curated + scored map of official MCP servers and agents for DevOps, Cloud, SRE, and Platform Engineering — every entry rated on production access, approval gates, and audit evidence.
githubRisorsacrescita misurata77stelle GitHubstabile - 3AttivoApri fonte ↗
tunelab
SkillClaude Code plugin for LLM fine-tuning, distillation, and evaluation — decide whether you need fine-tuning at all, distill your LLM logs into small local models (MLX/LoRA), evaluate with held-out discipline, and learn the why at every step.
githubInstalla
Risorsacrescita misurata/plugin marketplace add rchaz/tunelab6stelle GitHubstabile