Le descrizioni degli strumenti sono in inglese.
Osservabilità LLM
Monitoraggio, tracce, valutazione e qualità delle applicazioni LLM.
Utilizzo
Attività
Ordina per
122
Classifica degli strumenti
Classifica mista: la crescita misurata ha la precedenza. I valori grezzi mantengono la propria finestra; una variazione richiede almeno due rilevazioni in 7 giorni.
- 1Attivo
skill-receipts
SkillClaude Code skills where every entry ships receipts — accuracy-gated benchmarks against baseline and placebo, rejects published
githubcrescita misurataApri fonte ↗
Installa
git clone https://github.com/sjh9714/skill-receipts ~/.claude/skills/skill-receipts2stelle GitHubstabile - 2Attivo
Sentry instrumentation skill for system-behavior tracking
githubcrescita misurataApri fonte ↗
Installa
git clone https://github.com/tortastudios/sentry-instrumentation ~/.claude/skills/sentry-instrumentation24stelle GitHubstabile - 3Dormiente
CustoFlow
AltroMulti-agent customer support system with Google ADK & Gemini 2.5 Flash Lite. Kaggle capstone demonstrating 11+ concepts. Automates 80%+ queries, <10s response time.
githubcrescita misurataApri fonte ↗
2stelle GitHubstabile - 4Attivo
Deterministic, local-only audit reports for Claude Code AI agent sessions. Rust + SQLite. Zero network calls. MCP server for agents.
githubcrescita misurataApri fonte ↗
22stelle GitHub+2 (+10.0 %) - 5Attivo
nudge
MCPA typed, replayable, budget-aware programming language for LLM agents. Compiles to Python & TypeScript.
githubcrescita misurataApri fonte ↗
2stelle GitHubstabile - 6Attivo
Log Claude Code sessions to Opik, the open-source LLM observability and evaluation platform, built by Comet. Tracing, evaluation, and skills for observable AI applications.
githubcrescita misurataApri fonte ↗
22stelle GitHub+1 (+4.8 %) - 7Attivo
untell
AltroAI-detector auditing toolkit: measures how often a detector flags genuine human writing, how stable its verdict is across seeds, and whether that verdict survives meaning-preserving edits. Detector-in-the-loop measurement harness. Claude…
githubcrescita misurataApri fonte ↗
18stelle GitHubstabile - 8Dormiente
otel-agent-provenance
AgenteOpenTelemetry semantic conventions and instrumentation for agent provenance, derivation lineage, and acceptance criteria evaluation. Fills the Microsoft AI stack observability gap.
githubcrescita misurataApri fonte ↗
0stelle GitHubstabile - 9Dormiente
Agents-eval
AgenteA Multi-Agent System (MAS) evaluation framework using PydanticAI that generates and evaluates scientific paper reviews through a three-tiered assessment approach: traditional metrics, LLM-as-a-Judge, and graph-based complexity analysis.
githubcrescita misurataApri fonte ↗
2stelle GitHubstabile - 10Attivo
Multi-agent quality gate skill for Claude Code that researches, reviews, tests and challenges AI-generated work before the final answer.
githubcrescita misurataApri fonte ↗
Installa
git clone https://github.com/ma-nucho-pro/supervisor-skill-claude ~/.claude/skills/supervisor-skill-claude2stelle GitHubstabile - 11Attivo
rashomon
SkillMeasure prompt and skill improvements with blind A/B comparison.
githubcrescita misurataApri fonte ↗
Installa
/plugin marketplace add shinpr/rashomon18stelle GitHubstabile - 12Attivo
agent-stack
SkillProduction patterns for agent orchestrators, harnesses, evals, MCP/A2A interoperability, memory, provider routing, and LLM usage metering.
githubcrescita misurataApri fonte ↗
Installa
git clone https://github.com/ssheleg/agent-stack ~/.claude/skills/agent-stack2stelle GitHub+1 (+100.0 %) - 13Dormiente
MCP as a Judge: a behavioral MCP that strengthens AI coding assistants via explicit LLM evaluations
mcpcrescita misurataApri fonte ↗
Installa
claude mcp add mcp-as-a-judge -- uvx mcp-as-a-judge17stelle GitHubstabile - 14Attivo
sre-on-call
AgenteMulti-agent SRE on-call investigator that auto-triages Slack/Discord infrastructure alerts via AWS Bedrock AgentCore, fanning out to specialized agents (CloudWatch, EKS, Slack/Discord scanners) for parallel investigation.
githubcrescita misurataApri fonte ↗
3stelle GitHubstabile - 15Attivo
deslop-GPT
SkillDeletion-first Agent Skill for removing test bloat, verification theater, and speculative fallbacks while preserving behavior.
githubcrescita misurataApri fonte ↗
Installa
git clone https://github.com/MrZoyo/deslop-GPT ~/.claude/skills/deslop-GPT76stelle GitHub+48 (+171.4 %) - 16Attivo
Mini Program Engineering Skill Suite is an Agent Skill suite for evidence‑first mini‑program development. It helps agents bring WeChat or other mini‑program projects from vague intent to reliable engineering work: project intake, product…
githubcrescita misurataApri fonte ↗
Installa
git clone https://github.com/NocodeMrLi/mini-program-engineering-skill-suite ~/.claude/skills/mini-program-engineering-skill-suite43stelle GitHub+4 (+10.3 %) - 17Attivo
Build tested agent skills and govern their lifecycle through a user-defined marketplace: evidence, discovery, updates, rollback, quarantine, and 17-platform distribution.
githubmomentum stimatoApri fonte ↗
Installa
git clone https://github.com/FrancyJGLisboa/agent-skills-platform ~/.claude/skills/agent-skills-platform2 376stelle GitHub—
Risorse didattiche e di riferimento
Classifica per crescita misurata e normalizzata tra le fonti. Queste risorse restano accessibili separatamente e non partecipano alla classifica principale.
- 1AttivoApri fonte ↗
trigger_tree
SkillDocumentation-discovery telemetry for Claude Code — heat/cold maps, health grade, evidence-backed router fixes. 100% local, zero tokens. /tt
githubInstalla
Risorsacrescita misurata/plugin marketplace add Hedde/trigger_tree14stelle GitHub+1 (+7.7 %) - 2AttivoApri fonte ↗
Handbook técnico aberto sobre engenharia de IA em produção: ML tradicional, LLMs, RAG, agentes, segurança, observabilidade, FinOps e deployment.
githubRisorsacrescita misurata1stelle GitHubstabile - 3AttivoApri fonte ↗
50+ curated LLM observability tools PLUS 26 Agent Skills (several with runnable, unit-tested scripts) to build, evaluate, debug, secure & monitor reliable LLM apps. Tracing, evals, guardrails, LLMOps.
githubInstalla
Risorsacrescita misuratagit clone https://github.com/ContextJet-ai/awesome-llm-observability ~/.claude/skills/awesome-llm-observability33stelle GitHub+3 (+10.0 %) - 4AttivoApri fonte ↗
Agentic_AI_Engineer
AgenteMy complete journey to becoming an Agentic AI Engineer through structured learning, projects, experiments, and production-ready implementations of modern AI systems.
githubRisorsacrescita misurata18stelle GitHubstabile - 5AttivoApri fonte ↗
tunelab
SkillClaude Code plugin for LLM fine-tuning, distillation, and evaluation — decide whether you need fine-tuning at all, distill your LLM logs into small local models (MLX/LoRA), evaluate with held-out discipline, and learn the why at every step.
githubInstalla
Risorsacrescita misurata/plugin marketplace add rchaz/tunelab6stelle GitHubstabile