Die Werkzeugbeschreibungen sind auf Englisch.

LLM-Beobachtbarkeit

Monitoring, Traces, Bewertung und Qualität von LLM-Anwendungen.

Anwendungsfall

Aktivität

Sortieren nach

118

Werkzeug-Rangliste

Ranked by creation date, newest first; undated entries come last. Rohwerte behalten ihr eigenes Zeitfenster; eine Änderung erfordert mindestens zwei Messungen in 7 Tagen.

  1. 1
    Aktiv

    Governed local-LLM observability: model policy, prompt scanner, capture proxy, 21 tools.

    mcpgemessenes WachstumQuelle öffnen ↗

    Installieren claude mcp add ai-guardian -- uvx ai-guardian-aiops

    0GitHub-Sternestabil
  2. 2

    Sonstige
    Aktiv

    AI SRE AgenticOps for Kubernetes and cloud infrastructure.

    githubgemessenes WachstumQuelle öffnen ↗

    782GitHub-Sterne+1 (+0.13 %)
  3. 3

    Sonstige
    Aktiv

    🔑 HomeStream · 家园·流 — 零成本自托管多Agent协作框架,通往AI世界的那把钥匙 | Zero-cost self-hosted multi-agent framework — The key to AI world

    githubgemessenes WachstumQuelle öffnen ↗

    0GitHub-Sternestabil
  4. 4

    Skill
    Aktiv

    Turn one decision into a judged tournament of solutions, then pick the best — a Claude Code skill that generates candidates, auto-derives the rubric, judges independently, and returns a defensible winner.

    githubgemessenes WachstumQuelle öffnen ↗

    Installieren /plugin marketplace add CoriChui/bakeoff

    10GitHub-Sternestabil
  5. 5

    Skill
    Aktiv

    The Fable Workflow: how Claude Fable 5 worked, distilled into skills any model can run, with the eval that keeps it honest. Think / act / prove.

    githubgemessenes WachstumQuelle öffnen ↗

    Installieren git clone https://github.com/Sahir619/fable-method ~/.claude/skills/fable-method

    2 272GitHub-Sterne+16 (+0.71 %)
  6. 6

    Sonstige
    Aktiv

    Self-hosted AI SRE for Kubernetes — zero-instrumentation eBPF observability plus a copilot that fixes issues through guardrailed, self-verifying actions. BYO-LLM, air-gapped capable.

    githubgemessenes WachstumQuelle öffnen ↗

    131GitHub-Sterne-4 (-3.0 %)
  7. 7
    Aktiv

    Open benchmark for Claude Code SEO skills — real headless execution against fixture sites with planted-defect answer keys. Deterministic scoring, pre-registered rubric.

    githubgemessenes WachstumQuelle öffnen ↗

    Installieren git clone https://github.com/aleclindz/seo-skill-bench ~/.claude/skills/seo-skill-bench

    51GitHub-Sterne-5 (-8.9 %)
  8. 8
    Aktiv

    Make Claude Opus 4.8 behave like Claude Fable 5 — doctrine output style, drift-catching hooks, and an eval loop against golden Fable transcripts. Claude Code plugin.

    githubgemessenes WachstumQuelle öffnen ↗

    Installieren /plugin marketplace add rennf93/opus-fable-playbook

    33GitHub-Sternestabil
  9. 9
    Aktiv

    Claude Code plugin marketplace — agentic-engineering (spec-driven shape→decide→execute→measure→eval with adversarial review) + github-keeper (audit/elevate READMEs and make a repo open-source-ready).

    githubgemessenes WachstumQuelle öffnen ↗

    Installieren /plugin marketplace add GiustoPiedimonte/agentic-engineering-marketplace

    13GitHub-Sternestabil
  10. 10

    Sonstige
    Aktiv

    AI-detector auditing toolkit: measures how often a detector flags genuine human writing, how stable its verdict is across seeds, and whether that verdict survives meaning-preserving edits. Detector-in-the-loop measurement harness. Claude…

    githubgemessenes WachstumQuelle öffnen ↗

    18GitHub-Sternestabil
  11. 11
    Aktiv

    Multi-tier framework for evaluating AI agent skills with quality gates, semantic overlap detection, synthetic evaluation dataset generation, and live agent evaluation that measures how skills affect agent behavior.

    githubgemessenes WachstumQuelle öffnen ↗

    Installieren git clone https://github.com/NVIDIA/SkillEvaluator ~/.claude/skills/SkillEvaluator

    389GitHub-Sterne+57 (+17.2 %)
  12. 12
    Aktiv

    Agent 降智检测与自愈公评网络 — an immune system for the AI agent society

    githubgemessenes WachstumQuelle öffnen ↗

    1GitHub-Sternestabil
  13. 13

    Agent
    Aktiv

    DataBuff is an AI-native APM built on Opentelemetry,with multi-agent troubleshooting out of the box.

    githubgemessenes WachstumQuelle öffnen ↗

    642GitHub-Sterne+32 (+5.2 %)
  14. 14

    Agent
    Aktiv

    A local A2A event and continuity service for agent applications, with structured history, provenance-safe transcripts, literal search, and natural-language queries.

    githubgemessenes WachstumQuelle öffnen ↗

    1GitHub-Sterne+1
  15. 15

    Sonstige
    Aktiv

    Terminal-first validation harness for reviewing PRDs, RFCs, strategy docs, and experiment plans.

    githubgemessenes WachstumQuelle öffnen ↗

    1GitHub-Sternestabil
  16. 16
    Aktiv

    Optimize any AI agent’s skills, tools/MCP, and prompts against your own evals.

    githubgemessenes WachstumQuelle öffnen ↗

    47GitHub-Sterne+1 (+2.2 %)
  17. 17

    Agent
    Aktiv

    A lightweight Python library that decouples agentic runtime from applications it builds

    githubgemessenes WachstumQuelle öffnen ↗

    1GitHub-Sternestabil
  18. 18
    Aktiv

    Agent Skill evaluation harness for paired variants, trace artifacts, and runner adapters

    githubgemessenes WachstumQuelle öffnen ↗

    Installieren git clone https://github.com/adewale/skill-eval-harness ~/.claude/skills/skill-eval-harness

    73GitHub-Sterne+4 (+5.8 %)
  19. 19

    Agent
    Aktiv

    The agent framework where the model never holds the trigger — every consequential action clears your policy first, waits for a human when it matters, and lands on a record you can verify. Build on it, or put it around the agent you already…

    githubgemessenes WachstumQuelle öffnen ↗

    1GitHub-Sternestabil
  20. 20

    Skill
    Aktiv

    Axiom is a curated marketplace of shared plugins for Claude Code and Codex.

    githubgemessenes WachstumQuelle öffnen ↗

    Installieren git clone https://github.com/netopsengineer/axiom ~/.claude/skills/axiom

    5GitHub-Sternestabil
  21. 21

    Sonstige
    Aktiv

    Local-first Agentic AI Infrastructure Platform with LLM Gateway, Agent DAG Runtime, MCP Tool Hub, A2A Agent Mesh, Local RAG, Tool Sandbox and Observability.

    githubgemessenes WachstumQuelle öffnen ↗

    1GitHub-Sternestabil
  22. 22

    Skill
    Aktiv

    Anchor — the production-grade AGENTS.md template for AI coding agents. 51 sections of battle-tested rules. Keep your agents grounded.

    githubgemessenes WachstumQuelle öffnen ↗

    Installieren git clone https://github.com/peva3/anchor ~/.claude/skills/anchor

    8GitHub-Sternestabil

Lern- und Referenzressourcen

Ranked by creation date, newest first; undated entries come last. Diese Ressourcen bleiben getrennt zugänglich und fließen nicht in die Hauptwertung ein.

  1. 1
    AktivQuelle öffnen ↗

    50+ curated LLM observability tools PLUS 26 Agent Skills (several with runnable, unit-tested scripts) to build, evaluate, debug, secure & monitor reliable LLM apps. Tracing, evals, guardrails, LLMOps.

    github

    Installieren git clone https://github.com/ContextJet-ai/awesome-llm-observability ~/.claude/skills/awesome-llm-observability

    Ressourcegemessenes Wachstum
    32GitHub-Sterne+2 (+6.7 %)
  2. 2
    AktivQuelle öffnen ↗

    Curated + scored map of official MCP servers and agents for DevOps, Cloud, SRE, and Platform Engineering — every entry rated on production access, approval gates, and audit evidence.

    githubRessourcegemessenes Wachstum
    77GitHub-Sternestabil
  3. 3

    tunelab

    Skill
    AktivQuelle öffnen ↗

    Claude Code plugin for LLM fine-tuning, distillation, and evaluation — decide whether you need fine-tuning at all, distill your LLM logs into small local models (MLX/LoRA), evaluate with held-out discipline, and learn the why at every step.

    github

    Installieren /plugin marketplace add rchaz/tunelab

    Ressourcegemessenes Wachstum
    6GitHub-Sternestabil