Die Werkzeugbeschreibungen sind auf Englisch.

LLM-Beobachtbarkeit

Monitoring, Traces, Bewertung und Qualität von LLM-Anwendungen.

Anwendungsfall

Aktivität

Sortieren nach

122

Werkzeug-Rangliste

Ranked by normalized popularity across sources. Rohwerte behalten ihr eigenes Zeitfenster; eine Änderung erfordert mindestens zwei Messungen in 7 Tagen.

  1. 1
    Aktiv

    Mini Program Engineering Skill Suite is an Agent Skill suite for evidence‑first mini‑program development. It helps agents bring WeChat or other mini‑program projects from vague intent to reliable engineering work: project intake, product…

    githubgeschätztes MomentumQuelle öffnen ↗

    Installieren git clone https://github.com/NocodeMrLi/mini-program-engineering-skill-suite ~/.claude/skills/mini-program-engineering-skill-suite

    39GitHub-Sterne
  2. 2

    Skill
    Aktiv

    Real-time execution trace and cost intelligence for Claude Code

    githubgemessenes WachstumQuelle öffnen ↗

    Installieren /plugin marketplace add DeibyGS/claudestat

    34GitHub-Sterne+1 (+3.0 %)
  3. 3
    Aktiv

    Make Claude Opus 4.8 behave like Claude Fable 5 — doctrine output style, drift-catching hooks, and an eval loop against golden Fable transcripts. Claude Code plugin.

    githubgemessenes WachstumQuelle öffnen ↗

    Installieren /plugin marketplace add rennf93/opus-fable-playbook

    33GitHub-Sternestabil
  4. 4
    Ruhend

    🔍 AI observability skill for Claude Code. Debug LangChain/LangGraph agents by fetching execution traces from LangSmith Studio directly in your terminal.

    githubgemessenes WachstumQuelle öffnen ↗

    Installieren git clone https://github.com/OthmanAdi/langsmith-fetch-skill ~/.claude/skills/langsmith-fetch-skill

    27GitHub-Sternestabil
  5. 5

    Agent
    Ruhend

    Policy-enforced observability and fail-closed guardrails for MCP/A2A multi-agent systems.

    githubgemessenes WachstumQuelle öffnen ↗

    26GitHub-Sternestabil
  6. 6
    Aktiv

    Sentry instrumentation skill for system-behavior tracking

    githubgemessenes WachstumQuelle öffnen ↗

    Installieren git clone https://github.com/tortastudios/sentry-instrumentation ~/.claude/skills/sentry-instrumentation

    24GitHub-Sternestabil
  7. 7
    Aktiv

    Deterministic, local-only audit reports for Claude Code AI agent sessions. Rust + SQLite. Zero network calls. MCP server for agents.

    githubgemessenes WachstumQuelle öffnen ↗

    22GitHub-Sterne+2 (+10.0 %)
  8. 8
    Aktiv

    Log Claude Code sessions to Opik, the open-source LLM observability and evaluation platform, built by Comet. Tracing, evaluation, and skills for observable AI applications.

    githubgemessenes WachstumQuelle öffnen ↗

    22GitHub-Sterne+1 (+4.8 %)
  9. 9

    Sonstige
    Aktiv

    AI-detector auditing toolkit: measures how often a detector flags genuine human writing, how stable its verdict is across seeds, and whether that verdict survives meaning-preserving edits. Detector-in-the-loop measurement harness. Claude…

    githubgemessenes WachstumQuelle öffnen ↗

    18GitHub-Sternestabil
  10. 10

    Skill
    Aktiv

    Measure prompt and skill improvements with blind A/B comparison.

    githubgemessenes WachstumQuelle öffnen ↗

    Installieren /plugin marketplace add shinpr/rashomon

    18GitHub-Sternestabil
  11. 11
    Aktiv

    Two Claude Code skills for adversarial code review (single-model PAR and multi-model MMAR with cross-critique to catch hallucinations), plus a fixture-based eval suite.

    githubgemessenes WachstumQuelle öffnen ↗

    Installieren /plugin marketplace add prime-radiant-inc/parallel-adversarial-review

    17GitHub-Sternestabil
  12. 12
    Ruhend

    MCP as a Judge: a behavioral MCP that strengthens AI coding assistants via explicit LLM evaluations

    mcpgemessenes WachstumQuelle öffnen ↗

    Installieren claude mcp add mcp-as-a-judge -- uvx mcp-as-a-judge

    17GitHub-Sternestabil
  13. 13
    Aktiv

    Make Claude write clearly, for everyone.

    githubgemessenes WachstumQuelle öffnen ↗

    Installieren /plugin marketplace add stefanobaghino/simple-output-styles

    16GitHub-Sternestabil
  14. 14

    Sonstige
    Aktiv

    A command-line tool to scaffold and manage enterprise-ready AI Agents powered by the A2A (Agent-to-Agent) protocol

    githubgemessenes WachstumQuelle öffnen ↗

    14GitHub-Sterne+1 (+7.7 %)
  15. 15
    Aktiv

    Claude Code plugin marketplace — agentic-engineering (spec-driven shape→decide→execute→measure→eval with adversarial review) + github-keeper (audit/elevate READMEs and make a repo open-source-ready).

    githubgemessenes WachstumQuelle öffnen ↗

    Installieren /plugin marketplace add GiustoPiedimonte/agentic-engineering-marketplace

    13GitHub-Sternestabil
  16. 16

    Sonstige
    Ruhend

    A Claude Code skill that adds a rubric-based eval layer to any agent project. Framework-agnostic — generates rubric, test cases, judge prompt, and harness. Returns a weighted score plus a judge-leniency signal.

    githubgemessenes WachstumQuelle öffnen ↗

    13GitHub-Sternestabil
  17. 17
    Aktiv

    Query Spanlens LLM observability from Cursor, Claude Desktop, or Continue via MCP.

    mcpgemessenes WachstumQuelle öffnen ↗

    Installieren claude mcp add mcp-server -- npx @spanlens/mcp-server

    12GitHub-Sternestabil
  18. 18

    Skill
    Aktiv

    Production-grade AI coding rules for Cursor and Claude Code. 15 rules + 9 doc templates + skills + agents + MCP setup. Drop into any project.

    githubgemessenes WachstumQuelle öffnen ↗

    Installieren git clone https://github.com/aiagentwithdhruv/ai-dev-stack ~/.claude/skills/ai-dev-stack

    10GitHub-Sterne+1 (+11.1 %)
  19. 19
    Aktiv

    Audit which Claude Code skills you actually use — surface dead installs and hallucinated invocations from your session logs.

    githubgemessenes WachstumQuelle öffnen ↗

    Installieren /plugin marketplace add sfrangulov/skill-graveyard

    10GitHub-Sterne+1 (+11.1 %)
  20. 20

    Skill
    Aktiv

    Turn one decision into a judged tournament of solutions, then pick the best — a Claude Code skill that generates candidates, auto-derives the rubric, judges independently, and returns a defensible winner.

    githubgemessenes WachstumQuelle öffnen ↗

    Installieren /plugin marketplace add CoriChui/bakeoff

    10GitHub-Sternestabil
  21. 21

    Sonstige
    Aktiv

    A Go-native framework for LLM agents, with OpenTelemetry observability built in.

    githubgemessenes WachstumQuelle öffnen ↗

    10GitHub-Sternestabil
  22. 22
    Ruhend

    Local open-source dev tool to debug, secure, and evaluate LLM agents. Provides static analysis, dynamic security checks, and runtime monitoring - integrates with Cursor and Claude Code.

    githubgemessenes WachstumQuelle öffnen ↗

    Installieren /plugin marketplace add cylestio/agent-inspector

    9GitHub-Sternestabil

Lern- und Referenzressourcen

Ranked by normalized popularity across sources. Diese Ressourcen bleiben getrennt zugänglich und fließen nicht in die Hauptwertung ein.

  1. 1
    AktivQuelle öffnen ↗

    50+ curated LLM observability tools PLUS 26 Agent Skills (several with runnable, unit-tested scripts) to build, evaluate, debug, secure & monitor reliable LLM apps. Tracing, evals, guardrails, LLMOps.

    github

    Installieren git clone https://github.com/ContextJet-ai/awesome-llm-observability ~/.claude/skills/awesome-llm-observability

    Ressourcegemessenes Wachstum
    32GitHub-Sterne+2 (+6.7 %)
  2. 2
    AktivQuelle öffnen ↗

    My complete journey to becoming an Agentic AI Engineer through structured learning, projects, experiments, and production-ready implementations of modern AI systems.

    githubRessourcegemessenes Wachstum
    18GitHub-Sternestabil
  3. 3
    AktivQuelle öffnen ↗

    Documentation-discovery telemetry for Claude Code — heat/cold maps, health grade, evidence-backed router fixes. 100% local, zero tokens. /tt

    github

    Installieren /plugin marketplace add Hedde/trigger_tree

    Ressourcegemessenes Wachstum
    14GitHub-Sterne+1 (+7.7 %)