Die Werkzeugbeschreibungen sind auf Englisch.

LLM-Beobachtbarkeit

Monitoring, Traces, Bewertung und Qualität von LLM-Anwendungen.

Anwendungsfall

Aktivität

Sortieren nach

122

Werkzeug-Rangliste

Ranked by known GitHub stars, highest first; archived repositories come after maintained repositories. Rohwerte behalten ihr eigenes Zeitfenster; eine Änderung erfordert mindestens zwei Messungen in 7 Tagen.

  1. 1
    Aktiv

    Multi-tier framework for evaluating AI agent skills with quality gates, semantic overlap detection, synthetic evaluation dataset generation, and live agent evaluation that measures how skills affect agent behavior.

    githubgemessenes WachstumQuelle öffnen ↗

    Installieren git clone https://github.com/NVIDIA/SkillEvaluator ~/.claude/skills/SkillEvaluator

    399GitHub-Sterne+51 (+14.7 %)
  2. 2
    Aktiv

    Dashboard for monitoring claude code sessions.

    githubgemessenes WachstumQuelle öffnen ↗

    Installieren /plugin marketplace add JayantDevkar/claude-code-karma

    323GitHub-Sterne+2 (+0.62 %)
  3. 3

    MCP
    Aktiv

    AURA is a production-tested SRE agent platform you can deploy in minutes. AURA handles the guardrails, APIs, state management, streaming, and failure handling required to put AI to work safely on production infrastructure.

    githubgemessenes WachstumQuelle öffnen ↗

    291GitHub-Sterne+33 (+12.8 %)
  4. 4
    Aktiv

    🪢 Langfuse documentation -- Langfuse is the open source LLM Engineering Platform. Observability, evals, prompt management, playground and metrics to debug and improve LLM apps

    githubgemessenes WachstumQuelle öffnen ↗

    Installieren git clone https://github.com/langfuse/langfuse-docs ~/.claude/skills/langfuse-docs

    239GitHub-Sterne+3 (+1.3 %)
  5. 5
    Aktiv

    🟪 Open-source runtime that ships any LangGraph or Google ADK agent as a production-ready FastAPI service. Bundled , AG-UI copilotkit API, chat UI, 15+ guardrails, MCP, OpenTelemetry, OIDC. One pip install. Self-hosted, no vendor lock-in.

    githubgemessenes WachstumQuelle öffnen ↗

    Installieren git clone https://github.com/Idun-Group/idun-agent-platform ~/.claude/skills/idun-agent-platform

    199GitHub-Sternestabil
  6. 6

    Agent
    Aktiv

    The Operating System for Scalable Enterprise AI Agents - Run, orchestrate, and deploy Compliant Enterprise AI Agents at scale across frameworks, without lock-in, rewrites or fragile glue code. Native support for MCP, A2A. Interface with all…

    githubgemessenes WachstumQuelle öffnen ↗

    166GitHub-Sterne+29 (+21.2 %)
  7. 7

    Sonstige
    Aktiv

    Evidence-first infrastructure for reproducible, isolated executable evaluation and online rewards.

    githubgemessenes WachstumQuelle öffnen ↗

    163GitHub-Sterne+47 (+40.5 %)
  8. 8

    Skill
    Aktiv

    Research-backed, eval-driven skills for AI agents

    githubgemessenes WachstumQuelle öffnen ↗

    Installieren git clone https://github.com/ZSeven-W/craft-skills ~/.claude/skills/craft-skills

    145GitHub-Sterne+14 (+10.7 %)
  9. 9
    Aktiv

    Skills, prompts, and instructions for building AI agents on top of Dynatrace production context

    githubgemessenes WachstumQuelle öffnen ↗

    Installieren git clone https://github.com/Dynatrace/dynatrace-for-ai ~/.claude/skills/dynatrace-for-ai

    135GitHub-Sterne+4 (+3.1 %)
  10. 10
    Aktiv

    🐙 ADLC Team Skills — Agentic SDLC for Engineering Teams

    githubgemessenes WachstumQuelle öffnen ↗

    Installieren git clone https://github.com/tikalk/adlc-team-skills ~/.claude/skills/adlc-team-skills

    132GitHub-Sterne+1 (+0.76 %)
  11. 11

    Sonstige
    Aktiv

    Self-hosted AI SRE for Kubernetes — zero-instrumentation eBPF observability plus a copilot that fixes issues through guardrailed, self-verifying actions. BYO-LLM, air-gapped capable.

    githubgemessenes WachstumQuelle öffnen ↗

    131GitHub-Sterne-4 (-3.0 %)
  12. 12
    Aktiv

    74 open-source Agent Skills for Claude Code and Codex: AI SEO, AEO and GEO, code review with an A-F ship grade, CI gates, AI evals, design systems, conversion copy, Instagram growth, iOS and Android app shipping, creator rights, and…

    githubgemessenes WachstumQuelle öffnen ↗

    129GitHub-Sterne-4 (-3.0 %)
  13. 13
    Aktiv

    Self improving agents through iterations

    githubgemessenes WachstumQuelle öffnen ↗

    105GitHub-Sterne+1 (+0.96 %)
  14. 14
    Aktiv

    A Model Context Protocol (MCP) server for Langfuse, enabling AI agents to query Langfuse trace data for enhanced debugging and observability

    githubgemessenes WachstumQuelle öffnen ↗

    105GitHub-Sternestabil
  15. 15
    Aktiv

    Vendor-neutral OpenTelemetry skills for AI coding agents, grounded in upstream sources

    githubgemessenes WachstumQuelle öffnen ↗

    Installieren git clone https://github.com/ollygarden/opentelemetry-agent-skills ~/.claude/skills/opentelemetry-agent-skills

    97GitHub-Sterne+1 (+1.0 %)
  16. 16

    Skill
    Ruhend

    Don't make LLMs honest. Make every factual claim auditable. — An LLM Claim Auditing Layer with T1-T7 truth gradients. 98.1% business effectiveness on LiarBench v0.2.

    githubgemessenes WachstumQuelle öffnen ↗

    Installieren git clone https://github.com/lc198707/anti-lie ~/.claude/skills/anti-lie

    89GitHub-Sternestabil
  17. 17
    Aktiv

    MCP server for Langfuse LLM observability — trace and observation analysis.

    mcpgemessenes WachstumQuelle öffnen ↗

    Installieren claude mcp add langfuse -- npx langfuse-observability-mcp-server

    80GitHub-Sterne+2 (+2.6 %)
  18. 18

    Skill
    Aktiv

    local-first analytics for AI agent skills

    githubgemessenes WachstumQuelle öffnen ↗

    Installieren git clone https://github.com/crafter-station/skill-kit ~/.claude/skills/skill-kit

    77GitHub-Sterne+1 (+1.3 %)
  19. 19

    Skill
    Aktiv

    Deletion-first Agent Skill for removing test bloat, verification theater, and speculative fallbacks while preserving behavior.

    githubgemessenes WachstumQuelle öffnen ↗

    Installieren git clone https://github.com/MrZoyo/deslop-GPT ~/.claude/skills/deslop-GPT

    76GitHub-Sterne+48 (+171.4 %)
  20. 20
    Aktiv

    Agent Skill evaluation harness for paired variants, trace artifacts, and runner adapters

    githubgemessenes WachstumQuelle öffnen ↗

    Installieren git clone https://github.com/adewale/skill-eval-harness ~/.claude/skills/skill-eval-harness

    73GitHub-Sterne+4 (+5.8 %)
  21. 21

    Sonstige
    Aktiv

    AgentX python SDK. Build multi-agent AI workforce. Run evaluation. Trace your agent. Full Observability.

    githubgemessenes WachstumQuelle öffnen ↗

    68GitHub-Sternestabil
  22. 22
    Aktiv

    Open benchmark for Claude Code SEO skills — real headless execution against fixture sites with planted-defect answer keys. Deterministic scoring, pre-registered rubric.

    githubgemessenes WachstumQuelle öffnen ↗

    Installieren git clone https://github.com/aleclindz/seo-skill-bench ~/.claude/skills/seo-skill-bench

    51GitHub-Sterne-5 (-8.9 %)
  23. 23

    MCP
    Aktiv

    Open-source, self-hosted control plane for OpenClaw and Hermes AI-agent fleets on Docker/Kubernetes — REST, CLI, and MCP.

    githubgemessenes WachstumQuelle öffnen ↗

    50GitHub-Sterne+2 (+4.2 %)
  24. 24

    Skill
    Aktiv

    Agent skills for Arize — datasets, experiments, and traces via the ax CLI

    githubgemessenes WachstumQuelle öffnen ↗

    Installieren git clone https://github.com/Arize-ai/arize-skills ~/.claude/skills/arize-skills

    47GitHub-Sternestabil

Lern- und Referenzressourcen

Ranked by known GitHub stars, highest first; archived repositories come after maintained repositories. Diese Ressourcen bleiben getrennt zugänglich und fließen nicht in die Hauptwertung ein.

  1. 1
    AktivQuelle öffnen ↗

    Curated + scored map of official MCP servers and agents for DevOps, Cloud, SRE, and Platform Engineering — every entry rated on production access, approval gates, and audit evidence.

    githubRessourcegemessenes Wachstum
    77GitHub-Sternestabil