Die Werkzeugbeschreibungen sind auf Englisch.

LLM-Beobachtbarkeit

Monitoring, Traces, Bewertung und Qualität von LLM-Anwendungen.

Anwendungsfall

Aktivität

Sortieren nach

76

Werkzeug-Rangliste

Ranked by creation date, newest first; undated entries come last. Rohwerte behalten ihr eigenes Zeitfenster; eine Änderung erfordert mindestens zwei Messungen in 7 Tagen.

  1. 1
    Aktiv

    Open benchmark for Claude Code SEO skills — real headless execution against fixture sites with planted-defect answer keys. Deterministic scoring, pre-registered rubric.

    githubgemessenes WachstumQuelle öffnen ↗

    Installieren git clone https://github.com/aleclindz/seo-skill-bench ~/.claude/skills/seo-skill-bench

    51GitHub-Sterne-5 (-8.9 %)
  2. 2

    Sonstige
    Aktiv

    AI-detector auditing toolkit: measures how often a detector flags genuine human writing, how stable its verdict is across seeds, and whether that verdict survives meaning-preserving edits. Detector-in-the-loop measurement harness. Claude…

    githubgemessenes WachstumQuelle öffnen ↗

    18GitHub-Sternestabil
  3. 3
    Aktiv

    Multi-tier framework for evaluating AI agent skills with quality gates, semantic overlap detection, synthetic evaluation dataset generation, and live agent evaluation that measures how skills affect agent behavior.

    githubgemessenes WachstumQuelle öffnen ↗

    Installieren git clone https://github.com/NVIDIA/SkillEvaluator ~/.claude/skills/SkillEvaluator

    399GitHub-Sterne+51 (+14.7 %)
  4. 4

    Agent
    Aktiv

    DataBuff is an AI-native APM built on Opentelemetry,with multi-agent troubleshooting out of the box.

    githubgemessenes WachstumQuelle öffnen ↗

    649GitHub-Sterne+39 (+6.4 %)
  5. 5
    Aktiv

    Optimize any AI agent’s skills, tools/MCP, and prompts against your own evals.

    githubgeschätztes MomentumQuelle öffnen ↗

    47GitHub-Sterne
  6. 6
    Aktiv

    Agent Skill evaluation harness for paired variants, trace artifacts, and runner adapters

    githubgemessenes WachstumQuelle öffnen ↗

    Installieren git clone https://github.com/adewale/skill-eval-harness ~/.claude/skills/skill-eval-harness

    73GitHub-Sterne+4 (+5.8 %)
  7. 7

    Agent
    Aktiv

    The agent framework where the model never holds the trigger — every consequential action clears your policy first, waits for a human when it matters, and lands on a record you can verify. Build on it, or put it around the agent you already…

    githubgemessenes WachstumQuelle öffnen ↗

    1GitHub-Sternestabil
  8. 8

    Skill
    Aktiv

    Axiom is a curated marketplace of shared plugins for Claude Code and Codex.

    githubgemessenes WachstumQuelle öffnen ↗

    Installieren git clone https://github.com/netopsengineer/axiom ~/.claude/skills/axiom

    5GitHub-Sternestabil
  9. 9

    Sonstige
    Aktiv

    Local-first Agentic AI Infrastructure Platform with LLM Gateway, Agent DAG Runtime, MCP Tool Hub, A2A Agent Mesh, Local RAG, Tool Sandbox and Observability.

    githubgemessenes WachstumQuelle öffnen ↗

    1GitHub-Sternestabil
  10. 10

    Skill
    Aktiv

    Anchor — the production-grade AGENTS.md template for AI coding agents. 51 sections of battle-tested rules. Keep your agents grounded.

    githubgemessenes WachstumQuelle öffnen ↗

    Installieren git clone https://github.com/peva3/anchor ~/.claude/skills/anchor

    8GitHub-Sternestabil
  11. 11

    Sonstige
    Aktiv

    Open-source observability tool that uses AI agents to self-heal your software

    githubgemessenes WachstumQuelle öffnen ↗

    1 404GitHub-Sterne+1 (+0.07 %)
  12. 12

    Skill
    Aktiv

    Glanceable Claude Code and Codex state in your terminal tabs: white=idle, blue=working, orange=waiting. Multi-terminal (iTerm2, WezTerm, AI Power Term), a live status bar with Anthropic usage limits, /sfl and /nil window save-and-restore,…

    githubgemessenes WachstumQuelle öffnen ↗

    Installieren git clone https://github.com/wasulajr/headsup ~/.claude/skills/headsup

    1GitHub-Sternestabil
  13. 13
    Aktiv

    74 open-source Agent Skills for Claude Code and Codex: AI SEO, AEO and GEO, code review with an A-F ship grade, CI gates, AI evals, design systems, conversion copy, Instagram growth, iOS and Android app shipping, creator rights, and…

    githubgemessenes WachstumQuelle öffnen ↗

    129GitHub-Sterne-4 (-3.0 %)
  14. 14

    Sonstige
    Aktiv

    A Go-native framework for LLM agents, with OpenTelemetry observability built in.

    githubgemessenes WachstumQuelle öffnen ↗

    10GitHub-Sternestabil
  15. 15

    Skill
    Aktiv

    Real-time execution trace and cost intelligence for Claude Code

    githubgemessenes WachstumQuelle öffnen ↗

    Installieren /plugin marketplace add DeibyGS/claudestat

    34GitHub-Sternestabil
  16. 16
    Aktiv

    Query Spanlens LLM observability from Cursor, Claude Desktop, or Continue via MCP.

    mcpgemessenes WachstumQuelle öffnen ↗

    Installieren claude mcp add mcp-server -- npx @spanlens/mcp-server

    12GitHub-Sternestabil
  17. 17

    Agent
    Aktiv

    MemroOS / memroos: memory OS and governance layer for AI agents, agent workflows, dispatch, proof, and context continuity.

    githubgemessenes WachstumQuelle öffnen ↗

    7GitHub-Sternestabil
  18. 18
    Aktiv

    YAO = Yielding AI Outcomes. A rigorous engineering, evaluation, governance, and portability system for reusable agent skills.

    githubgemessenes WachstumQuelle öffnen ↗

    Installieren git clone https://github.com/yaojingang/yao-meta-skill ~/.claude/skills/yao-meta-skill

    2 589GitHub-Sterne+38 (+1.5 %)
  19. 19
    Aktiv

    Skills, prompts, and instructions for building AI agents on top of Dynatrace production context

    githubgemessenes WachstumQuelle öffnen ↗

    Installieren git clone https://github.com/Dynatrace/dynatrace-for-ai ~/.claude/skills/dynatrace-for-ai

    135GitHub-Sterne+4 (+3.1 %)
  20. 20
    Aktiv

    AgentGateway — independent third-party profile of a public API surface, by API Evangelist. AgentGateway is an open-source, AI-native proxy and gateway for routing, observing, and governing traffic to and from AI agents, LLM providers, and…

    githubgemessenes WachstumQuelle öffnen ↗

    0GitHub-Sternestabil
  21. 21

    Skill
    Aktiv

    Marketing Mix Model expert agent plugin for Claude Code - pymc-marketing v0.18.2+

    githubgemessenes WachstumQuelle öffnen ↗

    Installieren /plugin marketplace add Yakoub-ai/agent-mmm

    4GitHub-Sternestabil
  22. 22

    Agent
    Aktiv

    Provide clear documentation for AgentStack’s MCP protocol, plugins, and ecosystem API with usage examples and tool references.

    githubgemessenes WachstumQuelle öffnen ↗

    0GitHub-Sternestabil
  23. 23

    MCP
    Aktiv

    Open-source, self-hosted control plane for OpenClaw and Hermes AI-agent fleets on Docker/Kubernetes — REST, CLI, and MCP.

    githubgemessenes WachstumQuelle öffnen ↗

    50GitHub-Sterne+2 (+4.2 %)
  24. 24
    Aktiv

    Vendor-neutral OpenTelemetry skills for AI coding agents, grounded in upstream sources

    githubgemessenes WachstumQuelle öffnen ↗

    Installieren git clone https://github.com/ollygarden/opentelemetry-agent-skills ~/.claude/skills/opentelemetry-agent-skills

    97GitHub-Sterne+1 (+1.0 %)

Lern- und Referenzressourcen

Ranked by creation date, newest first; undated entries come last. Diese Ressourcen bleiben getrennt zugänglich und fließen nicht in die Hauptwertung ein.

  1. 1
    AktivQuelle öffnen ↗

    Curated + scored map of official MCP servers and agents for DevOps, Cloud, SRE, and Platform Engineering — every entry rated on production access, approval gates, and audit evidence.

    githubRessourcegemessenes Wachstum
    77GitHub-Sternestabil