Die Werkzeugbeschreibungen sind auf Englisch.
LLM-Beobachtbarkeit
Monitoring, Traces, Bewertung und Qualität von LLM-Anwendungen.
Anwendungsfall
Aktivität
Sortieren nach
122
Werkzeug-Rangliste
Nach gemessenem Wachstum und quellenübergreifend normalisiert sortiert. Rohwerte behalten ihr eigenes Zeitfenster; eine Änderung erfordert mindestens zwei Messungen in 7 Tagen.
- 1Aktiv
aura
MCPAURA is a production-tested SRE agent platform you can deploy in minutes. AURA handles the guardrails, APIs, state management, streaming, and failure handling required to put AI to work safely on production infrastructure.
githubgemessenes WachstumQuelle öffnen ↗
254GitHub-Sterne-3 (-1.2 %) - 2Aktiv
idun-agent-platform
Skill🟪 Open-source runtime that ships any LangGraph or Google ADK agent as a production-ready FastAPI service. Bundled , AG-UI copilotkit API, chat UI, 15+ guardrails, MCP, OpenTelemetry, OIDC. One pip install. Self-hosted, no vendor lock-in.
githubgemessenes WachstumQuelle öffnen ↗
Installieren
git clone https://github.com/Idun-Group/idun-agent-platform ~/.claude/skills/idun-agent-platform198GitHub-Sterne-1 (-0.50 %) - 3Aktiv
rocketplaneIO
SonstigeSelf-hosted AI SRE for Kubernetes — zero-instrumentation eBPF observability plus a copilot that fixes issues through guardrailed, self-verifying actions. BYO-LLM, air-gapped capable.
githubgemessenes WachstumQuelle öffnen ↗
131GitHub-Sterne-4 (-3.0 %) - 4Aktiv
74 open-source Agent Skills for Claude Code and Codex: AI SEO, AEO and GEO, code review with an A-F ship grade, CI gates, AI evals, design systems, conversion copy, Instagram growth, iOS and Android app shipping, creator rights, and…
githubgemessenes WachstumQuelle öffnen ↗
129GitHub-Sterne-4 (-3.0 %) - 5Aktiv
langfuse-mcp
MCPA Model Context Protocol (MCP) server for Langfuse, enabling AI agents to query Langfuse trace data for enhanced debugging and observability
githubgemessenes WachstumQuelle öffnen ↗
105GitHub-Sternestabil - 6Ruhend
anti-lie
SkillDon't make LLMs honest. Make every factual claim auditable. — An LLM Claim Auditing Layer with T1-T7 truth gradients. 98.1% business effectiveness on LiarBench v0.2.
githubgemessenes WachstumQuelle öffnen ↗
Installieren
git clone https://github.com/lc198707/anti-lie ~/.claude/skills/anti-lie89GitHub-Sternestabil - 7Aktiv
MCP server for Langfuse LLM observability — trace and observation analysis.
mcpgemessenes WachstumQuelle öffnen ↗
Installieren
claude mcp add langfuse -- npx langfuse-observability-mcp-server78GitHub-Sternestabil - 8Aktiv
AgentX-Python
SonstigeAgentX python SDK. Build multi-agent AI workforce. Run evaluation. Trace your agent. Full Observability.
githubgemessenes WachstumQuelle öffnen ↗
68GitHub-Sternestabil - 9Aktiv
seo-skill-bench
SkillOpen benchmark for Claude Code SEO skills — real headless execution against fixture sites with planted-defect answer keys. Deterministic scoring, pre-registered rubric.
githubgemessenes WachstumQuelle öffnen ↗
Installieren
git clone https://github.com/aleclindz/seo-skill-bench ~/.claude/skills/seo-skill-bench51GitHub-Sterne-5 (-8.9 %) - 10Aktiv
arize-skills
SkillAgent skills for Arize — datasets, experiments, and traces via the ax CLI
githubgemessenes WachstumQuelle öffnen ↗
Installieren
git clone https://github.com/Arize-ai/arize-skills ~/.claude/skills/arize-skills47GitHub-Sternestabil - 11Aktiv
opus-fable-playbook
SkillMake Claude Opus 4.8 behave like Claude Fable 5 — doctrine output style, drift-catching hooks, and an eval loop against golden Fable transcripts. Claude Code plugin.
githubgemessenes WachstumQuelle öffnen ↗
Installieren
/plugin marketplace add rennf93/opus-fable-playbook33GitHub-Sternestabil - 12Ruhend
🔍 AI observability skill for Claude Code. Debug LangChain/LangGraph agents by fetching execution traces from LangSmith Studio directly in your terminal.
githubgemessenes WachstumQuelle öffnen ↗
Installieren
git clone https://github.com/OthmanAdi/langsmith-fetch-skill ~/.claude/skills/langsmith-fetch-skill27GitHub-Sternestabil - 13Ruhend
astragraph
AgentPolicy-enforced observability and fail-closed guardrails for MCP/A2A multi-agent systems.
githubgemessenes WachstumQuelle öffnen ↗
26GitHub-Sternestabil - 14Aktiv
Sentry instrumentation skill for system-behavior tracking
githubgemessenes WachstumQuelle öffnen ↗
Installieren
git clone https://github.com/tortastudios/sentry-instrumentation ~/.claude/skills/sentry-instrumentation24GitHub-Sternestabil - 15Aktiv
untell
SonstigeAI-detector auditing toolkit: measures how often a detector flags genuine human writing, how stable its verdict is across seeds, and whether that verdict survives meaning-preserving edits. Detector-in-the-loop measurement harness. Claude…
githubgemessenes WachstumQuelle öffnen ↗
18GitHub-Sternestabil - 16Aktiv
rashomon
SkillMeasure prompt and skill improvements with blind A/B comparison.
githubgemessenes WachstumQuelle öffnen ↗
Installieren
/plugin marketplace add shinpr/rashomon18GitHub-Sternestabil - 17Aktiv
Two Claude Code skills for adversarial code review (single-model PAR and multi-model MMAR with cross-critique to catch hallucinations), plus a fixture-based eval suite.
githubgemessenes WachstumQuelle öffnen ↗
Installieren
/plugin marketplace add prime-radiant-inc/parallel-adversarial-review17GitHub-Sternestabil - 18Ruhend
MCP as a Judge: a behavioral MCP that strengthens AI coding assistants via explicit LLM evaluations
mcpgemessenes WachstumQuelle öffnen ↗
Installieren
claude mcp add mcp-as-a-judge -- uvx mcp-as-a-judge17GitHub-Sternestabil - 19Aktiv
simple-output-styles
SkillMake Claude write clearly, for everyone.
githubgemessenes WachstumQuelle öffnen ↗
Installieren
/plugin marketplace add stefanobaghino/simple-output-styles16GitHub-Sternestabil - 20Aktiv
Claude Code plugin marketplace — agentic-engineering (spec-driven shape→decide→execute→measure→eval with adversarial review) + github-keeper (audit/elevate READMEs and make a repo open-source-ready).
githubgemessenes WachstumQuelle öffnen ↗
Installieren
/plugin marketplace add GiustoPiedimonte/agentic-engineering-marketplace13GitHub-Sternestabil - 21Ruhend
eval-layer
SonstigeA Claude Code skill that adds a rubric-based eval layer to any agent project. Framework-agnostic — generates rubric, test cases, judge prompt, and harness. Returns a weighted score plus a judge-leniency signal.
githubgemessenes WachstumQuelle öffnen ↗
13GitHub-Sternestabil - 22Aktiv
Query Spanlens LLM observability from Cursor, Claude Desktop, or Continue via MCP.
mcpgemessenes WachstumQuelle öffnen ↗
Installieren
claude mcp add mcp-server -- npx @spanlens/mcp-server12GitHub-Sternestabil - 23Aktiv
bakeoff
SkillTurn one decision into a judged tournament of solutions, then pick the best — a Claude Code skill that generates candidates, auto-derives the rubric, judges independently, and returns a defensible winner.
githubgemessenes WachstumQuelle öffnen ↗
Installieren
/plugin marketplace add CoriChui/bakeoff10GitHub-Sternestabil
Lern- und Referenzressourcen
Nach gemessenem Wachstum und quellenübergreifend normalisiert sortiert. Diese Ressourcen bleiben getrennt zugänglich und fließen nicht in die Hauptwertung ein.
- 1AktivQuelle öffnen ↗
Curated + scored map of official MCP servers and agents for DevOps, Cloud, SRE, and Platform Engineering — every entry rated on production access, approval gates, and audit evidence.
githubRessourcegemessenes Wachstum77GitHub-Sternestabil - 2AktivQuelle öffnen ↗
Agentic_AI_Engineer
AgentMy complete journey to becoming an Agentic AI Engineer through structured learning, projects, experiments, and production-ready implementations of modern AI systems.
githubRessourcegemessenes Wachstum18GitHub-Sternestabil