Die Werkzeugbeschreibungen sind auf Englisch.
LLM-Beobachtbarkeit
Monitoring, Traces, Bewertung und Qualität von LLM-Anwendungen.
Anwendungsfall
Aktivität
Sortieren nach
123
Werkzeug-Rangliste
Nach gemessenem Wachstum und quellenübergreifend normalisiert sortiert. Rohwerte behalten ihr eigenes Zeitfenster; eine Änderung erfordert mindestens zwei Messungen in 7 Tagen.
- 1Aktiv
SkillCorpus
SkillOpen-source infrastructure that turns scattered SKILL.md files into curated, retrieval-ready agent-skill corpora—with retrieval and evaluation tooling included.
githubgemessenes WachstumQuelle öffnen ↗
Installieren
git clone https://github.com/EverMind-AI/SkillCorpus ~/.claude/skills/SkillCorpus360GitHub-Sterne+168 (+87.5 %) - 2Aktiv
SkillEvaluator
SkillMulti-tier framework for evaluating AI agent skills with quality gates, semantic overlap detection, synthetic evaluation dataset generation, and live agent evaluation that measures how skills affect agent behavior.
githubgemessenes WachstumQuelle öffnen ↗
Installieren
git clone https://github.com/NVIDIA/SkillEvaluator ~/.claude/skills/SkillEvaluator363GitHub-Sterne+58 (+19.0 %) - 3Aktiv
craft-skills
SkillResearch-backed, eval-driven skills for AI agents
githubgemessenes WachstumQuelle öffnen ↗
Installieren
git clone https://github.com/ZSeven-W/craft-skills ~/.claude/skills/craft-skills138GitHub-Sterne+21 (+17.9 %) - 4Aktiv
agent-kernel
AgentThe Operating System for Scalable Enterprise AI Agents - Run, orchestrate, and deploy Compliant Enterprise AI Agents at scale across frameworks, without lock-in, rewrites or fragile glue code. Native support for MCP, A2A. Interface with all…
githubgemessenes WachstumQuelle öffnen ↗
146GitHub-Sterne+13 (+9.8 %) - 5Aktiv
skill-eval-harness
SkillAgent Skill evaluation harness for paired variants, trace artifacts, and runner adapters
githubgemessenes WachstumQuelle öffnen ↗
Installieren
git clone https://github.com/adewale/skill-eval-harness ~/.claude/skills/skill-eval-harness72GitHub-Sterne+4 (+5.9 %) - 6Aktiv
dynatrace-for-ai
SkillSkills, prompts, and instructions for building AI agents on top of Dynatrace production context
githubgemessenes WachstumQuelle öffnen ↗
Installieren
git clone https://github.com/Dynatrace/dynatrace-for-ai ~/.claude/skills/dynatrace-for-ai131GitHub-Sterne+5 (+4.0 %) - 7Aktiv
databuff
AgentDataBuff is an AI-native APM built on Opentelemetry,with multi-agent troubleshooting out of the box.
githubgemessenes WachstumQuelle öffnen ↗
627GitHub-Sterne+22 (+3.6 %) - 8Aktiv
Vendor-neutral OpenTelemetry skills for AI coding agents, grounded in upstream sources
githubgemessenes WachstumQuelle öffnen ↗
Installieren
git clone https://github.com/ollygarden/opentelemetry-agent-skills ~/.claude/skills/opentelemetry-agent-skills96GitHub-Sterne+3 (+3.2 %) - 9Aktiv
yao-meta-skill
SkillYAO = Yielding AI Outcomes. A rigorous engineering, evaluation, governance, and portability system for reusable agent skills.
githubgemessenes WachstumQuelle öffnen ↗
Installieren
git clone https://github.com/yaojingang/yao-meta-skill ~/.claude/skills/yao-meta-skill2 565GitHub-Sterne+58 (+2.3 %) - 10Aktiv
agent-skills-eval
SkillA test runner for agentskills.io-style AI agent skills
githubgemessenes WachstumQuelle öffnen ↗
Installieren
git clone https://github.com/darkrishabh/agent-skills-eval ~/.claude/skills/agent-skills-eval713GitHub-Sterne+12 (+1.7 %) - 11Aktiv
skill-kit
Skilllocal-first analytics for AI agent skills
githubgemessenes WachstumQuelle öffnen ↗
Installieren
git clone https://github.com/crafter-station/skill-kit ~/.claude/skills/skill-kit77GitHub-Sterne+1 (+1.3 %) - 12Aktiv
MCP server for Langfuse LLM observability — trace and observation analysis.
mcpgemessenes WachstumQuelle öffnen ↗
Installieren
claude mcp add langfuse -- npx langfuse-observability-mcp-server78GitHub-Sterne+1 (+1.3 %) - 13Aktiv
A Go framework for building production agent systems with graph workflows, tools, memory, A2A, AG-UI, MCP, evaluation, and observability.
githubgemessenes WachstumQuelle öffnen ↗
1 752GitHub-Sterne+22 (+1.3 %) - 14Aktiv
agent-skill-creator
SkillBuild tested agent skills and govern their lifecycle through a user-defined marketplace: evidence, discovery, updates, rollback, quarantine, and 17-platform distribution.
githubgemessenes WachstumQuelle öffnen ↗
Installieren
git clone https://github.com/FrancyJGLisboa/agent-skill-creator ~/.claude/skills/agent-skill-creator2 358GitHub-Sterne+29 (+1.2 %) - 15Aktiv
fable-method
SkillThe Fable Workflow: how Claude Fable 5 worked, distilled into skills any model can run, with the eval that keeps it honest. Think / act / prove.
githubgemessenes WachstumQuelle öffnen ↗
Installieren
git clone https://github.com/Sahir619/fable-method ~/.claude/skills/fable-method2 269GitHub-Sterne+22 (+0.98 %) - 16105GitHub-Sterne+1 (+0.96 %)
- 17Aktiv
OpenJudge
SkillOpenJudge: A Unified Framework for Holistic Evaluation and Quality Rewards
githubgemessenes WachstumQuelle öffnen ↗
Installieren
git clone https://github.com/agentscope-ai/OpenJudge ~/.claude/skills/OpenJudge807GitHub-Sterne+7 (+0.88 %) - 18Aktiv
langfuse-docs
Skill🪢 Langfuse documentation -- Langfuse is the open source LLM Engineering Platform. Observability, evals, prompt management, playground and metrics to debug and improve LLM apps
githubgemessenes WachstumQuelle öffnen ↗
Installieren
git clone https://github.com/langfuse/langfuse-docs ~/.claude/skills/langfuse-docs237GitHub-Sterne+2 (+0.85 %) - 19Aktiv
aura
MCPAURA is a production-tested SRE agent platform you can deploy in minutes. AURA handles the guardrails, APIs, state management, streaming, and failure handling required to put AI to work safely on production infrastructure.
githubgemessenes WachstumQuelle öffnen ↗
258GitHub-Sterne+2 (+0.78 %) - 20Aktiv
claude-code-karma
SkillDashboard for monitoring claude code sessions.
githubgemessenes WachstumQuelle öffnen ↗
Installieren
/plugin marketplace add JayantDevkar/claude-code-karma321GitHub-Sterne+2 (+0.63 %) - 21Aktiv
SkillForge
SkillA skill creator that proves its skills work. Evidence-driven skill creation for Claude Code and Codex: baseline-tested generation, per-skill regression evals, ecosystem doctor, cross-runtime compile, and an opt-in proactive advisor.
githubgemessenes WachstumQuelle öffnen ↗
Installieren
git clone https://github.com/tripleyak/SkillForge ~/.claude/skills/SkillForge884GitHub-Sterne+5 (+0.57 %) - 22Aktiv
promptfoo
SonstigeTest your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare performance of GPT, Claude, Gemini, DeepSeek, and more. Simple declarative configs with command line and CI/CD integration. Used by OpenAI…
githubgemessenes WachstumQuelle öffnen ↗
24 710GitHub-Sterne+137 (+0.56 %) - 23Aktiv
openobserve
SonstigeOpen source observability platform for logs, metrics, traces, frontend monitoring, pipelines and LLM observability. A sophisticated, simple and highly performant alternative to Datadog, Splunk, and Elasticsearch with 140x lower storage…
githubgemessenes WachstumQuelle öffnen ↗
21 594GitHub-Sterne+119 (+0.55 %) - 249 079GitHub-Sterne+46 (+0.51 %)
Lern- und Referenzressourcen
Nach gemessenem Wachstum und quellenübergreifend normalisiert sortiert. Diese Ressourcen bleiben getrennt zugänglich und fließen nicht in die Hauptwertung ein.
- 1AktivQuelle öffnen ↗
Curated + scored map of official MCP servers and agents for DevOps, Cloud, SRE, and Platform Engineering — every entry rated on production access, approval gates, and audit evidence.
githubRessourcegemessenes Wachstum77GitHub-Sterne+1 (+1.3 %)