Die Werkzeugbeschreibungen sind auf Englisch.
LLM-Beobachtbarkeit
Monitoring, Traces, Bewertung und Qualität von LLM-Anwendungen.
Anwendungsfall
Aktivität
Sortieren nach
123
Werkzeug-Rangliste
Nach gemessenem Wachstum und quellenübergreifend normalisiert sortiert. Rohwerte behalten ihr eigenes Zeitfenster; eine Änderung erfordert mindestens zwei Messungen in 7 Tagen.
- 1Aktiv
SkillCorpus
SkillOpen-source infrastructure that turns scattered SKILL.md files into curated, retrieval-ready agent-skill corpora—with retrieval and evaluation tooling included.
githubgemessenes WachstumQuelle öffnen ↗
Installieren
git clone https://github.com/EverMind-AI/SkillCorpus ~/.claude/skills/SkillCorpus360GitHub-Sterne+168 (+87.5 %) - 2Aktiv
SkillEvaluator
SkillMulti-tier framework for evaluating AI agent skills with quality gates, semantic overlap detection, synthetic evaluation dataset generation, and live agent evaluation that measures how skills affect agent behavior.
githubgemessenes WachstumQuelle öffnen ↗
Installieren
git clone https://github.com/NVIDIA/SkillEvaluator ~/.claude/skills/SkillEvaluator363GitHub-Sterne+58 (+19.0 %) - 3Aktiv
craft-skills
SkillResearch-backed, eval-driven skills for AI agents
githubgemessenes WachstumQuelle öffnen ↗
Installieren
git clone https://github.com/ZSeven-W/craft-skills ~/.claude/skills/craft-skills138GitHub-Sterne+21 (+17.9 %) - 4Aktiv
deslop-GPT
SkillDeletion-first Agent Skill for removing test bloat, verification theater, and speculative fallbacks while preserving behavior.
githubgemessenes WachstumQuelle öffnen ↗
Installieren
git clone https://github.com/MrZoyo/deslop-GPT ~/.claude/skills/deslop-GPT38GitHub-Sterne+10 (+35.7 %) - 5Aktiv
promptfoo
SonstigeTest your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare performance of GPT, Claude, Gemini, DeepSeek, and more. Simple declarative configs with command line and CI/CD integration. Used by OpenAI…
githubgemessenes WachstumQuelle öffnen ↗
24 710GitHub-Sterne+137 (+0.56 %) - 6Aktiv
mastra
SonstigeMastra is the modern TypeScript framework for AI-powered applications and agents.
githubgemessenes WachstumQuelle öffnen ↗
27 601GitHub-Sterne+128 (+0.47 %) - 7Aktiv
openobserve
SonstigeOpen source observability platform for logs, metrics, traces, frontend monitoring, pipelines and LLM observability. A sophisticated, simple and highly performant alternative to Datadog, Splunk, and Elasticsearch with 140x lower storage…
githubgemessenes WachstumQuelle öffnen ↗
21 594GitHub-Sterne+119 (+0.55 %) - 8Aktiv
yao-meta-skill
SkillYAO = Yielding AI Outcomes. A rigorous engineering, evaluation, governance, and portability system for reusable agent skills.
githubgemessenes WachstumQuelle öffnen ↗
Installieren
git clone https://github.com/yaojingang/yao-meta-skill ~/.claude/skills/yao-meta-skill2 565GitHub-Sterne+58 (+2.3 %) - 9Aktiv
nudge
MCPA typed, replayable, budget-aware programming language for LLM agents. Compiles to Python & TypeScript.
githubgemessenes WachstumQuelle öffnen ↗
2GitHub-Sterne+1 (+100.0 %) - 10Aktiv
agent-stack
SkillProduction patterns for agent orchestrators, harnesses, evals, MCP/A2A interoperability, memory, provider routing, and LLM usage metering.
githubgemessenes WachstumQuelle öffnen ↗
Installieren
git clone https://github.com/ssheleg/agent-stack ~/.claude/skills/agent-stack2GitHub-Sterne+1 (+100.0 %) - 11Aktiv
historian
AgentA local A2A event and continuity service for agent applications, with structured history, provenance-safe transcripts, literal search, and natural-language queries.
githubgemessenes WachstumQuelle öffnen ↗
1GitHub-Sterne+1 - 12Aktiv
agent-kernel
AgentThe Operating System for Scalable Enterprise AI Agents - Run, orchestrate, and deploy Compliant Enterprise AI Agents at scale across frameworks, without lock-in, rewrites or fragile glue code. Native support for MCP, A2A. Interface with all…
githubgemessenes WachstumQuelle öffnen ↗
146GitHub-Sterne+13 (+9.8 %) - 13Aktiv
mlflow
AgentThe open source AI engineering platform for agents, LLMs, and ML models. MLflow enables teams of all sizes to debug, evaluate, monitor, and optimize production-quality AI applications while controlling costs and managing access to models…
githubgemessenes WachstumQuelle öffnen ↗
27 759GitHub-Sterne+78 (+0.28 %) - 14Aktiv
databuff
AgentDataBuff is an AI-native APM built on Opentelemetry,with multi-agent troubleshooting out of the box.
githubgemessenes WachstumQuelle öffnen ↗
627GitHub-Sterne+22 (+3.6 %) - 15Aktiv
netdata
SonstigeThe fastest path to AI-powered full stack observability, even for lean teams.
githubgemessenes WachstumQuelle öffnen ↗
80 382GitHub-Sterne+80 (+0.10 %) - 16Aktiv
prefect
BibliothekPrefect is a workflow orchestration framework for building resilient data pipelines in Python.
githubgemessenes WachstumQuelle öffnen ↗
23 741GitHub-Sterne+55 (+0.23 %) - 179 079GitHub-Sterne+46 (+0.51 %)
- 18Aktiv
agent-skill-creator
SkillBuild tested agent skills and govern their lifecycle through a user-defined marketplace: evidence, discovery, updates, rollback, quarantine, and 17-platform distribution.
githubgemessenes WachstumQuelle öffnen ↗
Installieren
git clone https://github.com/FrancyJGLisboa/agent-skill-creator ~/.claude/skills/agent-skill-creator2 358GitHub-Sterne+29 (+1.2 %) - 19Aktiv
signoz
SonstigeSigNoz is an open-source, OpenTelemetry-native observability platform for your team and their AI agents. Get logs, metrics, and traces in one tool with features like APM, distributed tracing, log management, infra monitoring, etc. Combined…
githubgemessenes WachstumQuelle öffnen ↗
31 983GitHub-Sterne+48 (+0.15 %) - 20Aktiv
vector
SonstigeA high-performance observability data pipeline.
githubgemessenes WachstumQuelle öffnen ↗
22 497GitHub-Sterne+41 (+0.18 %) - 21Aktiv
A Go framework for building production agent systems with graph workflows, tools, memory, A2A, AG-UI, MCP, evaluation, and observability.
githubgemessenes WachstumQuelle öffnen ↗
1 752GitHub-Sterne+22 (+1.3 %) - 22Aktiv
fable-method
SkillThe Fable Workflow: how Claude Fable 5 worked, distilled into skills any model can run, with the eval that keeps it honest. Think / act / prove.
githubgemessenes WachstumQuelle öffnen ↗
Installieren
git clone https://github.com/Sahir619/fable-method ~/.claude/skills/fable-method2 269GitHub-Sterne+22 (+0.98 %) - 23Aktiv
Amazon Bedrock AgentCore enterprise platform accelerator with AWS CDK, Terraform organization guardrails, MCP/A2A agents, security, memory, and observability.
githubgemessenes WachstumQuelle öffnen ↗
5GitHub-Sterne+1 (+25.0 %) - 24Ruhend
kiboserve
MCPOpen-source Python framework to deploy AI agents via HTTP, A2A, and MCP with built-in observability
githubgemessenes WachstumQuelle öffnen ↗
5GitHub-Sterne+1 (+25.0 %) - 25Aktiv
openstatus
MCP🫖 Status page with uptime monitoring & API monitoring as code 🫖
githubgemessenes WachstumQuelle öffnen ↗
9 045GitHub-Sterne+28 (+0.31 %)