Die Werkzeugbeschreibungen sind auf Englisch.
LLM-Beobachtbarkeit
Monitoring, Traces, Bewertung und Qualität von LLM-Anwendungen.
Anwendungsfall
Aktivität
Sortieren nach
122
Werkzeug-Rangliste
Nach gemessenem Wachstum und quellenübergreifend normalisiert sortiert. Rohwerte behalten ihr eigenes Zeitfenster; eine Änderung erfordert mindestens zwei Messungen in 7 Tagen.
- 1Aktiv
bakeoff
SkillTurn one decision into a judged tournament of solutions, then pick the best — a Claude Code skill that generates candidates, auto-derives the rubric, judges independently, and returns a defensible winner.
githubgemessenes WachstumQuelle öffnen ↗
Installieren
/plugin marketplace add CoriChui/bakeoff10GitHub-Sternestabil - 2Aktiv
nora
MCPOpen-source, self-hosted control plane for OpenClaw and Hermes AI-agent fleets on Docker/Kubernetes — REST, CLI, and MCP.
githubgemessenes WachstumQuelle öffnen ↗
50GitHub-Sterne+1 (+2.0 %) - 3Aktiv
Skill instalable para Claude Code que mide si tu agente cumple las reglas establecidas. Lee el historial de sesiones de tu disco y devuelve un porcentaje por regla de cumplimiento. Sin instalar nada, sin red, biblioteca estándar.
githubgemessenes WachstumQuelle öffnen ↗
Installieren
/plugin marketplace add jleonceo/skill-adherencia-reglas1GitHub-Sternestabil - 4Ruhend
forge-skills
SkillAn assembly line for AI software development. 35 skills, 11 agent personas, 29 commands. From raw idea to shipped code.
githubgemessenes WachstumQuelle öffnen ↗
Installieren
/plugin marketplace add aneja5/forge-skills3GitHub-Sternestabil - 5Aktiv
skill-graveyard
SkillAudit which Claude Code skills you actually use — surface dead installs and hallucinated invocations from your session logs.
githubgemessenes WachstumQuelle öffnen ↗
Installieren
/plugin marketplace add sfrangulov/skill-graveyard10GitHub-Sternestabil - 6Ruhend
A Python proof-of-concept for tracing multi-turn Agent-to-Agent (A2A) conversations as a single unified MLflow trace for LLM observability and evaluation.
githubgemessenes WachstumQuelle öffnen ↗
0GitHub-Sternestabil - 7Aktiv
ai-dev-stack
SkillProduction-grade AI coding rules for Cursor and Claude Code. 15 rules + 9 doc templates + skills + agents + MCP setup. Drop into any project.
githubgemessenes WachstumQuelle öffnen ↗
Installieren
git clone https://github.com/aiagentwithdhruv/ai-dev-stack ~/.claude/skills/ai-dev-stack10GitHub-Sternestabil - 8Aktiv
DriftSentinel
AgentAgent 降智检测与自愈公评网络 — an immune system for the AI agent society
githubgemessenes WachstumQuelle öffnen ↗
1GitHub-Sternestabil - 9Aktiv
galdor
SonstigeA Go-native framework for LLM agents, with OpenTelemetry observability built in.
githubgemessenes WachstumQuelle öffnen ↗
10GitHub-Sternestabil - 10Ruhend
kiboserve
MCPOpen-source Python framework to deploy AI agents via HTTP, A2A, and MCP with built-in observability
githubgemessenes WachstumQuelle öffnen ↗
5GitHub-Sternestabil - 11Aktiv
boundary-bench
AgentDeterministic crash tests for agent control planes: paired worlds, hard stops, budgets, scopes, approvals, and hash-linked receipts.
githubgemessenes WachstumQuelle öffnen ↗
1GitHub-Sternestabil - 12Aktiv
dsh-plugins
AgentGeneric DeepSeek Harness (dsh) plugins: A2A protocol server, session storage mirror, and Langfuse observability.
githubgemessenes WachstumQuelle öffnen ↗
1GitHub-Sternestabil - 13Aktiv
Multi-agent quality gate skill for Claude Code that researches, reviews, tests and challenges AI-generated work before the final answer.
githubgemessenes WachstumQuelle öffnen ↗
Installieren
git clone https://github.com/ma-nucho-pro/supervisor-skill-claude ~/.claude/skills/supervisor-skill-claude2GitHub-Sternestabil - 14Aktiv
historian
AgentA local A2A event and continuity service for agent applications, with structured history, provenance-safe transcripts, literal search, and natural-language queries.
githubgemessenes WachstumQuelle öffnen ↗
1GitHub-Sternestabil - 15Aktiv
DeepSeek-Infra
SonstigeLocal-first Agentic AI Infrastructure Platform with LLM Gateway, Agent DAG Runtime, MCP Tool Hub, A2A Agent Mesh, Local RAG, Tool Sandbox and Observability.
githubgemessenes WachstumQuelle öffnen ↗
1GitHub-Sternestabil - 16Aktiv
sre-on-call
AgentMulti-agent SRE on-call investigator that auto-triages Slack/Discord infrastructure alerts via AWS Bedrock AgentCore, fanning out to specialized agents (CloudWatch, EKS, Slack/Discord scanners) for parallel investigation.
githubgemessenes WachstumQuelle öffnen ↗
3GitHub-Sternestabil - 17Aktiv
Query Spanlens LLM observability from Cursor, Claude Desktop, or Continue via MCP.
mcpgemessenes WachstumQuelle öffnen ↗
Installieren
claude mcp add mcp-server -- npx @spanlens/mcp-server12GitHub-Sternestabil - 18Aktiv
AI Guardian
MCPGoverned local-LLM observability: model policy, prompt scanner, capture proxy, 21 tools.
mcpgemessenes WachstumQuelle öffnen ↗
Installieren
claude mcp add ai-guardian -- uvx ai-guardian-aiops0GitHub-Sternestabil - 19Aktiv
pyxen
AgentA lightweight Python library that decouples agentic runtime from applications it builds
githubgemessenes WachstumQuelle öffnen ↗
1GitHub-Sternestabil - 20Aktiv
Amazon Bedrock AgentCore enterprise platform accelerator with AWS CDK, Terraform organization guardrails, MCP/A2A agents, security, memory, and observability.
githubgemessenes WachstumQuelle öffnen ↗
6GitHub-Sterne+1 (+20.0 %) - 21Aktiv
Claude Code plugin marketplace — agentic-engineering (spec-driven shape→decide→execute→measure→eval with adversarial review) + github-keeper (audit/elevate READMEs and make a repo open-source-ready).
githubgemessenes WachstumQuelle öffnen ↗
Installieren
/plugin marketplace add GiustoPiedimonte/agentic-engineering-marketplace13GitHub-Sternestabil - 22Aktiv
rashomon
SkillMeasure prompt and skill improvements with blind A/B comparison.
githubgemessenes WachstumQuelle öffnen ↗
Installieren
/plugin marketplace add shinpr/rashomon18GitHub-Sternestabil - 23Aktiv
shokunin-review
SonstigeTerminal-first validation harness for reviewing PRDs, RFCs, strategy docs, and experiment plans.
githubgemessenes WachstumQuelle öffnen ↗
1GitHub-Sternestabil - 24Ruhend
primal-core
AgentThe reliability and interoperability layer for AI agents. A2A-native. MCP-bridged.
githubgemessenes WachstumQuelle öffnen ↗
1GitHub-Sternestabil - 25Aktiv
nudge
MCPA typed, replayable, budget-aware programming language for LLM agents. Compiles to Python & TypeScript.
githubgemessenes WachstumQuelle öffnen ↗
2GitHub-Sternestabil