Die Werkzeugbeschreibungen sind auf Englisch.
LLM-Beobachtbarkeit
Monitoring, Traces, Bewertung und Qualität von LLM-Anwendungen.
Anwendungsfall
Aktivität
Sortieren nach
121
Werkzeug-Rangliste
Nach gemessenem Wachstum und quellenübergreifend normalisiert sortiert. Rohwerte behalten ihr eigenes Zeitfenster; eine Änderung erfordert mindestens zwei Messungen in 7 Tagen.
- 1Aktiv
ai-dev-stack
SkillProduction-grade AI coding rules for Cursor and Claude Code. 15 rules + 9 doc templates + skills + agents + MCP setup. Drop into any project.
githubgemessenes WachstumQuelle öffnen ↗
Installieren
git clone https://github.com/aiagentwithdhruv/ai-dev-stack ~/.claude/skills/ai-dev-stack10GitHub-Sternestabil - 2Aktiv
Claude Code plugin marketplace — agentic-engineering (spec-driven shape→decide→execute→measure→eval with adversarial review) + github-keeper (audit/elevate READMEs and make a repo open-source-ready).
githubgemessenes WachstumQuelle öffnen ↗
Installieren
/plugin marketplace add GiustoPiedimonte/agentic-engineering-marketplace13GitHub-Sternestabil - 3Aktiv
nudge
MCPA typed, replayable, budget-aware programming language for LLM agents. Compiles to Python & TypeScript.
githubgemessenes WachstumQuelle öffnen ↗
2GitHub-Sternestabil - 4Aktiv
agent-mmm
SkillMarketing Mix Model expert agent plugin for Claude Code - pymc-marketing v0.18.2+
githubgemessenes WachstumQuelle öffnen ↗
Installieren
/plugin marketplace add Yakoub-ai/agent-mmm4GitHub-Sternestabil - 5Aktiv
rashomon
SkillMeasure prompt and skill improvements with blind A/B comparison.
githubgemessenes WachstumQuelle öffnen ↗
Installieren
/plugin marketplace add shinpr/rashomon18GitHub-Sternestabil - 6Aktiv
agentgateway
MCPAgentGateway — independent third-party profile of a public API surface, by API Evangelist. AgentGateway is an open-source, AI-native proxy and gateway for routing, observing, and governing traffic to and from AI agents, LLM providers, and…
githubgemessenes WachstumQuelle öffnen ↗
0GitHub-Sternestabil - 7Aktiv
evoagent-os
AgentLocal-first control plane for durable, governed agent teams: DAG orchestration, approvals, memory, signed skills, trace contracts, and realtime.
githubgemessenes WachstumQuelle öffnen ↗
0GitHub-Sternestabil - 8Aktiv
Production reliability, benchmarking, and evaluation layer for MCP agent runtimes — protocol negotiation, A2A interop, circuit breakers, connection pooling, adaptive streaming, and autoscaling, with measured throughput/latency and a…
githubgemessenes WachstumQuelle öffnen ↗
3GitHub-Sternestabil - 9Aktiv
dsh-plugins
AgentGeneric DeepSeek Harness (dsh) plugins: A2A protocol server, session storage mirror, and Langfuse observability.
githubgemessenes WachstumQuelle öffnen ↗
1GitHub-Sternestabil - 10Ruhend
anti-lie
SkillDon't make LLMs honest. Make every factual claim auditable. — An LLM Claim Auditing Layer with T1-T7 truth gradients. 98.1% business effectiveness on LiarBench v0.2.
githubgemessenes WachstumQuelle öffnen ↗
Installieren
git clone https://github.com/lc198707/anti-lie ~/.claude/skills/anti-lie89GitHub-Sternestabil - 11Aktiv
skill-graveyard
SkillAudit which Claude Code skills you actually use — surface dead installs and hallucinated invocations from your session logs.
githubgemessenes WachstumQuelle öffnen ↗
Installieren
/plugin marketplace add sfrangulov/skill-graveyard10GitHub-Sternestabil - 12Aktiv
adherencia-reglas
SkillMide si tu agente de IA cumple las reglas que le escribiste. Lee el historial local de Claude Code y devuelve un porcentaje por regla. Sin instalar nada, sin red, biblioteca estándar.
githubgemessenes WachstumQuelle öffnen ↗
Installieren
git clone https://github.com/jleonceo/adherencia-reglas ~/.claude/skills/adherencia-reglas1GitHub-Sternestabil - 13Aktiv
DeepSeek-Infra
SonstigeLocal-first Agentic AI Infrastructure Platform with LLM Gateway, Agent DAG Runtime, MCP Tool Hub, A2A Agent Mesh, Local RAG, Tool Sandbox and Observability.
githubgemessenes WachstumQuelle öffnen ↗
1GitHub-Sternestabil - 14Ruhend
primal-core
AgentThe reliability and interoperability layer for AI agents. A2A-native. MCP-bridged.
githubgemessenes WachstumQuelle öffnen ↗
1GitHub-Sternestabil - 15Ruhend
An AI-powered multi-agent system that demonstrates clinical triage, OTC medication recommendations, and e-pharmacy integration for respiratory conditions. Built with modular agents that collaborate to provide safe, intelligent healthcare…
githubgemessenes WachstumQuelle öffnen ↗
0GitHub-Sternestabil - 16Ruhend
agent-inspector
SkillLocal open-source dev tool to debug, secure, and evaluate LLM agents. Provides static analysis, dynamic security checks, and runtime monitoring - integrates with Cursor and Claude Code.
githubgemessenes WachstumQuelle öffnen ↗
Installieren
/plugin marketplace add cylestio/agent-inspector9GitHub-Sternestabil - 17Ruhend
MCP as a Judge: a behavioral MCP that strengthens AI coding assistants via explicit LLM evaluations
mcpgemessenes WachstumQuelle öffnen ↗
Installieren
claude mcp add mcp-as-a-judge -- uvx mcp-as-a-judge17GitHub-Sternestabil - 18Aktiv
8 MCP SMB products — standalone AI servers for SMBs: Guardrails, FinOps, Observability, Router, Trust Score, Memory, ThinkSecure, A2A Lite. 37 tools, TNC credits billing.
githubgemessenes WachstumQuelle öffnen ↗
1GitHub-Sternestabil - 19Aktiv
skill-receipts
SkillClaude Code skills where every entry ships receipts — accuracy-gated benchmarks against baseline and placebo, rejects published
githubgemessenes WachstumQuelle öffnen ↗
Installieren
git clone https://github.com/sjh9714/skill-receipts ~/.claude/skills/skill-receipts2GitHub-Sternestabil - 20Ruhend
eval-layer
SonstigeA Claude Code skill that adds a rubric-based eval layer to any agent project. Framework-agnostic — generates rubric, test cases, judge prompt, and harness. Returns a weighted score plus a judge-leniency signal.
githubgemessenes WachstumQuelle öffnen ↗
13GitHub-Sternestabil - 21Ruhend
A safety-first multi-agent mental health companion with real-time distress tracking, triple-layer guardrails, and evidence-based grounding techniques. Built for Kaggle × Google Agents Intensive 2025 Capstone (Agents for Good Track)
githubgemessenes WachstumQuelle öffnen ↗
1GitHub-Sternestabil - 22Aktiv
anchor
SkillAnchor — the production-grade AGENTS.md template for AI coding agents. 51 sections of battle-tested rules. Keep your agents grounded.
githubgemessenes WachstumQuelle öffnen ↗
Installieren
git clone https://github.com/peva3/anchor ~/.claude/skills/anchor8GitHub-Sternestabil - 23Ruhend
Two Claude Code skills for adversarial code review (single-model PAR and multi-model MMAR with cross-critique to catch hallucinations), plus a fixture-based eval suite.
githubgemessenes WachstumQuelle öffnen ↗
Installieren
/plugin marketplace add prime-radiant-inc/parallel-adversarial-review17GitHub-Sternestabil - 24Ruhend
agentops
AgentAgentOps: Multi-agent infrastructure remediation platform. A2A protocol, agent coordination, HITL approval, auto-rollback.
githubgemessenes WachstumQuelle öffnen ↗
0GitHub-Sternestabil - 25Aktiv
agent-stack
SkillProduction patterns for agent orchestrators, harnesses, evals, MCP/A2A interoperability, memory, provider routing, and LLM usage metering.
githubgemessenes WachstumQuelle öffnen ↗
Installieren
git clone https://github.com/ssheleg/agent-stack ~/.claude/skills/agent-stack2GitHub-Sternestabil