Die Werkzeugbeschreibungen sind auf Englisch.

LLM-Beobachtbarkeit

Monitoring, Traces, Bewertung und Qualität von LLM-Anwendungen.

Anwendungsfall

Aktivität

Sortieren nach

121

Werkzeug-Rangliste

Nach gemessenem Wachstum und quellenübergreifend normalisiert sortiert. Rohwerte behalten ihr eigenes Zeitfenster; eine Änderung erfordert mindestens zwei Messungen in 7 Tagen.

  1. 1

    Skill
    Aktiv

    Production-grade AI coding rules for Cursor and Claude Code. 15 rules + 9 doc templates + skills + agents + MCP setup. Drop into any project.

    githubgemessenes WachstumQuelle öffnen ↗

    Installieren git clone https://github.com/aiagentwithdhruv/ai-dev-stack ~/.claude/skills/ai-dev-stack

    10GitHub-Sternestabil
  2. 2
    Aktiv

    Claude Code plugin marketplace — agentic-engineering (spec-driven shape→decide→execute→measure→eval with adversarial review) + github-keeper (audit/elevate READMEs and make a repo open-source-ready).

    githubgemessenes WachstumQuelle öffnen ↗

    Installieren /plugin marketplace add GiustoPiedimonte/agentic-engineering-marketplace

    13GitHub-Sternestabil
  3. 3

    MCP
    Aktiv

    A typed, replayable, budget-aware programming language for LLM agents. Compiles to Python & TypeScript.

    githubgemessenes WachstumQuelle öffnen ↗

    2GitHub-Sternestabil
  4. 4

    Skill
    Aktiv

    Marketing Mix Model expert agent plugin for Claude Code - pymc-marketing v0.18.2+

    githubgemessenes WachstumQuelle öffnen ↗

    Installieren /plugin marketplace add Yakoub-ai/agent-mmm

    4GitHub-Sternestabil
  5. 5

    Skill
    Aktiv

    Measure prompt and skill improvements with blind A/B comparison.

    githubgemessenes WachstumQuelle öffnen ↗

    Installieren /plugin marketplace add shinpr/rashomon

    18GitHub-Sternestabil
  6. 6
    Aktiv

    AgentGateway — independent third-party profile of a public API surface, by API Evangelist. AgentGateway is an open-source, AI-native proxy and gateway for routing, observing, and governing traffic to and from AI agents, LLM providers, and…

    githubgemessenes WachstumQuelle öffnen ↗

    0GitHub-Sternestabil
  7. 7

    Agent
    Aktiv

    Local-first control plane for durable, governed agent teams: DAG orchestration, approvals, memory, signed skills, trace contracts, and realtime.

    githubgemessenes WachstumQuelle öffnen ↗

    0GitHub-Sternestabil
  8. 8
    Aktiv

    Production reliability, benchmarking, and evaluation layer for MCP agent runtimes — protocol negotiation, A2A interop, circuit breakers, connection pooling, adaptive streaming, and autoscaling, with measured throughput/latency and a…

    githubgemessenes WachstumQuelle öffnen ↗

    3GitHub-Sternestabil
  9. 9

    Agent
    Aktiv

    Generic DeepSeek Harness (dsh) plugins: A2A protocol server, session storage mirror, and Langfuse observability.

    githubgemessenes WachstumQuelle öffnen ↗

    1GitHub-Sternestabil
  10. 10

    Skill
    Ruhend

    Don't make LLMs honest. Make every factual claim auditable. — An LLM Claim Auditing Layer with T1-T7 truth gradients. 98.1% business effectiveness on LiarBench v0.2.

    githubgemessenes WachstumQuelle öffnen ↗

    Installieren git clone https://github.com/lc198707/anti-lie ~/.claude/skills/anti-lie

    89GitHub-Sternestabil
  11. 11
    Aktiv

    Audit which Claude Code skills you actually use — surface dead installs and hallucinated invocations from your session logs.

    githubgemessenes WachstumQuelle öffnen ↗

    Installieren /plugin marketplace add sfrangulov/skill-graveyard

    10GitHub-Sternestabil
  12. 12
    Aktiv

    Mide si tu agente de IA cumple las reglas que le escribiste. Lee el historial local de Claude Code y devuelve un porcentaje por regla. Sin instalar nada, sin red, biblioteca estándar.

    githubgemessenes WachstumQuelle öffnen ↗

    Installieren git clone https://github.com/jleonceo/adherencia-reglas ~/.claude/skills/adherencia-reglas

    1GitHub-Sternestabil
  13. 13

    Sonstige
    Aktiv

    Local-first Agentic AI Infrastructure Platform with LLM Gateway, Agent DAG Runtime, MCP Tool Hub, A2A Agent Mesh, Local RAG, Tool Sandbox and Observability.

    githubgemessenes WachstumQuelle öffnen ↗

    1GitHub-Sternestabil
  14. 14

    Agent
    Ruhend

    The reliability and interoperability layer for AI agents. A2A-native. MCP-bridged.

    githubgemessenes WachstumQuelle öffnen ↗

    1GitHub-Sternestabil
  15. 15
    Ruhend

    An AI-powered multi-agent system that demonstrates clinical triage, OTC medication recommendations, and e-pharmacy integration for respiratory conditions. Built with modular agents that collaborate to provide safe, intelligent healthcare…

    githubgemessenes WachstumQuelle öffnen ↗

    0GitHub-Sternestabil
  16. 16
    Ruhend

    Local open-source dev tool to debug, secure, and evaluate LLM agents. Provides static analysis, dynamic security checks, and runtime monitoring - integrates with Cursor and Claude Code.

    githubgemessenes WachstumQuelle öffnen ↗

    Installieren /plugin marketplace add cylestio/agent-inspector

    9GitHub-Sternestabil
  17. 17
    Ruhend

    MCP as a Judge: a behavioral MCP that strengthens AI coding assistants via explicit LLM evaluations

    mcpgemessenes WachstumQuelle öffnen ↗

    Installieren claude mcp add mcp-as-a-judge -- uvx mcp-as-a-judge

    17GitHub-Sternestabil
  18. 18
    Aktiv

    8 MCP SMB products — standalone AI servers for SMBs: Guardrails, FinOps, Observability, Router, Trust Score, Memory, ThinkSecure, A2A Lite. 37 tools, TNC credits billing.

    githubgemessenes WachstumQuelle öffnen ↗

    1GitHub-Sternestabil
  19. 19
    Aktiv

    Claude Code skills where every entry ships receipts — accuracy-gated benchmarks against baseline and placebo, rejects published

    githubgemessenes WachstumQuelle öffnen ↗

    Installieren git clone https://github.com/sjh9714/skill-receipts ~/.claude/skills/skill-receipts

    2GitHub-Sternestabil
  20. 20

    Sonstige
    Ruhend

    A Claude Code skill that adds a rubric-based eval layer to any agent project. Framework-agnostic — generates rubric, test cases, judge prompt, and harness. Returns a weighted score plus a judge-leniency signal.

    githubgemessenes WachstumQuelle öffnen ↗

    13GitHub-Sternestabil
  21. 21
    Ruhend

    A safety-first multi-agent mental health companion with real-time distress tracking, triple-layer guardrails, and evidence-based grounding techniques. Built for Kaggle × Google Agents Intensive 2025 Capstone (Agents for Good Track)

    githubgemessenes WachstumQuelle öffnen ↗

    1GitHub-Sternestabil
  22. 22

    Skill
    Aktiv

    Anchor — the production-grade AGENTS.md template for AI coding agents. 51 sections of battle-tested rules. Keep your agents grounded.

    githubgemessenes WachstumQuelle öffnen ↗

    Installieren git clone https://github.com/peva3/anchor ~/.claude/skills/anchor

    8GitHub-Sternestabil
  23. 23
    Ruhend

    Two Claude Code skills for adversarial code review (single-model PAR and multi-model MMAR with cross-critique to catch hallucinations), plus a fixture-based eval suite.

    githubgemessenes WachstumQuelle öffnen ↗

    Installieren /plugin marketplace add prime-radiant-inc/parallel-adversarial-review

    17GitHub-Sternestabil
  24. 24

    Agent
    Ruhend

    AgentOps: Multi-agent infrastructure remediation platform. A2A protocol, agent coordination, HITL approval, auto-rollback.

    githubgemessenes WachstumQuelle öffnen ↗

    0GitHub-Sternestabil
  25. 25

    Skill
    Aktiv

    Production patterns for agent orchestrators, harnesses, evals, MCP/A2A interoperability, memory, provider routing, and LLM usage metering.

    githubgemessenes WachstumQuelle öffnen ↗

    Installieren git clone https://github.com/ssheleg/agent-stack ~/.claude/skills/agent-stack

    2GitHub-Sternestabil