Die Werkzeugbeschreibungen sind auf Englisch.

LLM-Beobachtbarkeit

Monitoring, Traces, Bewertung und Qualität von LLM-Anwendungen.

Anwendungsfall

Aktivität

Sortieren nach

122

Werkzeug-Rangliste

Ranked by creation date, newest first; undated entries come last. Rohwerte behalten ihr eigenes Zeitfenster; eine Änderung erfordert mindestens zwei Messungen in 7 Tagen.

  1. 1

    Skill
    Ruhend

    An assembly line for AI software development. 35 skills, 11 agent personas, 29 commands. From raw idea to shipped code.

    githubgemessenes WachstumQuelle öffnen ↗

    Installieren /plugin marketplace add aneja5/forge-skills

    3GitHub-Sternestabil
  2. 2

    MCP
    Aktiv

    Open-source, self-hosted control plane for OpenClaw and Hermes AI-agent fleets on Docker/Kubernetes — REST, CLI, and MCP.

    githubgemessenes WachstumQuelle öffnen ↗

    50GitHub-Sternestabil
  3. 3
    Ruhend

    A self-improving harness router for Claude Code.

    githubgemessenes WachstumQuelle öffnen ↗

    Installieren /plugin marketplace add SeongwoongCho/adaptive-harness

    8GitHub-Sternestabil
  4. 4
    Ruhend

    OpenTelemetry semantic conventions and instrumentation for agent provenance, derivation lineage, and acceptance criteria evaluation. Fills the Microsoft AI stack observability gap.

    githubgemessenes WachstumQuelle öffnen ↗

    0GitHub-Sternestabil
  5. 5

    Skill
    Aktiv

    Production-grade AI coding rules for Cursor and Claude Code. 15 rules + 9 doc templates + skills + agents + MCP setup. Drop into any project.

    githubgemessenes WachstumQuelle öffnen ↗

    Installieren git clone https://github.com/aiagentwithdhruv/ai-dev-stack ~/.claude/skills/ai-dev-stack

    10GitHub-Sternestabil
  6. 6
    Aktiv

    Vendor-neutral OpenTelemetry skills for AI coding agents, grounded in upstream sources

    githubgemessenes WachstumQuelle öffnen ↗

    Installieren git clone https://github.com/ollygarden/opentelemetry-agent-skills ~/.claude/skills/opentelemetry-agent-skills

    97GitHub-Sternestabil
  7. 7

    MCP
    Aktiv

    AURA is a production-tested SRE agent platform you can deploy in minutes. AURA handles the guardrails, APIs, state management, streaming, and failure handling required to put AI to work safely on production infrastructure.

    githubgemessenes WachstumQuelle öffnen ↗

    348GitHub-Sterne+94 (+37.0 %)
  8. 8

    MCP
    Ruhend

    Open-source Python framework to deploy AI agents via HTTP, A2A, and MCP with built-in observability

    githubgemessenes WachstumQuelle öffnen ↗

    5GitHub-Sternestabil
  9. 9

    Agent
    Ruhend

    AgentOps: Multi-agent infrastructure remediation platform. A2A protocol, agent coordination, HITL approval, auto-rollback.

    githubgemessenes WachstumQuelle öffnen ↗

    0GitHub-Sternestabil
  10. 10

    Skill
    Aktiv

    local-first analytics for AI agent skills

    githubgemessenes WachstumQuelle öffnen ↗

    Installieren git clone https://github.com/crafter-station/skill-kit ~/.claude/skills/skill-kit

    78GitHub-Sterne+1 (+1.3 %)
  11. 11

    Sonstige
    Ruhend

    Official Python SDK for GT8004 — AI agent observability with MCP, A2A, x402 payment tracking. FastAPI, Flask, FastMCP middleware included.

    githubgemessenes WachstumQuelle öffnen ↗

    1GitHub-Sternestabil
  12. 12
    Aktiv

    Log Claude Code sessions to Opik, the open-source LLM observability and evaluation platform, built by Comet. Tracing, evaluation, and skills for observable AI applications.

    githubgemessenes WachstumQuelle öffnen ↗

    22GitHub-Sternestabil
  13. 13

    Agent
    Ruhend

    Policy-enforced observability and fail-closed guardrails for MCP/A2A multi-agent systems.

    githubgemessenes WachstumQuelle öffnen ↗

    26GitHub-Sternestabil
  14. 14
    Ruhend

    A Python proof-of-concept for tracing multi-turn Agent-to-Agent (A2A) conversations as a single unified MLflow trace for LLM observability and evaluation.

    githubgemessenes WachstumQuelle öffnen ↗

    0GitHub-Sternestabil
  15. 15

    Skill
    Aktiv

    Measure prompt and skill improvements with blind A/B comparison.

    githubgemessenes WachstumQuelle öffnen ↗

    Installieren /plugin marketplace add shinpr/rashomon

    18GitHub-Sternestabil
  16. 16
    Aktiv

    Dashboard for monitoring claude code sessions.

    githubgemessenes WachstumQuelle öffnen ↗

    Installieren /plugin marketplace add JayantDevkar/claude-code-karma

    325GitHub-Sterne+2 (+0.62 %)
  17. 17

    Skill
    Aktiv

    A skill creator that proves its skills work. Evidence-driven skill creation for Claude Code and Codex: baseline-tested generation, per-skill regression evals, ecosystem doctor, cross-runtime compile, and an opt-in proactive advisor.

    githubgemessenes WachstumQuelle öffnen ↗

    Installieren git clone https://github.com/tripleyak/SkillForge ~/.claude/skills/SkillForge

    887GitHub-Sterne+3 (+0.34 %)
  18. 18
    Ruhend

    🔍 AI observability skill for Claude Code. Debug LangChain/LangGraph agents by fetching execution traces from LangSmith Studio directly in your terminal.

    githubgemessenes WachstumQuelle öffnen ↗

    Installieren git clone https://github.com/OthmanAdi/langsmith-fetch-skill ~/.claude/skills/langsmith-fetch-skill

    27GitHub-Sternestabil
  19. 19
    Aktiv

    Self improving agents through iterations

    githubgemessenes WachstumQuelle öffnen ↗

    105GitHub-Sternestabil
  20. 20
    Ruhend

    An AI-powered multi-agent system that demonstrates clinical triage, OTC medication recommendations, and e-pharmacy integration for respiratory conditions. Built with modular agents that collaborate to provide safe, intelligent healthcare…

    githubgemessenes WachstumQuelle öffnen ↗

    0GitHub-Sternestabil
  21. 21
    Ruhend

    A safety-first multi-agent mental health companion with real-time distress tracking, triple-layer guardrails, and evidence-based grounding techniques. Built for Kaggle × Google Agents Intensive 2025 Capstone (Agents for Good Track)

    githubgemessenes WachstumQuelle öffnen ↗

    1GitHub-Sternestabil
  22. 22

    Sonstige
    Ruhend

    Multi-agent customer support system with Google ADK & Gemini 2.5 Flash Lite. Kaggle capstone demonstrating 11+ concepts. Automates 80%+ queries, <10s response time.

    githubgemessenes WachstumQuelle öffnen ↗

    2GitHub-Sternestabil
  23. 23
    Ruhend

    Local open-source dev tool to debug, secure, and evaluate LLM agents. Provides static analysis, dynamic security checks, and runtime monitoring - integrates with Cursor and Claude Code.

    githubgemessenes WachstumQuelle öffnen ↗

    Installieren /plugin marketplace add cylestio/agent-inspector

    9GitHub-Sternestabil
  24. 24
    Aktiv

    Build tested agent skills and govern their lifecycle through a user-defined marketplace: evidence, discovery, updates, rollback, quarantine, and 17-platform distribution.

    githubgemessenes WachstumQuelle öffnen ↗

    Installieren git clone https://github.com/FrancyJGLisboa/agent-skills-platform ~/.claude/skills/agent-skills-platform

    2 379GitHub-Sterne+3 (+0.13 %)
  25. 25
    Aktiv

    Build tested agent skills and govern their lifecycle through a user-defined marketplace: evidence, discovery, updates, rollback, quarantine, and 17-platform distribution.

    githubgeschätztes MomentumQuelle öffnen ↗

    Installieren git clone https://github.com/FrancyJGLisboa/agent-skill-creator ~/.claude/skills/agent-skill-creator

    2 367GitHub-Sterne·