Die Werkzeugbeschreibungen sind auf Englisch.

LLM-Beobachtbarkeit

Monitoring, Traces, Bewertung und Qualität von LLM-Anwendungen.

Anwendungsfall

Aktivität

Sortieren nach

122

Werkzeug-Rangliste

Gemischte Rangfolge: Gemessenes Wachstum hat Vorrang. Rohwerte behalten ihr eigenes Zeitfenster; eine Änderung erfordert mindestens zwei Messungen in 7 Tagen.

  1. 1

    Sonstige
    Ruhend

    A Claude Code skill that adds a rubric-based eval layer to any agent project. Framework-agnostic — generates rubric, test cases, judge prompt, and harness. Returns a weighted score plus a judge-leniency signal.

    githubgemessenes WachstumQuelle öffnen ↗

    13GitHub-Sternestabil
  2. 2

    Agent
    Aktiv

    Local-first control plane for durable, governed agent teams: DAG orchestration, approvals, memory, signed skills, trace contracts, and realtime.

    githubgemessenes WachstumQuelle öffnen ↗

    0GitHub-Sternestabil
  3. 3

    Agent
    Aktiv

    MemroOS / memroos: memory OS and governance layer for AI agents, agent workflows, dispatch, proof, and context continuity.

    githubgemessenes WachstumQuelle öffnen ↗

    7GitHub-Sternestabil
  4. 4

    Sonstige
    Aktiv

    🔑 HomeStream · 家园·流 — 零成本自托管多Agent协作框架,通往AI世界的那把钥匙 | Zero-cost self-hosted multi-agent framework — The key to AI world

    githubgemessenes WachstumQuelle öffnen ↗

    0GitHub-Sternestabil
  5. 5

    Sonstige
    Aktiv

    A command-line tool to scaffold and manage enterprise-ready AI Agents powered by the A2A (Agent-to-Agent) protocol

    githubgemessenes WachstumQuelle öffnen ↗

    14GitHub-Sterne+1 (+7.7 %)
  6. 6
    Aktiv

    Make Claude write clearly, for everyone.

    githubgemessenes WachstumQuelle öffnen ↗

    Installieren /plugin marketplace add stefanobaghino/simple-output-styles

    16GitHub-Sternestabil
  7. 7

    Skill
    Aktiv

    Marketing Mix Model expert agent plugin for Claude Code - pymc-marketing v0.18.2+

    githubgemessenes WachstumQuelle öffnen ↗

    Installieren /plugin marketplace add Yakoub-ai/agent-mmm

    4GitHub-Sternestabil
  8. 8

    Agent
    Aktiv

    The agent framework where the model never holds the trigger — every consequential action clears your policy first, waits for a human when it matters, and lands on a record you can verify. Build on it, or put it around the agent you already…

    githubgemessenes WachstumQuelle öffnen ↗

    1GitHub-Sternestabil
  9. 9

    Agent
    Ruhend

    A Multi-Agent System (MAS) evaluation framework using PydanticAI that generates and evaluates scientific paper reviews through a three-tiered assessment approach: traditional metrics, LLM-as-a-Judge, and graph-based complexity analysis.

    githubgemessenes WachstumQuelle öffnen ↗

    2GitHub-Sternestabil
  10. 10
    Ruhend

    MCP as a Judge: a behavioral MCP that strengthens AI coding assistants via explicit LLM evaluations

    mcpgemessenes WachstumQuelle öffnen ↗

    Installieren claude mcp add mcp-as-a-judge -- uvx mcp-as-a-judge

    17GitHub-Sternestabil
  11. 11
    Ruhend

    An AI-powered multi-agent system that demonstrates clinical triage, OTC medication recommendations, and e-pharmacy integration for respiratory conditions. Built with modular agents that collaborate to provide safe, intelligent healthcare…

    githubgemessenes WachstumQuelle öffnen ↗

    0GitHub-Sternestabil
  12. 12

    Sonstige
    Ruhend

    Multi-agent customer support system with Google ADK & Gemini 2.5 Flash Lite. Kaggle capstone demonstrating 11+ concepts. Automates 80%+ queries, <10s response time.

    githubgemessenes WachstumQuelle öffnen ↗

    2GitHub-Sternestabil
  13. 13
    Aktiv

    Architecture-first Python scaffold for an auditable multi-agent corporate credit desk using A2A, MCP, deterministic credit policies, model routing, and OpenTelemetry.

    githubgemessenes WachstumQuelle öffnen ↗

    0GitHub-Sternestabil
  14. 14
    Ruhend

    Two Claude Code skills for adversarial code review (single-model PAR and multi-model MMAR with cross-critique to catch hallucinations), plus a fixture-based eval suite.

    githubgemessenes WachstumQuelle öffnen ↗

    Installieren /plugin marketplace add prime-radiant-inc/parallel-adversarial-review

    17GitHub-Sternestabil
  15. 15
    Aktiv

    Mini Program Engineering Skill Suite is an Agent Skill suite for evidence‑first mini‑program development. It helps agents bring WeChat or other mini‑program projects from vague intent to reliable engineering work: project intake, product…

    githubgemessenes WachstumQuelle öffnen ↗

    Installieren git clone https://github.com/NocodeMrLi/mini-program-engineering-skill-suite ~/.claude/skills/mini-program-engineering-skill-suite

    45GitHub-Sterne+6 (+15.4 %)
  16. 16
    Aktiv

    Run many AIs on one board and keep control of all of it. Deterministic code decides who acts — never a model. A privacy floor keeps sensitive work on your machine, your own tests decide what counts as done, and every action lands on a…

    githubgemessenes WachstumQuelle öffnen ↗

    Installieren git clone https://github.com/sandhusukhdeep2/sc-prism-releases ~/.claude/skills/sc-prism-releases

    1GitHub-Sternestabil
  17. 17

    Sonstige
    Aktiv

    Open-source observability tool that uses AI agents to self-heal your software

    githubgeschätztes MomentumQuelle öffnen ↗

    1 404GitHub-Sterne

Lern- und Referenzressourcen

Nach gemessenem Wachstum und quellenübergreifend normalisiert sortiert. Diese Ressourcen bleiben getrennt zugänglich und fließen nicht in die Hauptwertung ein.

  1. 1
    AktivQuelle öffnen ↗

    My complete journey to becoming an Agentic AI Engineer through structured learning, projects, experiments, and production-ready implementations of modern AI systems.

    githubRessourcegemessenes Wachstum
    18GitHub-Sternestabil
  2. 2
    AktivQuelle öffnen ↗

    50+ curated LLM observability tools PLUS 26 Agent Skills (several with runnable, unit-tested scripts) to build, evaluate, debug, secure & monitor reliable LLM apps. Tracing, evals, guardrails, LLMOps.

    github

    Installieren git clone https://github.com/ContextJet-ai/awesome-llm-observability ~/.claude/skills/awesome-llm-observability

    Ressourcegemessenes Wachstum
    33GitHub-Sterne+2 (+6.5 %)
  3. 3
    AktivQuelle öffnen ↗

    Handbook técnico aberto sobre engenharia de IA em produção: ML tradicional, LLMs, RAG, agentes, segurança, observabilidade, FinOps e deployment.

    githubRessourcegemessenes Wachstum
    1GitHub-Sternestabil
  4. 4
    AktivQuelle öffnen ↗

    Documentation-discovery telemetry for Claude Code — heat/cold maps, health grade, evidence-backed router fixes. 100% local, zero tokens. /tt

    github

    Installieren /plugin marketplace add Hedde/trigger_tree

    Ressourcegemessenes Wachstum
    14GitHub-Sternestabil
  5. 5

    tunelab

    Skill
    AktivQuelle öffnen ↗

    Claude Code plugin for LLM fine-tuning, distillation, and evaluation — decide whether you need fine-tuning at all, distill your LLM logs into small local models (MLX/LoRA), evaluate with held-out discipline, and learn the why at every step.

    github

    Installieren /plugin marketplace add rchaz/tunelab

    Ressourcegemessenes Wachstum
    6GitHub-Sternestabil