Die Werkzeugbeschreibungen sind auf Englisch.
LLM-Beobachtbarkeit
Monitoring, Traces, Bewertung und Qualität von LLM-Anwendungen.
Anwendungsfall
Aktivität
Sortieren nach
122
Werkzeug-Rangliste
Gemischte Rangfolge: Gemessenes Wachstum hat Vorrang. Rohwerte behalten ihr eigenes Zeitfenster; eine Änderung erfordert mindestens zwei Messungen in 7 Tagen.
- 1Ruhend
eval-layer
SonstigeA Claude Code skill that adds a rubric-based eval layer to any agent project. Framework-agnostic — generates rubric, test cases, judge prompt, and harness. Returns a weighted score plus a judge-leniency signal.
githubgemessenes WachstumQuelle öffnen ↗
13GitHub-Sternestabil - 2Aktiv
evoagent-os
AgentLocal-first control plane for durable, governed agent teams: DAG orchestration, approvals, memory, signed skills, trace contracts, and realtime.
githubgemessenes WachstumQuelle öffnen ↗
0GitHub-Sternestabil - 3Aktiv
memroos
AgentMemroOS / memroos: memory OS and governance layer for AI agents, agent workflows, dispatch, proof, and context continuity.
githubgemessenes WachstumQuelle öffnen ↗
7GitHub-Sternestabil - 4Aktiv
homestream
Sonstige🔑 HomeStream · 家园·流 — 零成本自托管多Agent协作框架,通往AI世界的那把钥匙 | Zero-cost self-hosted multi-agent framework — The key to AI world
githubgemessenes WachstumQuelle öffnen ↗
0GitHub-Sternestabil - 5Aktiv
adl-cli
SonstigeA command-line tool to scaffold and manage enterprise-ready AI Agents powered by the A2A (Agent-to-Agent) protocol
githubgemessenes WachstumQuelle öffnen ↗
14GitHub-Sterne+1 (+7.7 %) - 6Aktiv
simple-output-styles
SkillMake Claude write clearly, for everyone.
githubgemessenes WachstumQuelle öffnen ↗
Installieren
/plugin marketplace add stefanobaghino/simple-output-styles16GitHub-Sternestabil - 7Aktiv
agent-mmm
SkillMarketing Mix Model expert agent plugin for Claude Code - pymc-marketing v0.18.2+
githubgemessenes WachstumQuelle öffnen ↗
Installieren
/plugin marketplace add Yakoub-ai/agent-mmm4GitHub-Sternestabil - 8Aktiv
tulip-agents
AgentThe agent framework where the model never holds the trigger — every consequential action clears your policy first, waits for a human when it matters, and lands on a record you can verify. Build on it, or put it around the agent you already…
githubgemessenes WachstumQuelle öffnen ↗
1GitHub-Sternestabil - 9Ruhend
Agents-eval
AgentA Multi-Agent System (MAS) evaluation framework using PydanticAI that generates and evaluates scientific paper reviews through a three-tiered assessment approach: traditional metrics, LLM-as-a-Judge, and graph-based complexity analysis.
githubgemessenes WachstumQuelle öffnen ↗
2GitHub-Sternestabil - 10Ruhend
MCP as a Judge: a behavioral MCP that strengthens AI coding assistants via explicit LLM evaluations
mcpgemessenes WachstumQuelle öffnen ↗
Installieren
claude mcp add mcp-as-a-judge -- uvx mcp-as-a-judge17GitHub-Sternestabil - 11Ruhend
An AI-powered multi-agent system that demonstrates clinical triage, OTC medication recommendations, and e-pharmacy integration for respiratory conditions. Built with modular agents that collaborate to provide safe, intelligent healthcare…
githubgemessenes WachstumQuelle öffnen ↗
0GitHub-Sternestabil - 12Ruhend
CustoFlow
SonstigeMulti-agent customer support system with Google ADK & Gemini 2.5 Flash Lite. Kaggle capstone demonstrating 11+ concepts. Automates 80%+ queries, <10s response time.
githubgemessenes WachstumQuelle öffnen ↗
2GitHub-Sternestabil - 13Aktiv
Architecture-first Python scaffold for an auditable multi-agent corporate credit desk using A2A, MCP, deterministic credit policies, model routing, and OpenTelemetry.
githubgemessenes WachstumQuelle öffnen ↗
0GitHub-Sternestabil - 14Ruhend
Two Claude Code skills for adversarial code review (single-model PAR and multi-model MMAR with cross-critique to catch hallucinations), plus a fixture-based eval suite.
githubgemessenes WachstumQuelle öffnen ↗
Installieren
/plugin marketplace add prime-radiant-inc/parallel-adversarial-review17GitHub-Sternestabil - 15Aktiv
Mini Program Engineering Skill Suite is an Agent Skill suite for evidence‑first mini‑program development. It helps agents bring WeChat or other mini‑program projects from vague intent to reliable engineering work: project intake, product…
githubgemessenes WachstumQuelle öffnen ↗
Installieren
git clone https://github.com/NocodeMrLi/mini-program-engineering-skill-suite ~/.claude/skills/mini-program-engineering-skill-suite45GitHub-Sterne+6 (+15.4 %) - 16Aktiv
sc-prism-releases
SkillRun many AIs on one board and keep control of all of it. Deterministic code decides who acts — never a model. A privacy floor keeps sensitive work on your machine, your own tests decide what counts as done, and every action lands on a…
githubgemessenes WachstumQuelle öffnen ↗
Installieren
git clone https://github.com/sandhusukhdeep2/sc-prism-releases ~/.claude/skills/sc-prism-releases1GitHub-Sternestabil - 17Aktiv
superlog
SonstigeOpen-source observability tool that uses AI agents to self-heal your software
githubgeschätztes MomentumQuelle öffnen ↗
1 404GitHub-Sterne—
Lern- und Referenzressourcen
Nach gemessenem Wachstum und quellenübergreifend normalisiert sortiert. Diese Ressourcen bleiben getrennt zugänglich und fließen nicht in die Hauptwertung ein.
- 1AktivQuelle öffnen ↗
Agentic_AI_Engineer
AgentMy complete journey to becoming an Agentic AI Engineer through structured learning, projects, experiments, and production-ready implementations of modern AI systems.
githubRessourcegemessenes Wachstum18GitHub-Sternestabil - 2AktivQuelle öffnen ↗
50+ curated LLM observability tools PLUS 26 Agent Skills (several with runnable, unit-tested scripts) to build, evaluate, debug, secure & monitor reliable LLM apps. Tracing, evals, guardrails, LLMOps.
githubInstallieren
Ressourcegemessenes Wachstumgit clone https://github.com/ContextJet-ai/awesome-llm-observability ~/.claude/skills/awesome-llm-observability33GitHub-Sterne+2 (+6.5 %) - 3AktivQuelle öffnen ↗
Handbook técnico aberto sobre engenharia de IA em produção: ML tradicional, LLMs, RAG, agentes, segurança, observabilidade, FinOps e deployment.
githubRessourcegemessenes Wachstum1GitHub-Sternestabil - 4AktivQuelle öffnen ↗
trigger_tree
SkillDocumentation-discovery telemetry for Claude Code — heat/cold maps, health grade, evidence-backed router fixes. 100% local, zero tokens. /tt
githubInstallieren
Ressourcegemessenes Wachstum/plugin marketplace add Hedde/trigger_tree14GitHub-Sternestabil - 5AktivQuelle öffnen ↗
tunelab
SkillClaude Code plugin for LLM fine-tuning, distillation, and evaluation — decide whether you need fine-tuning at all, distill your LLM logs into small local models (MLX/LoRA), evaluate with held-out discipline, and learn the why at every step.
githubInstallieren
Ressourcegemessenes Wachstum/plugin marketplace add rchaz/tunelab6GitHub-Sternestabil