Die Werkzeugbeschreibungen sind auf Englisch.

LLM-Beobachtbarkeit

Monitoring, Traces, Bewertung und Qualität von LLM-Anwendungen.

Anwendungsfall

Aktivität

Sortieren nach

122

Werkzeug-Rangliste

Nach gemessenem Wachstum und quellenübergreifend normalisiert sortiert. Rohwerte behalten ihr eigenes Zeitfenster; eine Änderung erfordert mindestens zwei Messungen in 7 Tagen.

  1. 1

    Skill
    Aktiv

    Open-source infrastructure that turns scattered SKILL.md files into curated, retrieval-ready agent-skill corpora—with retrieval and evaluation tooling included.

    githubgemessenes WachstumQuelle öffnen ↗

    Installieren git clone https://github.com/EverMind-AI/SkillCorpus ~/.claude/skills/SkillCorpus

    424GitHub-Sterne+170 (+66.9 %)
  2. 2

    Sonstige
    Aktiv

    Test your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare performance of GPT, Claude, Gemini, DeepSeek, and more. Simple declarative configs with command line and CI/CD integration. Used by OpenAI…

    githubgemessenes WachstumQuelle öffnen ↗

    24 768GitHub-Sterne+142 (+0.58 %)
  3. 3

    Sonstige
    Aktiv

    Mastra is the modern TypeScript framework for AI-powered applications and agents.

    githubgemessenes WachstumQuelle öffnen ↗

    27 658GitHub-Sterne+128 (+0.46 %)
  4. 4

    Sonstige
    Aktiv

    Open source observability platform for logs, metrics, traces, frontend monitoring, pipelines and LLM observability. A sophisticated, simple and highly performant alternative to Datadog, Splunk, and Elasticsearch with 140x lower storage…

    githubgemessenes WachstumQuelle öffnen ↗

    21 615GitHub-Sterne+96 (+0.45 %)
  5. 5

    Sonstige
    Aktiv

    The fastest path to AI-powered full stack observability, even for lean teams.

    githubgemessenes WachstumQuelle öffnen ↗

    80 412GitHub-Sterne+85 (+0.11 %)
  6. 6

    Agent
    Aktiv

    The open source AI engineering platform for agents, LLMs, and ML models. MLflow enables teams of all sizes to debug, evaluate, monitor, and optimize production-quality AI applications while controlling costs and managing access to models…

    githubgemessenes WachstumQuelle öffnen ↗

    27 783GitHub-Sterne+82 (+0.30 %)
  7. 7

    Bibliothek
    Aktiv

    Prefect is a workflow orchestration framework for building resilient data pipelines in Python.

    githubgemessenes WachstumQuelle öffnen ↗

    23 766GitHub-Sterne+66 (+0.28 %)
  8. 8
    Aktiv

    Multi-tier framework for evaluating AI agent skills with quality gates, semantic overlap detection, synthetic evaluation dataset generation, and live agent evaluation that measures how skills affect agent behavior.

    githubgemessenes WachstumQuelle öffnen ↗

    Installieren git clone https://github.com/NVIDIA/SkillEvaluator ~/.claude/skills/SkillEvaluator

    389GitHub-Sterne+57 (+17.2 %)
  9. 9

    Sonstige
    Aktiv

    SigNoz is an open-source, OpenTelemetry-native observability platform for your team and their AI agents. Get logs, metrics, and traces in one tool with features like APM, distributed tracing, log management, infra monitoring, etc. Combined…

    githubgemessenes WachstumQuelle öffnen ↗

    32 003GitHub-Sterne+50 (+0.16 %)
  10. 10

    Sonstige
    Aktiv

    A high-performance observability data pipeline.

    githubgemessenes WachstumQuelle öffnen ↗

    22 509GitHub-Sterne+40 (+0.18 %)
  11. 11
    Aktiv

    YAO = Yielding AI Outcomes. A rigorous engineering, evaluation, governance, and portability system for reusable agent skills.

    githubgemessenes WachstumQuelle öffnen ↗

    Installieren git clone https://github.com/yaojingang/yao-meta-skill ~/.claude/skills/yao-meta-skill

    2 577GitHub-Sterne+35 (+1.4 %)
  12. 12

    Agent
    Aktiv

    DataBuff is an AI-native APM built on Opentelemetry,with multi-agent troubleshooting out of the box.

    githubgemessenes WachstumQuelle öffnen ↗

    642GitHub-Sterne+32 (+5.2 %)
  13. 13
    Aktiv

    🫖 Status page with uptime monitoring & API monitoring as code 🫖

    githubgemessenes WachstumQuelle öffnen ↗

    9 056GitHub-Sterne+31 (+0.34 %)
  14. 14

    Sonstige
    Aktiv

    the LLM vulnerability scanner

    githubgemessenes WachstumQuelle öffnen ↗

    9 079GitHub-Sterne+31 (+0.34 %)
  15. 15

    Sonstige
    Aktiv

    eBPF-based Networking, Security, and Observability

    githubgemessenes WachstumQuelle öffnen ↗

    25 047GitHub-Sterne+29 (+0.12 %)
  16. 16
    Aktiv

    Build tested agent skills and govern their lifecycle through a user-defined marketplace: evidence, discovery, updates, rollback, quarantine, and 17-platform distribution.

    githubgemessenes WachstumQuelle öffnen ↗

    Installieren git clone https://github.com/FrancyJGLisboa/agent-skill-creator ~/.claude/skills/agent-skill-creator

    2 367GitHub-Sterne+28 (+1.2 %)
  17. 17

    Skill
    Aktiv

    Deletion-first Agent Skill for removing test bloat, verification theater, and speculative fallbacks while preserving behavior.

    githubgemessenes WachstumQuelle öffnen ↗

    Installieren git clone https://github.com/MrZoyo/deslop-GPT ~/.claude/skills/deslop-GPT

    55GitHub-Sterne+27 (+96.4 %)
  18. 18

    Sonstige
    Aktiv

    Evidence-first infrastructure for reproducible, isolated executable evaluation and online rewards.

    githubgemessenes WachstumQuelle öffnen ↗

    140GitHub-Sterne+24 (+20.7 %)
  19. 19

    Agent
    Aktiv

    The Operating System for Scalable Enterprise AI Agents - Run, orchestrate, and deploy Compliant Enterprise AI Agents at scale across frameworks, without lock-in, rewrites or fragile glue code. Native support for MCP, A2A. Interface with all…

    githubgemessenes WachstumQuelle öffnen ↗

    156GitHub-Sterne+19 (+13.9 %)
  20. 20

    Skill
    Aktiv

    Research-backed, eval-driven skills for AI agents

    githubgemessenes WachstumQuelle öffnen ↗

    Installieren git clone https://github.com/ZSeven-W/craft-skills ~/.claude/skills/craft-skills

    142GitHub-Sterne+18 (+14.5 %)
  21. 21
    Aktiv

    A Go framework for building production agent systems with graph workflows, tools, memory, A2A, AG-UI, MCP, evaluation, and observability.

    githubgemessenes WachstumQuelle öffnen ↗

    1 759GitHub-Sterne+18 (+1.0 %)
  22. 22

    Skill
    Aktiv

    The Fable Workflow: how Claude Fable 5 worked, distilled into skills any model can run, with the eval that keeps it honest. Think / act / prove.

    githubgemessenes WachstumQuelle öffnen ↗

    Installieren git clone https://github.com/Sahir619/fable-method ~/.claude/skills/fable-method

    2 272GitHub-Sterne+16 (+0.71 %)
  23. 23
    Aktiv

    A test runner for agentskills.io-style AI agent skills

    githubgemessenes WachstumQuelle öffnen ↗

    Installieren git clone https://github.com/darkrishabh/agent-skills-eval ~/.claude/skills/agent-skills-eval

    719GitHub-Sterne+13 (+1.8 %)
  24. 24

    Agent
    Aktiv

    eBPF-powered network observability for Kubernetes. Indexes L4/L7 traffic with full K8s context, decrypts TLS without keys. Queryable by AI agents via MCP and humans via dashboard.

    githubgemessenes WachstumQuelle öffnen ↗

    12 068GitHub-Sterne+8 (+0.07 %)
  25. 25

    Skill
    Aktiv

    The container platform tailored for Kubernetes multi-cloud, datacenter, and edge management ⎈ 🖥 ☁️

    githubgemessenes WachstumQuelle öffnen ↗

    Installieren git clone https://github.com/kubesphere/kubesphere ~/.claude/skills/kubesphere

    17 037GitHub-Sterne+8 (+0.05 %)