LLM observability: tools for developers
Monitoring, traces, evaluation and quality tracking for LLM applications. An entry can appear under several use cases when its topics or description provide several explicit signals.
Use case
Activity
Sort by
123 entries in this view.
Tool ranking
Ranked by normalized popularity across sources. Raw values keep their source-specific window, shown under each tool; a change requires at least two snapshots from the last 7 days.
- 1Active
promptfoo
OtherTest your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare performance of GPT, Claude, Gemini, DeepSeek, and more. Simple declarative configs with command line and CI/CD integration. Used by OpenAI…
githubmeasured growthOpen source ↗
24 738GitHub stars+131 (+0.53 %) - 2Active
signoz
OtherSigNoz is an open-source, OpenTelemetry-native observability platform for your team and their AI agents. Get logs, metrics, and traces in one tool with features like APM, distributed tracing, log management, infra monitoring, etc. Combined…
githubmeasured growthOpen source ↗
31 992GitHub stars+62 (+0.19 %) - 3Active
mlflow
AgentThe open source AI engineering platform for agents, LLMs, and ML models. MLflow enables teams of all sizes to debug, evaluate, monitor, and optimize production-quality AI applications while controlling costs and managing access to models…
githubmeasured growthOpen source ↗
27 768GitHub stars+80 (+0.29 %) - 4Active
prefect
LibraryPrefect is a workflow orchestration framework for building resilient data pipelines in Python.
githubmeasured growthOpen source ↗
23 759GitHub stars+66 (+0.28 %) - 522 507GitHub stars+45 (+0.20 %)
- 6Active
openobserve
OtherOpen source observability platform for logs, metrics, traces, frontend monitoring, pipelines and LLM observability. A sophisticated, simple and highly performant alternative to Datadog, Splunk, and Elasticsearch with 140x lower storage…
githubmeasured growthOpen source ↗
21 608GitHub stars+115 (+0.54 %) - 7Active
mastra
OtherMastra is the modern TypeScript framework for AI-powered applications and agents.
githubmeasured growthOpen source ↗
27 624GitHub stars+121 (+0.44 %) - 8Active
netdata
OtherThe fastest path to AI-powered full stack observability, even for lean teams.
githubmeasured growthOpen source ↗
80 402GitHub stars+91 (+0.11 %) - 9Active
cilium
OthereBPF-based Networking, Security, and Observability
githubmeasured growthOpen source ↗
25 041GitHub stars+27 (+0.11 %) - 10Active
kubeshark
AgenteBPF-powered network observability for Kubernetes. Indexes L4/L7 traffic with full K8s context, decrypts TLS without keys. Queryable by AI agents via MCP and humans via dashboard.
githubmeasured growthOpen source ↗
12 066GitHub stars+7 (+0.06 %) - 11Active
kubesphere
SkillThe container platform tailored for Kubernetes multi-cloud, datacenter, and edge management ⎈ 🖥 ☁️
githubmeasured growthOpen source ↗
Install
git clone https://github.com/kubesphere/kubesphere ~/.claude/skills/kubesphere17 037GitHub stars+9 (+0.05 %) - 129 079GitHub stars+43 (+0.48 %)
- 13Active
openstatus
MCP🫖 Status page with uptime monitoring & API monitoring as code 🫖
githubmeasured growthOpen source ↗
9 052GitHub stars+28 (+0.31 %) - 14Active
yao-meta-skill
SkillYAO = Yielding AI Outcomes. A rigorous engineering, evaluation, governance, and portability system for reusable agent skills.
githubmeasured growthOpen source ↗
Install
git clone https://github.com/yaojingang/yao-meta-skill ~/.claude/skills/yao-meta-skill2 577GitHub stars+54 (+2.1 %) - 15Active
agent-skill-creator
SkillBuild tested agent skills and govern their lifecycle through a user-defined marketplace: evidence, discovery, updates, rollback, quarantine, and 17-platform distribution.
githubmeasured growthOpen source ↗
Install
git clone https://github.com/FrancyJGLisboa/agent-skill-creator ~/.claude/skills/agent-skill-creator2 367GitHub stars+30 (+1.3 %) - 16Active
fable-method
SkillThe Fable Workflow: how Claude Fable 5 worked, distilled into skills any model can run, with the eval that keeps it honest. Think / act / prove.
githubmeasured growthOpen source ↗
Install
git clone https://github.com/Sahir619/fable-method ~/.claude/skills/fable-method2 272GitHub stars+19 (+0.84 %) - 17Active
A Go framework for building production agent systems with graph workflows, tools, memory, A2A, AG-UI, MCP, evaluation, and observability.
githubmeasured growthOpen source ↗
1 759GitHub stars+24 (+1.4 %) - 18Active
superlog
OtherOpen-source observability tool that uses AI agents to self-heal your software
githubmeasured growthOpen source ↗
1 404GitHub stars+2 (+0.14 %) - 19Active
SkillForge
SkillA skill creator that proves its skills work. Evidence-driven skill creation for Claude Code and Codex: baseline-tested generation, per-skill regression evals, ecosystem doctor, cross-runtime compile, and an opt-in proactive advisor.
githubmeasured growthOpen source ↗
Install
git clone https://github.com/tripleyak/SkillForge ~/.claude/skills/SkillForge884GitHub stars+3 (+0.34 %) - 20Active
OpenJudge
SkillOpenJudge: A Unified Framework for Holistic Evaluation and Quality Rewards
githubmeasured growthOpen source ↗
Install
git clone https://github.com/agentscope-ai/OpenJudge ~/.claude/skills/OpenJudge809GitHub stars+10 (+1.3 %) - 21Active
Flawless
OtherAI SRE AgenticOps for Kubernetes and cloud infrastructure.
githubmeasured growthOpen source ↗
782GitHub stars+1 (+0.13 %) - 22Active
agent-skills-eval
SkillA test runner for agentskills.io-style AI agent skills
githubmeasured growthOpen source ↗
Install
git clone https://github.com/darkrishabh/agent-skills-eval ~/.claude/skills/agent-skills-eval719GitHub stars+15 (+2.1 %) - 23Active
databuff
AgentDataBuff is an AI-native APM built on Opentelemetry,with multi-agent troubleshooting out of the box.
githubmeasured growthOpen source ↗
634GitHub stars+26 (+4.3 %) - 24Active
SkillCorpus
SkillOpen-source infrastructure that turns scattered SKILL.md files into curated, retrieval-ready agent-skill corpora—with retrieval and evaluation tooling included.
githubmeasured growthOpen source ↗
Install
git clone https://github.com/EverMind-AI/SkillCorpus ~/.claude/skills/SkillCorpus424GitHub stars+204 (+92.7 %) - 25Active
SkillEvaluator
SkillMulti-tier framework for evaluating AI agent skills with quality gates, semantic overlap detection, synthetic evaluation dataset generation, and live agent evaluation that measures how skills affect agent behavior.
githubmeasured growthOpen source ↗
Install
git clone https://github.com/NVIDIA/SkillEvaluator ~/.claude/skills/SkillEvaluator389GitHub stars+63 (+19.3 %)