LLM observability: tools for developers
Monitoring, traces, evaluation and quality tracking for LLM applications. An entry can appear under several use cases when its topics or description provide several explicit signals.
Use case
Activity
Sort by
118 entries in this view.
Tool ranking
Ranked by measured growth, normalized across sources. Raw values keep their source-specific window, shown under each tool; a change requires at least two snapshots from the last 7 days.
- 1Active
SkillCorpus
SkillOpen-source infrastructure that turns scattered SKILL.md files into curated, retrieval-ready agent-skill corpora—with retrieval and evaluation tooling included.
githubmeasured growthOpen source ↗
Install
git clone https://github.com/EverMind-AI/SkillCorpus ~/.claude/skills/SkillCorpus360GitHub stars+200 (+125.0 %) - 2Active
craft-skills
SkillResearch-backed, eval-driven skills for AI agents
githubmeasured growthOpen source ↗
Install
git clone https://github.com/ZSeven-W/craft-skills ~/.claude/skills/craft-skills138GitHub stars+56 (+68.3 %) - 3Active
SkillEvaluator
SkillMulti-tier framework for evaluating AI agent skills with quality gates, semantic overlap detection, synthetic evaluation dataset generation, and live agent evaluation that measures how skills affect agent behavior.
githubmeasured growthOpen source ↗
Install
git clone https://github.com/NVIDIA/SkillEvaluator ~/.claude/skills/SkillEvaluator363GitHub stars+70 (+23.9 %) - 4Active
deslop-GPT
SkillDeletion-first Agent Skill for removing test bloat, verification theater, and speculative fallbacks while preserving behavior.
githubmeasured growthOpen source ↗
Install
git clone https://github.com/MrZoyo/deslop-GPT ~/.claude/skills/deslop-GPT38GitHub stars+10 (+35.7 %) - 5Active
yao-meta-skill
SkillYAO = Yielding AI Outcomes. A rigorous engineering, evaluation, governance, and portability system for reusable agent skills.
githubmeasured growthOpen source ↗
Install
git clone https://github.com/yaojingang/yao-meta-skill ~/.claude/skills/yao-meta-skill2 565GitHub stars+69 (+2.8 %) - 6Active
promptfoo
OtherTest your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare performance of GPT, Claude, Gemini, DeepSeek, and more. Simple declarative configs with command line and CI/CD integration. Used by OpenAI…
githubmeasured growthOpen source ↗
24 683GitHub stars+139 (+0.57 %) - 7Active
mastra
OtherMastra is the modern TypeScript framework for AI-powered applications and agents.
githubmeasured growthOpen source ↗
27 580GitHub stars+139 (+0.51 %) - 8Active
openobserve
OtherOpen source observability platform for logs, metrics, traces, frontend monitoring, pipelines and LLM observability. A sophisticated, simple and highly performant alternative to Datadog, Splunk, and Elasticsearch with 140x lower storage…
githubmeasured growthOpen source ↗
21 575GitHub stars+128 (+0.60 %) - 9Active
agent-kernel
AgentThe Operating System for Scalable Enterprise AI Agents - Run, orchestrate, and deploy Compliant Enterprise AI Agents at scale across frameworks, without lock-in, rewrites or fragile glue code. Native support for MCP, A2A. Interface with all…
githubmeasured growthOpen source ↗
146GitHub stars+16 (+12.3 %) - 10Active
nudge
MCPA typed, replayable, budget-aware programming language for LLM agents. Compiles to Python & TypeScript.
githubmeasured growthOpen source ↗
2GitHub stars+1 (+100.0 %) - 11Active
agent-stack
SkillProduction patterns for agent orchestrators, harnesses, evals, MCP/A2A interoperability, memory, provider routing, and LLM usage metering.
githubmeasured growthOpen source ↗
Install
git clone https://github.com/ssheleg/agent-stack ~/.claude/skills/agent-stack2GitHub stars+1 (+100.0 %) - 12Active
historian
AgentA local A2A event and continuity service for agent applications, with structured history, provenance-safe transcripts, literal search, and natural-language queries.
githubmeasured growthOpen source ↗
1GitHub stars+1 - 13Active
skill-graveyard
SkillAudit which Claude Code skills you actually use — surface dead installs and hallucinated invocations from your session logs.
githubmeasured growthOpen source ↗
Install
/plugin marketplace add sfrangulov/skill-graveyard10GitHub stars+3 (+42.9 %) - 14Active
mlflow
AgentThe open source AI engineering platform for agents, LLMs, and ML models. MLflow enables teams of all sizes to debug, evaluate, monitor, and optimize production-quality AI applications while controlling costs and managing access to models…
githubmeasured growthOpen source ↗
27 741GitHub stars+78 (+0.28 %) - 159 079GitHub stars+56 (+0.62 %)
- 16Active
netdata
OtherThe fastest path to AI-powered full stack observability, even for lean teams.
githubmeasured growthOpen source ↗
80 361GitHub stars+77 (+0.10 %) - 17Active
prefect
LibraryPrefect is a workflow orchestration framework for building resilient data pipelines in Python.
githubmeasured growthOpen source ↗
23 727GitHub stars+56 (+0.24 %) - 18Active
agent-skill-creator
SkillBuild tested agent skills and govern their lifecycle through a user-defined marketplace: evidence, discovery, updates, rollback, quarantine, and 17-platform distribution.
githubmeasured growthOpen source ↗
Install
git clone https://github.com/FrancyJGLisboa/agent-skill-creator ~/.claude/skills/agent-skill-creator2 358GitHub stars+29 (+1.2 %) - 19Active
A Go framework for building production agent systems with graph workflows, tools, memory, A2A, AG-UI, MCP, evaluation, and observability.
githubmeasured growthOpen source ↗
1 752GitHub stars+26 (+1.5 %) - 2022 489GitHub stars+41 (+0.18 %)
- 21Active
signoz
OtherSigNoz is an open-source, OpenTelemetry-native observability platform for your team and their AI agents. Get logs, metrics, and traces in one tool with features like APM, distributed tracing, log management, infra monitoring, etc. Combined…
githubmeasured growthOpen source ↗
31 973GitHub stars+42 (+0.13 %) - 22Active
fable-method
SkillThe Fable Workflow: how Claude Fable 5 worked, distilled into skills any model can run, with the eval that keeps it honest. Think / act / prove.
githubmeasured growthOpen source ↗
Install
git clone https://github.com/Sahir619/fable-method ~/.claude/skills/fable-method2 269GitHub stars+24 (+1.1 %) - 23Active
agent-skills-eval
SkillA test runner for agentskills.io-style AI agent skills
githubmeasured growthOpen source ↗
Install
git clone https://github.com/darkrishabh/agent-skills-eval ~/.claude/skills/agent-skills-eval713GitHub stars+14 (+2.0 %) - 24Dormant
kiboserve
MCPOpen-source Python framework to deploy AI agents via HTTP, A2A, and MCP with built-in observability
githubmeasured growthOpen source ↗
5GitHub stars+1 (+25.0 %)
Learning resources
Ranked by measured growth, normalized across sources. These resources remain available separately and do not take part in the main tool ranking.
- 1ActiveOpen source ↗
trigger_tree
SkillDocumentation-discovery telemetry for Claude Code — heat/cold maps, health grade, evidence-backed router fixes. 100% local, zero tokens. /tt
githubInstall
Resourcemeasured growth/plugin marketplace add Hedde/trigger_tree14GitHub stars+2 (+16.7 %)