LLM observability: tools for developers
Monitoring, traces, evaluation and quality tracking for LLM applications. An entry can appear under several use cases when its topics or description provide several explicit signals.
Use case
Activity
Sort by
74 entries in this view.
Tool ranking
Ranked by normalized popularity across sources. Raw values keep their source-specific window, shown under each tool; a change requires at least two snapshots from the last 7 days.
- 1Active
agent-kernel
AgentThe Operating System for Scalable Enterprise AI Agents - Run, orchestrate, and deploy Compliant Enterprise AI Agents at scale across frameworks, without lock-in, rewrites or fragile glue code. Native support for MCP, A2A. Interface with all…
githubmeasured growthOpen source ↗
156GitHub stars+19 (+13.9 %) - 2Active
craft-skills
SkillResearch-backed, eval-driven skills for AI agents
githubmeasured growthOpen source ↗
Install
git clone https://github.com/ZSeven-W/craft-skills ~/.claude/skills/craft-skills142GitHub stars+18 (+14.5 %) - 3Active
VeriRun
OtherEvidence-first infrastructure for reproducible, isolated executable evaluation and online rewards.
githubmeasured growthOpen source ↗
140GitHub stars+24 (+20.7 %) - 4Active
adlc-team-skills
Skill🐙 ADLC Team Skills — Agentic SDLC for Engineering Teams
githubmeasured growthOpen source ↗
Install
git clone https://github.com/tikalk/adlc-team-skills ~/.claude/skills/adlc-team-skills132GitHub stars+1 (+0.76 %) - 5Active
dynatrace-for-ai
SkillSkills, prompts, and instructions for building AI agents on top of Dynatrace production context
githubmeasured growthOpen source ↗
Install
git clone https://github.com/Dynatrace/dynatrace-for-ai ~/.claude/skills/dynatrace-for-ai132GitHub stars+4 (+3.1 %) - 6Active
74 open-source Agent Skills for Claude Code and Codex: AI SEO, AEO and GEO, code review with an A-F ship grade, CI gates, AI evals, design systems, conversion copy, Instagram growth, iOS and Android app shipping, creator rights, and…
githubmeasured growthOpen source ↗
129GitHub stars-4 (-3.0 %) - 7Active
langfuse-mcp
MCPA Model Context Protocol (MCP) server for Langfuse, enabling AI agents to query Langfuse trace data for enhanced debugging and observability
githubmeasured growthOpen source ↗
105GitHub starsstable - 8105GitHub stars+1 (+0.96 %)
- 9Active
Vendor-neutral OpenTelemetry skills for AI coding agents, grounded in upstream sources
githubmeasured growthOpen source ↗
Install
git clone https://github.com/ollygarden/opentelemetry-agent-skills ~/.claude/skills/opentelemetry-agent-skills97GitHub stars+4 (+4.3 %) - 10Active
skill-kit
Skilllocal-first analytics for AI agent skills
githubmeasured growthOpen source ↗
Install
git clone https://github.com/crafter-station/skill-kit ~/.claude/skills/skill-kit77GitHub stars+1 (+1.3 %) - 11Active
skill-eval-harness
SkillAgent Skill evaluation harness for paired variants, trace artifacts, and runner adapters
githubmeasured growthOpen source ↗
Install
git clone https://github.com/adewale/skill-eval-harness ~/.claude/skills/skill-eval-harness73GitHub stars+4 (+5.8 %) - 12Active
AgentX-Python
OtherAgentX python SDK. Build multi-agent AI workforce. Run evaluation. Trace your agent. Full Observability.
githubmeasured growthOpen source ↗
68GitHub starsstable - 13Active
deslop-GPT
SkillDeletion-first Agent Skill for removing test bloat, verification theater, and speculative fallbacks while preserving behavior.
githubmeasured growthOpen source ↗
Install
git clone https://github.com/MrZoyo/deslop-GPT ~/.claude/skills/deslop-GPT55GitHub stars+27 (+96.4 %) - 14Active
seo-skill-bench
SkillOpen benchmark for Claude Code SEO skills — real headless execution against fixture sites with planted-defect answer keys. Deterministic scoring, pre-registered rubric.
githubmeasured growthOpen source ↗
Install
git clone https://github.com/aleclindz/seo-skill-bench ~/.claude/skills/seo-skill-bench51GitHub stars-5 (-8.9 %) - 15Active
nora
MCPOpen-source, self-hosted control plane for OpenClaw and Hermes AI-agent fleets on Docker/Kubernetes — REST, CLI, and MCP.
githubmeasured growthOpen source ↗
50GitHub stars+2 (+4.2 %) - 16Active
arize-skills
SkillAgent skills for Arize — datasets, experiments, and traces via the ax CLI
githubmeasured growthOpen source ↗
Install
git clone https://github.com/Arize-ai/arize-skills ~/.claude/skills/arize-skills47GitHub starsstable - 17Active
cap-evolve
MCPOptimize any AI agent’s skills, tools/MCP, and prompts against your own evals.
githubmeasured growthOpen source ↗
47GitHub stars+1 (+2.2 %) - 18Active
Mini Program Engineering Skill Suite is an Agent Skill suite for evidence‑first mini‑program development. It helps agents bring WeChat or other mini‑program projects from vague intent to reliable engineering work: project intake, product…
githubestimated momentumOpen source ↗
Install
git clone https://github.com/NocodeMrLi/mini-program-engineering-skill-suite ~/.claude/skills/mini-program-engineering-skill-suite39GitHub stars— - 19Active
claudestat
SkillReal-time execution trace and cost intelligence for Claude Code
githubmeasured growthOpen source ↗
Install
/plugin marketplace add DeibyGS/claudestat34GitHub stars+1 (+3.0 %) - 20Active
Deterministic, local-only audit reports for Claude Code AI agent sessions. Rust + SQLite. Zero network calls. MCP server for agents.
githubmeasured growthOpen source ↗
22GitHub stars+2 (+10.0 %) - 21Active
untell
OtherAI-detector auditing toolkit: measures how often a detector flags genuine human writing, how stable its verdict is across seeds, and whether that verdict survives meaning-preserving edits. Detector-in-the-loop measurement harness. Claude…
githubmeasured growthOpen source ↗
18GitHub starsstable - 22Active
rashomon
SkillMeasure prompt and skill improvements with blind A/B comparison.
githubmeasured growthOpen source ↗
Install
/plugin marketplace add shinpr/rashomon18GitHub starsstable
Learning resources
Ranked by normalized popularity across sources. These resources remain available separately and do not take part in the main tool ranking.
- 1ActiveOpen source ↗
Curated + scored map of official MCP servers and agents for DevOps, Cloud, SRE, and Platform Engineering — every entry rated on production access, approval gates, and audit evidence.
githubResourcemeasured growth77GitHub starsstable - 2ActiveOpen source ↗
50+ curated LLM observability tools PLUS 26 Agent Skills (several with runnable, unit-tested scripts) to build, evaluate, debug, secure & monitor reliable LLM apps. Tracing, evals, guardrails, LLMOps.
githubInstall
Resourcemeasured growthgit clone https://github.com/ContextJet-ai/awesome-llm-observability ~/.claude/skills/awesome-llm-observability32GitHub stars+2 (+6.7 %) - 3ActiveOpen source ↗
Agentic_AI_Engineer
AgentMy complete journey to becoming an Agentic AI Engineer through structured learning, projects, experiments, and production-ready implementations of modern AI systems.
githubResourcemeasured growth18GitHub starsstable