LLM observability: tools for developers
Monitoring, traces, evaluation and quality tracking for LLM applications. An entry can appear under several use cases when its topics or description provide several explicit signals.
Use case
Activity
Sort by
123 entries in this view.
Tool ranking
Ranked by measured growth, normalized across sources. Raw values keep their source-specific window, shown under each tool; a change requires at least two snapshots from the last 7 days.
- 1Active
Amazon Bedrock AgentCore enterprise platform accelerator with AWS CDK, Terraform organization guardrails, MCP/A2A agents, security, memory, and observability.
githubmeasured growthOpen source ↗
5GitHub stars+1 (+25.0 %) - 2Active
claudestat
SkillReal-time execution trace and cost intelligence for Claude Code
githubmeasured growthOpen source ↗
Install
/plugin marketplace add DeibyGS/claudestat34GitHub stars+3 (+9.7 %) - 3Active
openstatus
MCP🫖 Status page with uptime monitoring & API monitoring as code 🫖
githubmeasured growthOpen source ↗
9 036GitHub stars+25 (+0.28 %) - 4Active
rote
OtherA cron that remembers what it did—a CLI job scheduler with run history, captured output, and a live terminal dashboard.
githubmeasured growthOpen source ↗
216GitHub stars+7 (+3.3 %) - 5Active
untell
OtherAI-detector auditing toolkit: measures how often a detector flags genuine human writing, how stable its verdict is across seeds, and whether that verdict survives meaning-preserving edits. Detector-in-the-loop measurement harness. Claude…
githubmeasured growthOpen source ↗
19GitHub stars+2 (+11.8 %) - 6Active
skill-eval-harness
SkillAgent Skill evaluation harness for paired variants, trace artifacts, and runner adapters
githubmeasured growthOpen source ↗
Install
git clone https://github.com/adewale/skill-eval-harness ~/.claude/skills/skill-eval-harness72GitHub stars+4 (+5.9 %) - 7Active
cilium
OthereBPF-based Networking, Security, and Observability
githubmeasured growthOpen source ↗
25 028GitHub stars+24 (+0.10 %) - 8Active
databuff
AgentDataBuff is an AI-native APM built on Opentelemetry,with multi-agent troubleshooting out of the box.
githubmeasured growthOpen source ↗
615GitHub stars+10 (+1.7 %) - 9Active
Deterministic, local-only audit reports for Claude Code AI agent sessions. Rust + SQLite. Zero network calls. MCP server for agents.
githubmeasured growthOpen source ↗
22GitHub stars+2 (+10.0 %) - 10Active
dynatrace-for-ai
SkillSkills, prompts, and instructions for building AI agents on top of Dynatrace production context
githubmeasured growthOpen source ↗
Install
git clone https://github.com/Dynatrace/dynatrace-for-ai ~/.claude/skills/dynatrace-for-ai131GitHub stars+5 (+4.0 %) - 11Active
OpenJudge
SkillOpenJudge: A Unified Framework for Holistic Evaluation and Quality Rewards
githubmeasured growthOpen source ↗
Install
git clone https://github.com/agentscope-ai/OpenJudge ~/.claude/skills/OpenJudge807GitHub stars+10 (+1.3 %) - 12Active
Vendor-neutral OpenTelemetry skills for AI coding agents, grounded in upstream sources
githubmeasured growthOpen source ↗
Install
git clone https://github.com/ollygarden/opentelemetry-agent-skills ~/.claude/skills/opentelemetry-agent-skills96GitHub stars+4 (+4.3 %) - 13Active
ai-dev-stack
SkillProduction-grade AI coding rules for Cursor and Claude Code. 15 rules + 9 doc templates + skills + agents + MCP setup. Drop into any project.
githubmeasured growthOpen source ↗
Install
git clone https://github.com/aiagentwithdhruv/ai-dev-stack ~/.claude/skills/ai-dev-stack10GitHub stars+1 (+11.1 %) - 14Active
superlog
OtherOpen-source observability tool that uses AI agents to self-heal your software
githubmeasured growthOpen source ↗
1 404GitHub stars+10 (+0.72 %) - 15Dormant
eval-layer
OtherA Claude Code skill that adds a rubric-based eval layer to any agent project. Framework-agnostic — generates rubric, test cases, judge prompt, and harness. Returns a weighted score plus a judge-leniency signal.
githubmeasured growthOpen source ↗
13GitHub stars+1 (+8.3 %) - 16Active
SkillForge
SkillA skill creator that proves its skills work. Evidence-driven skill creation for Claude Code and Codex: baseline-tested generation, per-skill regression evals, ecosystem doctor, cross-runtime compile, and an opt-in proactive advisor.
githubmeasured growthOpen source ↗
Install
git clone https://github.com/tripleyak/SkillForge ~/.claude/skills/SkillForge884GitHub stars+6 (+0.68 %) - 17Active
kubesphere
SkillThe container platform tailored for Kubernetes multi-cloud, datacenter, and edge management ⎈ 🖥 ☁️
githubmeasured growthOpen source ↗
Install
git clone https://github.com/kubesphere/kubesphere ~/.claude/skills/kubesphere17 035GitHub stars+6 (+0.04 %) - 18Active
opus-fable-playbook
SkillMake Claude Opus 4.8 behave like Claude Fable 5 — doctrine output style, drift-catching hooks, and an eval loop against golden Fable transcripts. Claude Code plugin.
githubmeasured growthOpen source ↗
Install
/plugin marketplace add rennf93/opus-fable-playbook33GitHub stars+1 (+3.1 %) - 19Active
kubeshark
AgenteBPF-powered network observability for Kubernetes. Indexes L4/L7 traffic with full K8s context, decrypts TLS without keys. Queryable by AI agents via MCP and humans via dashboard.
githubmeasured growthOpen source ↗
12 063GitHub stars+5 (+0.04 %) - 20Active
cap-evolve
MCPOptimize any AI agent’s skills, tools/MCP, and prompts against your own evals.
githubmeasured growthOpen source ↗
47GitHub stars+1 (+2.2 %) - 21Active
nora
MCPOpen-source, self-hosted control plane for OpenClaw and Hermes AI-agent fleets on Docker/Kubernetes — REST, CLI, and MCP.
githubmeasured growthOpen source ↗
49GitHub stars+1 (+2.1 %) - 22Active
langfuse-docs
Skill🪢 Langfuse documentation -- Langfuse is the open source LLM Engineering Platform. Observability, evals, prompt management, playground and metrics to debug and improve LLM apps
githubmeasured growthOpen source ↗
Install
git clone https://github.com/langfuse/langfuse-docs ~/.claude/skills/langfuse-docs237GitHub stars+2 (+0.85 %) - 23Active
aura
MCPAURA is a production-tested SRE agent platform you can deploy in minutes. AURA handles the guardrails, APIs, state management, streaming, and failure handling required to put AI to work safely on production infrastructure.
githubmeasured growthOpen source ↗
258GitHub stars+2 (+0.78 %) - 24Active
claude-code-karma
SkillDashboard for monitoring claude code sessions.
githubmeasured growthOpen source ↗
Install
/plugin marketplace add JayantDevkar/claude-code-karma321GitHub stars+2 (+0.63 %)
Learning resources
Ranked by measured growth, normalized across sources. These resources remain available separately and do not take part in the main tool ranking.
- 1ActiveOpen source ↗
50+ curated LLM observability tools PLUS 26 Agent Skills (several with runnable, unit-tested scripts) to build, evaluate, debug, secure & monitor reliable LLM apps. Tracing, evals, guardrails, LLMOps.
githubInstall
Resourcemeasured growthgit clone https://github.com/ContextJet-ai/awesome-llm-observability ~/.claude/skills/awesome-llm-observability31GitHub stars+1 (+3.3 %)