LLM observability: tools for developers
Monitoring, traces, evaluation and quality tracking for LLM applications. An entry can appear under several use cases when its topics or description provide several explicit signals.
Use case
Activity
Sort by
117 entries in this view.
Tool ranking
Ranked by measured growth, normalized across sources. Raw values keep their source-specific window, shown under each tool; a change requires at least two snapshots from the last 7 days.
- 1Active
OpenJudge
SkillOpenJudge: A Unified Framework for Holistic Evaluation and Quality Rewards
githubmeasured growthOpen source ↗
Install
git clone https://github.com/agentscope-ai/OpenJudge ~/.claude/skills/OpenJudge812GitHub stars+7 (+0.87 %) - 2Active
dynatrace-for-ai
SkillSkills, prompts, and instructions for building AI agents on top of Dynatrace production context
githubmeasured growthOpen source ↗
Install
git clone https://github.com/Dynatrace/dynatrace-for-ai ~/.claude/skills/dynatrace-for-ai135GitHub stars+4 (+3.1 %) - 3Active
skill-eval-harness
SkillAgent Skill evaluation harness for paired variants, trace artifacts, and runner adapters
githubmeasured growthOpen source ↗
Install
git clone https://github.com/adewale/skill-eval-harness ~/.claude/skills/skill-eval-harness73GitHub stars+4 (+5.8 %) - 4Active
Mini Program Engineering Skill Suite is an Agent Skill suite for evidence‑first mini‑program development. It helps agents bring WeChat or other mini‑program projects from vague intent to reliable engineering work: project intake, product…
githubmeasured growthOpen source ↗
Install
git clone https://github.com/NocodeMrLi/mini-program-engineering-skill-suite ~/.claude/skills/mini-program-engineering-skill-suite43GitHub stars+4 (+10.3 %) - 5Active
langfuse-docs
Skill🪢 Langfuse documentation -- Langfuse is the open source LLM Engineering Platform. Observability, evals, prompt management, playground and metrics to debug and improve LLM apps
githubmeasured growthOpen source ↗
Install
git clone https://github.com/langfuse/langfuse-docs ~/.claude/skills/langfuse-docs239GitHub stars+3 (+1.3 %) - 6Active
kubesphere
SkillThe container platform tailored for Kubernetes multi-cloud, datacenter, and edge management ⎈ 🖥 ☁️
githubmeasured growthOpen source ↗
Install
git clone https://github.com/kubesphere/kubesphere ~/.claude/skills/kubesphere17 035GitHub stars+2 (+0.01 %) - 7Active
Amazon Bedrock AgentCore enterprise platform accelerator with AWS CDK, Terraform organization guardrails, MCP/A2A agents, security, memory, and observability.
githubmeasured growthOpen source ↗
6GitHub stars+2 (+50.0 %) - 8Active
claude-code-karma
SkillDashboard for monitoring claude code sessions.
githubmeasured growthOpen source ↗
Install
/plugin marketplace add JayantDevkar/claude-code-karma323GitHub stars+2 (+0.62 %) - 9Active
MCP server for Langfuse LLM observability — trace and observation analysis.
mcpmeasured growthOpen source ↗
Install
claude mcp add langfuse -- npx langfuse-observability-mcp-server80GitHub stars+2 (+2.6 %) - 10Active
nora
MCPOpen-source, self-hosted control plane for OpenClaw and Hermes AI-agent fleets on Docker/Kubernetes — REST, CLI, and MCP.
githubmeasured growthOpen source ↗
50GitHub stars+2 (+4.2 %) - 11Active
Deterministic, local-only audit reports for Claude Code AI agent sessions. Rust + SQLite. Zero network calls. MCP server for agents.
githubmeasured growthOpen source ↗
22GitHub stars+2 (+10.0 %) - 12Active
Flawless
OtherAI SRE AgenticOps for Kubernetes and cloud infrastructure.
githubmeasured growthOpen source ↗
782GitHub stars+1 (+0.13 %) - 13Active
adl-cli
OtherA command-line tool to scaffold and manage enterprise-ready AI Agents powered by the A2A (Agent-to-Agent) protocol
githubmeasured growthOpen source ↗
14GitHub stars+1 (+7.7 %) - 14Active
ai-dev-stack
SkillProduction-grade AI coding rules for Cursor and Claude Code. 15 rules + 9 doc templates + skills + agents + MCP setup. Drop into any project.
githubmeasured growthOpen source ↗
Install
git clone https://github.com/aiagentwithdhruv/ai-dev-stack ~/.claude/skills/ai-dev-stack10GitHub stars+1 (+11.1 %) - 15Active
skill-graveyard
SkillAudit which Claude Code skills you actually use — surface dead installs and hallucinated invocations from your session logs.
githubmeasured growthOpen source ↗
Install
/plugin marketplace add sfrangulov/skill-graveyard10GitHub stars+1 (+11.1 %) - 16Active
adlc-team-skills
Skill🐙 ADLC Team Skills — Agentic SDLC for Engineering Teams
githubmeasured growthOpen source ↗
Install
git clone https://github.com/tikalk/adlc-team-skills ~/.claude/skills/adlc-team-skills132GitHub stars+1 (+0.76 %) - 17105GitHub stars+1 (+0.96 %)
- 18Active
Vendor-neutral OpenTelemetry skills for AI coding agents, grounded in upstream sources
githubmeasured growthOpen source ↗
Install
git clone https://github.com/ollygarden/opentelemetry-agent-skills ~/.claude/skills/opentelemetry-agent-skills97GitHub stars+1 (+1.0 %) - 19Active
skill-kit
Skilllocal-first analytics for AI agent skills
githubmeasured growthOpen source ↗
Install
git clone https://github.com/crafter-station/skill-kit ~/.claude/skills/skill-kit77GitHub stars+1 (+1.3 %) - 20Active
opus-fable-playbook
SkillMake Claude Opus 4.8 behave like Claude Fable 5 — doctrine output style, drift-catching hooks, and an eval loop against golden Fable transcripts. Claude Code plugin.
githubmeasured growthOpen source ↗
Install
/plugin marketplace add rennf93/opus-fable-playbook34GitHub stars+1 (+3.0 %) - 21Active
Log Claude Code sessions to Opik, the open-source LLM observability and evaluation platform, built by Comet. Tracing, evaluation, and skills for observable AI applications.
githubmeasured growthOpen source ↗
22GitHub stars+1 (+4.8 %) - 22Active
agent-stack
SkillProduction patterns for agent orchestrators, harnesses, evals, MCP/A2A interoperability, memory, provider routing, and LLM usage metering.
githubmeasured growthOpen source ↗
Install
git clone https://github.com/ssheleg/agent-stack ~/.claude/skills/agent-stack2GitHub stars+1 (+100.0 %) - 23Active
superlog
OtherOpen-source observability tool that uses AI agents to self-heal your software
githubmeasured growthOpen source ↗
1 404GitHub stars+1 (+0.07 %)
Learning resources
Ranked by measured growth, normalized across sources. These resources remain available separately and do not take part in the main tool ranking.
- 1ActiveOpen source ↗
50+ curated LLM observability tools PLUS 26 Agent Skills (several with runnable, unit-tested scripts) to build, evaluate, debug, secure & monitor reliable LLM apps. Tracing, evals, guardrails, LLMOps.
githubInstall
Resourcemeasured growthgit clone https://github.com/ContextJet-ai/awesome-llm-observability ~/.claude/skills/awesome-llm-observability33GitHub stars+3 (+10.0 %) - 2ActiveOpen source ↗
trigger_tree
SkillDocumentation-discovery telemetry for Claude Code — heat/cold maps, health grade, evidence-backed router fixes. 100% local, zero tokens. /tt
githubInstall
Resourcemeasured growth/plugin marketplace add Hedde/trigger_tree14GitHub stars+1 (+7.7 %)