LLM observability: tools for developers
Monitoring, traces, evaluation and quality tracking for LLM applications. An entry can appear under several use cases when its topics or description provide several explicit signals.
Use case
Activity
Sort by
74 entries in this view.
Tool ranking
Ranked by measured growth, normalized across sources. Raw values keep their source-specific window, shown under each tool; a change requires at least two snapshots from the last 7 days.
- 1Active
skill-eval-harness
SkillAgent Skill evaluation harness for paired variants, trace artifacts, and runner adapters
githubmeasured growthOpen source ↗
Install
git clone https://github.com/adewale/skill-eval-harness ~/.claude/skills/skill-eval-harness73GitHub stars+4 (+5.8 %) - 2Active
langfuse-docs
Skill🪢 Langfuse documentation -- Langfuse is the open source LLM Engineering Platform. Observability, evals, prompt management, playground and metrics to debug and improve LLM apps
githubmeasured growthOpen source ↗
Install
git clone https://github.com/langfuse/langfuse-docs ~/.claude/skills/langfuse-docs239GitHub stars+3 (+1.3 %) - 3Active
nora
MCPOpen-source, self-hosted control plane for OpenClaw and Hermes AI-agent fleets on Docker/Kubernetes — REST, CLI, and MCP.
githubmeasured growthOpen source ↗
50GitHub stars+2 (+4.2 %) - 4Active
Deterministic, local-only audit reports for Claude Code AI agent sessions. Rust + SQLite. Zero network calls. MCP server for agents.
githubmeasured growthOpen source ↗
22GitHub stars+2 (+10.0 %) - 5Active
claude-code-karma
SkillDashboard for monitoring claude code sessions.
githubmeasured growthOpen source ↗
Install
/plugin marketplace add JayantDevkar/claude-code-karma323GitHub stars+2 (+0.62 %) - 6Active
superlog
OtherOpen-source observability tool that uses AI agents to self-heal your software
githubmeasured growthOpen source ↗
1 404GitHub stars+2 (+0.14 %) - 7Active
Flawless
OtherAI SRE AgenticOps for Kubernetes and cloud infrastructure.
githubmeasured growthOpen source ↗
782GitHub stars+1 (+0.13 %) - 8Active
agent-stack
SkillProduction patterns for agent orchestrators, harnesses, evals, MCP/A2A interoperability, memory, provider routing, and LLM usage metering.
githubmeasured growthOpen source ↗
Install
git clone https://github.com/ssheleg/agent-stack ~/.claude/skills/agent-stack2GitHub stars+1 (+100.0 %) - 9Active
adlc-team-skills
Skill🐙 ADLC Team Skills — Agentic SDLC for Engineering Teams
githubmeasured growthOpen source ↗
Install
git clone https://github.com/tikalk/adlc-team-skills ~/.claude/skills/adlc-team-skills132GitHub stars+1 (+0.76 %) - 10105GitHub stars+1 (+0.96 %)
- 11Active
skill-kit
Skilllocal-first analytics for AI agent skills
githubmeasured growthOpen source ↗
Install
git clone https://github.com/crafter-station/skill-kit ~/.claude/skills/skill-kit77GitHub stars+1 (+1.3 %) - 12Active
claudestat
SkillReal-time execution trace and cost intelligence for Claude Code
githubmeasured growthOpen source ↗
Install
/plugin marketplace add DeibyGS/claudestat34GitHub stars+1 (+3.0 %) - 13Active
adl-cli
OtherA command-line tool to scaffold and manage enterprise-ready AI Agents powered by the A2A (Agent-to-Agent) protocol
githubmeasured growthOpen source ↗
14GitHub stars+1 (+7.7 %) - 14Active
Amazon Bedrock AgentCore enterprise platform accelerator with AWS CDK, Terraform organization guardrails, MCP/A2A agents, security, memory, and observability.
githubmeasured growthOpen source ↗
5GitHub stars+1 (+25.0 %) - 15Active
cap-evolve
MCPOptimize any AI agent’s skills, tools/MCP, and prompts against your own evals.
githubmeasured growthOpen source ↗
47GitHub stars+1 (+2.2 %) - 16Active
Production reliability, benchmarking, and evaluation layer for MCP agent runtimes — protocol negotiation, A2A interop, circuit breakers, connection pooling, adaptive streaming, and autoscaling, with measured throughput/latency and a…
githubmeasured growthOpen source ↗
3GitHub starsstable - 17Active
Architecture-first Python scaffold for an auditable multi-agent corporate credit desk using A2A, MCP, deterministic credit policies, model routing, and OpenTelemetry.
githubmeasured growthOpen source ↗
0GitHub starsstable - 18Active
dsh-plugins
AgentGeneric DeepSeek Harness (dsh) plugins: A2A protocol server, session storage mirror, and Langfuse observability.
githubmeasured growthOpen source ↗
1GitHub starsstable - 19Active
nudge
MCPA typed, replayable, budget-aware programming language for LLM agents. Compiles to Python & TypeScript.
githubmeasured growthOpen source ↗
2GitHub starsstable - 20Active
Multi-agent quality gate skill for Claude Code that researches, reviews, tests and challenges AI-generated work before the final answer.
githubmeasured growthOpen source ↗
Install
git clone https://github.com/ma-nucho-pro/supervisor-skill-claude ~/.claude/skills/supervisor-skill-claude2GitHub starsstable - 21Active
boundary-bench
AgentDeterministic crash tests for agent control planes: paired worlds, hard stops, budgets, scopes, approvals, and hash-linked receipts.
githubmeasured growthOpen source ↗
1GitHub starsstable - 22Active
DeepSeek-Infra
OtherLocal-first Agentic AI Infrastructure Platform with LLM Gateway, Agent DAG Runtime, MCP Tool Hub, A2A Agent Mesh, Local RAG, Tool Sandbox and Observability.
githubmeasured growthOpen source ↗
1GitHub starsstable - 23Active
CodeFlow
Agent面向长程研发任务的 Java Agent Harness,支持 A2A 跨语言协作、可恢复执行、上下文工程与 Eval 驱动开发。
githubmeasured growthOpen source ↗
0GitHub starsstable
Learning resources
Ranked by measured growth, normalized across sources. These resources remain available separately and do not take part in the main tool ranking.
- 1ActiveOpen source ↗
50+ curated LLM observability tools PLUS 26 Agent Skills (several with runnable, unit-tested scripts) to build, evaluate, debug, secure & monitor reliable LLM apps. Tracing, evals, guardrails, LLMOps.
githubInstall
Resourcemeasured growthgit clone https://github.com/ContextJet-ai/awesome-llm-observability ~/.claude/skills/awesome-llm-observability32GitHub stars+2 (+6.7 %) - 2ActiveOpen source ↗
trigger_tree
SkillDocumentation-discovery telemetry for Claude Code — heat/cold maps, health grade, evidence-backed router fixes. 100% local, zero tokens. /tt
githubInstall
Resourcemeasured growth/plugin marketplace add Hedde/trigger_tree14GitHub stars+1 (+7.7 %)