LLM observability: tools for developers
Monitoring, traces, evaluation and quality tracking for LLM applications. An entry can appear under several use cases when its topics or description provide several explicit signals.
Use case
Activity
Sort by
75 entries in this view.
Tool ranking
Ranked by measured growth, normalized across sources. Raw values keep their source-specific window, shown under each tool; a change requires at least two snapshots from the last 7 days.
- 1Active
Query Spanlens LLM observability from Cursor, Claude Desktop, or Continue via MCP.
mcpmeasured growthOpen source ↗
Install
claude mcp add mcp-server -- npx @spanlens/mcp-server12GitHub starsstable - 2Active
promptfoo
OtherTest your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare performance of GPT, Claude, Gemini, DeepSeek, and more. Simple declarative configs with command line and CI/CD integration. Used by OpenAI…
githubmeasured growthOpen source ↗
24 768GitHub stars+142 (+0.58 %) - 3Active
Flawless
OtherAI SRE AgenticOps for Kubernetes and cloud infrastructure.
githubmeasured growthOpen source ↗
782GitHub stars+1 (+0.13 %) - 4Active
mlflow
AgentThe open source AI engineering platform for agents, LLMs, and ML models. MLflow enables teams of all sizes to debug, evaluate, monitor, and optimize production-quality AI applications while controlling costs and managing access to models…
githubmeasured growthOpen source ↗
27 783GitHub stars+82 (+0.30 %) - 5Active
prefect
LibraryPrefect is a workflow orchestration framework for building resilient data pipelines in Python.
githubmeasured growthOpen source ↗
23 766GitHub stars+66 (+0.28 %) - 622 509GitHub stars+40 (+0.18 %)
- 7Active
openobserve
OtherOpen source observability platform for logs, metrics, traces, frontend monitoring, pipelines and LLM observability. A sophisticated, simple and highly performant alternative to Datadog, Splunk, and Elasticsearch with 140x lower storage…
githubmeasured growthOpen source ↗
21 615GitHub stars+96 (+0.45 %) - 8Active
signoz
OtherSigNoz is an open-source, OpenTelemetry-native observability platform for your team and their AI agents. Get logs, metrics, and traces in one tool with features like APM, distributed tracing, log management, infra monitoring, etc. Combined…
githubmeasured growthOpen source ↗
32 003GitHub stars+50 (+0.16 %) - 9Active
mastra
OtherMastra is the modern TypeScript framework for AI-powered applications and agents.
githubmeasured growthOpen source ↗
27 658GitHub stars+128 (+0.46 %) - 10Active
netdata
OtherThe fastest path to AI-powered full stack observability, even for lean teams.
githubmeasured growthOpen source ↗
80 412GitHub stars+85 (+0.11 %) - 11Active
openstatus
MCP🫖 Status page with uptime monitoring & API monitoring as code 🫖
githubmeasured growthOpen source ↗
9 056GitHub stars+31 (+0.34 %) - 12Active
cilium
OthereBPF-based Networking, Security, and Observability
githubmeasured growthOpen source ↗
25 047GitHub stars+29 (+0.12 %) - 13Active
kubeshark
AgenteBPF-powered network observability for Kubernetes. Indexes L4/L7 traffic with full K8s context, decrypts TLS without keys. Queryable by AI agents via MCP and humans via dashboard.
githubmeasured growthOpen source ↗
12 068GitHub stars+8 (+0.07 %) - 14Active
databuff
AgentDataBuff is an AI-native APM built on Opentelemetry,with multi-agent troubleshooting out of the box.
githubmeasured growthOpen source ↗
642GitHub stars+32 (+5.2 %) - 15Active
VeriRun
OtherEvidence-first infrastructure for reproducible, isolated executable evaluation and online rewards.
githubmeasured growthOpen source ↗
140GitHub stars+24 (+20.7 %) - 16Active
agent-kernel
AgentThe Operating System for Scalable Enterprise AI Agents - Run, orchestrate, and deploy Compliant Enterprise AI Agents at scale across frameworks, without lock-in, rewrites or fragile glue code. Native support for MCP, A2A. Interface with all…
githubmeasured growthOpen source ↗
156GitHub stars+19 (+13.9 %) - 17Active
Production reliability, benchmarking, and evaluation layer for MCP agent runtimes — protocol negotiation, A2A interop, circuit breakers, connection pooling, adaptive streaming, and autoscaling, with measured throughput/latency and a…
githubmeasured growthOpen source ↗
3GitHub starsstable - 18Active
Architecture-first Python scaffold for an auditable multi-agent corporate credit desk using A2A, MCP, deterministic credit policies, model routing, and OpenTelemetry.
githubmeasured growthOpen source ↗
0GitHub starsstable - 19Active
aura
MCPAURA is a production-tested SRE agent platform you can deploy in minutes. AURA handles the guardrails, APIs, state management, streaming, and failure handling required to put AI to work safely on production infrastructure.
githubmeasured growthOpen source ↗
254GitHub stars-3 (-1.2 %) - 20Active
dsh-plugins
AgentGeneric DeepSeek Harness (dsh) plugins: A2A protocol server, session storage mirror, and Langfuse observability.
githubmeasured growthOpen source ↗
1GitHub starsstable - 21Active
nudge
MCPA typed, replayable, budget-aware programming language for LLM agents. Compiles to Python & TypeScript.
githubmeasured growthOpen source ↗
2GitHub starsstable - 22Active
agent-stack
SkillProduction patterns for agent orchestrators, harnesses, evals, MCP/A2A interoperability, memory, provider routing, and LLM usage metering.
githubmeasured growthOpen source ↗
Install
git clone https://github.com/ssheleg/agent-stack ~/.claude/skills/agent-stack2GitHub stars+1 (+100.0 %) - 23Active
Multi-agent quality gate skill for Claude Code that researches, reviews, tests and challenges AI-generated work before the final answer.
githubmeasured growthOpen source ↗
Install
git clone https://github.com/ma-nucho-pro/supervisor-skill-claude ~/.claude/skills/supervisor-skill-claude2GitHub starsstable - 24Active
langfuse-docs
Skill🪢 Langfuse documentation -- Langfuse is the open source LLM Engineering Platform. Observability, evals, prompt management, playground and metrics to debug and improve LLM apps
githubmeasured growthOpen source ↗
Install
git clone https://github.com/langfuse/langfuse-docs ~/.claude/skills/langfuse-docs239GitHub stars+3 (+1.3 %) - 25Active
boundary-bench
AgentDeterministic crash tests for agent control planes: paired worlds, hard stops, budgets, scopes, approvals, and hash-linked receipts.
githubmeasured growthOpen source ↗
1GitHub starsstable