LLM observability: tools for developers
Monitoring, traces, evaluation and quality tracking for LLM applications. An entry can appear under several use cases when its topics or description provide several explicit signals.
Use case
Activity
Sort by
75 entries in this view.
Tool ranking
Ranked by measured growth, normalized across sources. Raw values keep their source-specific window, shown under each tool; a change requires at least two snapshots from the last 7 days.
- 1Active
galdor
OtherA Go-native framework for LLM agents, with OpenTelemetry observability built in.
githubmeasured growthOpen source ↗
10GitHub starsstable - 2Active
Architecture-first Python scaffold for an auditable multi-agent corporate credit desk using A2A, MCP, deterministic credit policies, model routing, and OpenTelemetry.
githubmeasured growthOpen source ↗
0GitHub starsstable - 3Active
agent-mmm
SkillMarketing Mix Model expert agent plugin for Claude Code - pymc-marketing v0.18.2+
githubmeasured growthOpen source ↗
Install
/plugin marketplace add Yakoub-ai/agent-mmm4GitHub starsstable - 4Active
idun-agent-platform
Skill🟪 Open-source runtime that ships any LangGraph or Google ADK agent as a production-ready FastAPI service. Bundled , AG-UI copilotkit API, chat UI, 15+ guardrails, MCP, OpenTelemetry, OIDC. One pip install. Self-hosted, no vendor lock-in.
githubmeasured growthOpen source ↗
Install
git clone https://github.com/Idun-Group/idun-agent-platform ~/.claude/skills/idun-agent-platform199GitHub starsstable - 5Active
AI Guardian
MCPGoverned local-LLM observability: model policy, prompt scanner, capture proxy, 21 tools.
mcpmeasured growthOpen source ↗
Install
claude mcp add ai-guardian -- uvx ai-guardian-aiops0GitHub starsstable - 6Active
evoagent-os
AgentLocal-first control plane for durable, governed agent teams: DAG orchestration, approvals, memory, signed skills, trace contracts, and realtime.
githubmeasured growthOpen source ↗
0GitHub starsstable - 7Active
langfuse-mcp
MCPA Model Context Protocol (MCP) server for Langfuse, enabling AI agents to query Langfuse trace data for enhanced debugging and observability
githubmeasured growthOpen source ↗
105GitHub starsstable - 8Active
boundary-bench
AgentDeterministic crash tests for agent control planes: paired worlds, hard stops, budgets, scopes, approvals, and hash-linked receipts.
githubmeasured growthOpen source ↗
1GitHub starsstable - 9Active
headsup
SkillGlanceable Claude Code and Codex state in your terminal tabs: white=idle, blue=working, orange=waiting. Multi-terminal (iTerm2, WezTerm, AI Power Term), a live status bar with Anthropic usage limits, /sfl and /nil window save-and-restore,…
githubmeasured growthOpen source ↗
Install
git clone https://github.com/wasulajr/headsup ~/.claude/skills/headsup1GitHub starsstable - 10Active
dsh-plugins
AgentGeneric DeepSeek Harness (dsh) plugins: A2A protocol server, session storage mirror, and Langfuse observability.
githubmeasured growthOpen source ↗
1GitHub starsstable - 11Active
AgentX-Python
OtherAgentX python SDK. Build multi-agent AI workforce. Run evaluation. Trace your agent. Full Observability.
githubmeasured growthOpen source ↗
68GitHub starsstable - 12Active
a2a-otel-kit
MCPVendor-neutral OpenTelemetry tracing for A2A agents and MCP services, with W3C context propagation and privacy-safe telemetry.
githubmeasured growthOpen source ↗
1GitHub starsstable - 13Active
claudestat
SkillReal-time execution trace and cost intelligence for Claude Code
githubmeasured growthOpen source ↗
Install
/plugin marketplace add DeibyGS/claudestat34GitHub starsstable - 14Active
tulip-agents
AgentThe agent framework where the model never holds the trigger — every consequential action clears your policy first, waits for a human when it matters, and lands on a record you can verify. Build on it, or put it around the agent you already…
githubmeasured growthOpen source ↗
1GitHub starsstable - 15Active
DeepSeek-Infra
OtherLocal-first Agentic AI Infrastructure Platform with LLM Gateway, Agent DAG Runtime, MCP Tool Hub, A2A Agent Mesh, Local RAG, Tool Sandbox and Observability.
githubmeasured growthOpen source ↗
1GitHub starsstable - 16Active
agentgateway
MCPAgentGateway — independent third-party profile of a public API surface, by API Evangelist. AgentGateway is an open-source, AI-native proxy and gateway for routing, observing, and governing traffic to and from AI agents, LLM providers, and…
githubmeasured growthOpen source ↗
0GitHub starsstable - 17Active
nudge
MCPA typed, replayable, budget-aware programming language for LLM agents. Compiles to Python & TypeScript.
githubmeasured growthOpen source ↗
2GitHub starsstable - 18Active
untell
OtherAI-detector auditing toolkit: measures how often a detector flags genuine human writing, how stable its verdict is across seeds, and whether that verdict survives meaning-preserving edits. Detector-in-the-loop measurement harness. Claude…
githubmeasured growthOpen source ↗
18GitHub starsstable - 19Active
Multi-agent quality gate skill for Claude Code that researches, reviews, tests and challenges AI-generated work before the final answer.
githubmeasured growthOpen source ↗
Install
git clone https://github.com/ma-nucho-pro/supervisor-skill-claude ~/.claude/skills/supervisor-skill-claude2GitHub starsstable - 20Active
rashomon
SkillMeasure prompt and skill improvements with blind A/B comparison.
githubmeasured growthOpen source ↗
Install
/plugin marketplace add shinpr/rashomon18GitHub starsstable - 21Active
arize-skills
SkillAgent skills for Arize — datasets, experiments, and traces via the ax CLI
githubmeasured growthOpen source ↗
Install
git clone https://github.com/Arize-ai/arize-skills ~/.claude/skills/arize-skills47GitHub starsstable - 22Active
74 open-source Agent Skills for Claude Code and Codex: AI SEO, AEO and GEO, code review with an A-F ship grade, CI gates, AI evals, design systems, conversion copy, Instagram growth, iOS and Android app shipping, creator rights, and…
githubmeasured growthOpen source ↗
129GitHub stars-4 (-3.0 %) - 23Active
seo-skill-bench
SkillOpen benchmark for Claude Code SEO skills — real headless execution against fixture sites with planted-defect answer keys. Deterministic scoring, pre-registered rubric.
githubmeasured growthOpen source ↗
Install
git clone https://github.com/aleclindz/seo-skill-bench ~/.claude/skills/seo-skill-bench51GitHub stars-5 (-8.9 %)
Learning resources
Ranked by measured growth, normalized across sources. These resources remain available separately and do not take part in the main tool ranking.
- 1ActiveOpen source ↗
Curated + scored map of official MCP servers and agents for DevOps, Cloud, SRE, and Platform Engineering — every entry rated on production access, approval gates, and audit evidence.
githubResourcemeasured growth77GitHub starsstable - 2ActiveOpen source ↗
Agentic_AI_Engineer
AgentMy complete journey to becoming an Agentic AI Engineer through structured learning, projects, experiments, and production-ready implementations of modern AI systems.
githubResourcemeasured growth18GitHub starsstable