LLM observability: tools for developers
Monitoring, traces, evaluation and quality tracking for LLM applications. An entry can appear under several use cases when its topics or description provide several explicit signals.
Use case
Activity
Sort by
75 entries in this view.
Tool ranking
Ranked by creation date, newest first; undated entries come last. Raw values keep their source-specific window, shown under each tool; a change requires at least two snapshots from the last 7 days.
- 1Active
untell
OtherAI-detector auditing toolkit: measures how often a detector flags genuine human writing, how stable its verdict is across seeds, and whether that verdict survives meaning-preserving edits. Detector-in-the-loop measurement harness. Claude…
githubmeasured growthOpen source ↗
18GitHub starsstable - 2Active
SkillEvaluator
SkillMulti-tier framework for evaluating AI agent skills with quality gates, semantic overlap detection, synthetic evaluation dataset generation, and live agent evaluation that measures how skills affect agent behavior.
githubmeasured growthOpen source ↗
Install
git clone https://github.com/NVIDIA/SkillEvaluator ~/.claude/skills/SkillEvaluator389GitHub stars+57 (+17.2 %) - 3Active
databuff
AgentDataBuff is an AI-native APM built on Opentelemetry,with multi-agent troubleshooting out of the box.
githubmeasured growthOpen source ↗
642GitHub stars+32 (+5.2 %) - 4Active
cap-evolve
MCPOptimize any AI agent’s skills, tools/MCP, and prompts against your own evals.
githubmeasured growthOpen source ↗
47GitHub stars+1 (+2.2 %) - 5Active
skill-eval-harness
SkillAgent Skill evaluation harness for paired variants, trace artifacts, and runner adapters
githubmeasured growthOpen source ↗
Install
git clone https://github.com/adewale/skill-eval-harness ~/.claude/skills/skill-eval-harness73GitHub stars+4 (+5.8 %) - 6Active
tulip-agents
AgentThe agent framework where the model never holds the trigger — every consequential action clears your policy first, waits for a human when it matters, and lands on a record you can verify. Build on it, or put it around the agent you already…
githubmeasured growthOpen source ↗
1GitHub starsstable - 7Active
axiom
SkillAxiom is a curated marketplace of shared plugins for Claude Code and Codex.
githubmeasured growthOpen source ↗
Install
git clone https://github.com/netopsengineer/axiom ~/.claude/skills/axiom5GitHub starsstable - 8Active
DeepSeek-Infra
OtherLocal-first Agentic AI Infrastructure Platform with LLM Gateway, Agent DAG Runtime, MCP Tool Hub, A2A Agent Mesh, Local RAG, Tool Sandbox and Observability.
githubmeasured growthOpen source ↗
1GitHub starsstable - 9Active
anchor
SkillAnchor — the production-grade AGENTS.md template for AI coding agents. 51 sections of battle-tested rules. Keep your agents grounded.
githubmeasured growthOpen source ↗
Install
git clone https://github.com/peva3/anchor ~/.claude/skills/anchor8GitHub starsstable - 10Active
superlog
OtherOpen-source observability tool that uses AI agents to self-heal your software
githubmeasured growthOpen source ↗
1 404GitHub stars+2 (+0.14 %) - 11Active
headsup
SkillGlanceable Claude Code and Codex state in your terminal tabs: white=idle, blue=working, orange=waiting. Multi-terminal (iTerm2, WezTerm, AI Power Term), a live status bar with Anthropic usage limits, /sfl and /nil window save-and-restore,…
githubmeasured growthOpen source ↗
Install
git clone https://github.com/wasulajr/headsup ~/.claude/skills/headsup1GitHub starsstable - 12Active
74 open-source Agent Skills for Claude Code and Codex: AI SEO, AEO and GEO, code review with an A-F ship grade, CI gates, AI evals, design systems, conversion copy, Instagram growth, iOS and Android app shipping, creator rights, and…
githubmeasured growthOpen source ↗
129GitHub stars-4 (-3.0 %) - 13Active
galdor
OtherA Go-native framework for LLM agents, with OpenTelemetry observability built in.
githubmeasured growthOpen source ↗
10GitHub starsstable - 14Active
agent-skills-eval
SkillA test runner for agentskills.io-style AI agent skills
githubmeasured growthOpen source ↗
Install
git clone https://github.com/darkrishabh/agent-skills-eval ~/.claude/skills/agent-skills-eval719GitHub stars+13 (+1.8 %) - 15Active
claudestat
SkillReal-time execution trace and cost intelligence for Claude Code
githubmeasured growthOpen source ↗
Install
/plugin marketplace add DeibyGS/claudestat34GitHub stars+1 (+3.0 %) - 16Active
Query Spanlens LLM observability from Cursor, Claude Desktop, or Continue via MCP.
mcpmeasured growthOpen source ↗
Install
claude mcp add mcp-server -- npx @spanlens/mcp-server12GitHub starsstable - 17Active
memroos
AgentMemroOS / memroos: memory OS and governance layer for AI agents, agent workflows, dispatch, proof, and context continuity.
githubmeasured growthOpen source ↗
7GitHub starsstable - 18Active
yao-meta-skill
SkillYAO = Yielding AI Outcomes. A rigorous engineering, evaluation, governance, and portability system for reusable agent skills.
githubmeasured growthOpen source ↗
Install
git clone https://github.com/yaojingang/yao-meta-skill ~/.claude/skills/yao-meta-skill2 577GitHub stars+35 (+1.4 %) - 19Active
dynatrace-for-ai
SkillSkills, prompts, and instructions for building AI agents on top of Dynatrace production context
githubmeasured growthOpen source ↗
Install
git clone https://github.com/Dynatrace/dynatrace-for-ai ~/.claude/skills/dynatrace-for-ai132GitHub stars+4 (+3.1 %) - 20Active
agentgateway
MCPAgentGateway — independent third-party profile of a public API surface, by API Evangelist. AgentGateway is an open-source, AI-native proxy and gateway for routing, observing, and governing traffic to and from AI agents, LLM providers, and…
githubmeasured growthOpen source ↗
0GitHub starsstable - 21Active
AgentStack
AgentProvide clear documentation for AgentStack’s MCP protocol, plugins, and ecosystem API with usage examples and tool references.
githubmeasured growthOpen source ↗
0GitHub starsstable - 22Active
nora
MCPOpen-source, self-hosted control plane for OpenClaw and Hermes AI-agent fleets on Docker/Kubernetes — REST, CLI, and MCP.
githubmeasured growthOpen source ↗
50GitHub stars+2 (+4.2 %) - 23Active
Vendor-neutral OpenTelemetry skills for AI coding agents, grounded in upstream sources
githubmeasured growthOpen source ↗
Install
git clone https://github.com/ollygarden/opentelemetry-agent-skills ~/.claude/skills/opentelemetry-agent-skills97GitHub stars+4 (+4.3 %) - 24Active
arize-skills
SkillAgent skills for Arize — datasets, experiments, and traces via the ax CLI
githubmeasured growthOpen source ↗
Install
git clone https://github.com/Arize-ai/arize-skills ~/.claude/skills/arize-skills47GitHub starsstable
Learning resources
Ranked by creation date, newest first; undated entries come last. These resources remain available separately and do not take part in the main tool ranking.
- 1ActiveOpen source ↗
Curated + scored map of official MCP servers and agents for DevOps, Cloud, SRE, and Platform Engineering — every entry rated on production access, approval gates, and audit evidence.
githubResourcemeasured growth77GitHub starsstable