LLM observability: tools for developers
Monitoring, traces, evaluation and quality tracking for LLM applications. An entry can appear under several use cases when its topics or description provide several explicit signals.
Use case
Activity
Sort by
117 entries in this view.
Tool ranking
Ranked by measured growth, normalized across sources. Raw values keep their source-specific window, shown under each tool; a change requires at least two snapshots from the last 7 days.
- 1Active
AI Guardian
MCPGoverned local-LLM observability: model policy, prompt scanner, capture proxy, 21 tools.
mcpmeasured growthOpen source ↗
Install
claude mcp add ai-guardian -- uvx ai-guardian-aiops0GitHub starsstable - 2Active
evoagent-os
AgentLocal-first control plane for durable, governed agent teams: DAG orchestration, approvals, memory, signed skills, trace contracts, and realtime.
githubmeasured growthOpen source ↗
0GitHub starsstable - 3Active
historian
AgentA local A2A event and continuity service for agent applications, with structured history, provenance-safe transcripts, literal search, and natural-language queries.
githubmeasured growthOpen source ↗
1GitHub starsstable - 4Active
DriftSentinel
AgentAgent 降智检测与自愈公评网络 — an immune system for the AI agent society
githubmeasured growthOpen source ↗
1GitHub starsstable - 5Active
langfuse-mcp
MCPA Model Context Protocol (MCP) server for Langfuse, enabling AI agents to query Langfuse trace data for enhanced debugging and observability
githubmeasured growthOpen source ↗
105GitHub starsstable - 6Dormant
primal-core
AgentThe reliability and interoperability layer for AI agents. A2A-native. MCP-bridged.
githubmeasured growthOpen source ↗
1GitHub starsstable - 7Active
pyxen
AgentA lightweight Python library that decouples agentic runtime from applications it builds
githubmeasured growthOpen source ↗
1GitHub starsstable - 8Active
shokunin-review
OtherTerminal-first validation harness for reviewing PRDs, RFCs, strategy docs, and experiment plans.
githubmeasured growthOpen source ↗
1GitHub starsstable - 9Dormant
anti-lie
SkillDon't make LLMs honest. Make every factual claim auditable. — An LLM Claim Auditing Layer with T1-T7 truth gradients. 98.1% business effectiveness on LiarBench v0.2.
githubmeasured growthOpen source ↗
Install
git clone https://github.com/lc198707/anti-lie ~/.claude/skills/anti-lie89GitHub starsstable - 10Active
boundary-bench
AgentDeterministic crash tests for agent control planes: paired worlds, hard stops, budgets, scopes, approvals, and hash-linked receipts.
githubmeasured growthOpen source ↗
1GitHub starsstable - 11Active
headsup
SkillGlanceable Claude Code and Codex state in your terminal tabs: white=idle, blue=working, orange=waiting. Multi-terminal (iTerm2, WezTerm, AI Power Term), a live status bar with Anthropic usage limits, /sfl and /nil window save-and-restore,…
githubmeasured growthOpen source ↗
Install
git clone https://github.com/wasulajr/headsup ~/.claude/skills/headsup1GitHub starsstable - 12Active
dsh-plugins
AgentGeneric DeepSeek Harness (dsh) plugins: A2A protocol server, session storage mirror, and Langfuse observability.
githubmeasured growthOpen source ↗
1GitHub starsstable - 13Active
adherencia-reglas
SkillMide si tu agente de IA cumple las reglas que le escribiste. Lee el historial local de Claude Code y devuelve un porcentaje por regla. Sin instalar nada, sin red, biblioteca estándar.
githubmeasured growthOpen source ↗
Install
git clone https://github.com/jleonceo/adherencia-reglas ~/.claude/skills/adherencia-reglas1GitHub starsstable - 14Active
AgentX-Python
OtherAgentX python SDK. Build multi-agent AI workforce. Run evaluation. Trace your agent. Full Observability.
githubmeasured growthOpen source ↗
68GitHub starsstable - 15Active
a2a-otel-kit
MCPVendor-neutral OpenTelemetry tracing for A2A agents and MCP services, with W3C context propagation and privacy-safe telemetry.
githubmeasured growthOpen source ↗
1GitHub starsstable - 16Active
8 MCP SMB products — standalone AI servers for SMBs: Guardrails, FinOps, Observability, Router, Trust Score, Memory, ThinkSecure, A2A Lite. 37 tools, TNC credits billing.
githubmeasured growthOpen source ↗
1GitHub starsstable - 17Active
Skill instalable para Claude Code que mide si tu agente cumple las reglas establecidas. Lee el historial de sesiones de tu disco y devuelve un porcentaje por regla de cumplimiento. Sin instalar nada, sin red, biblioteca estándar.
githubmeasured growthOpen source ↗
Install
/plugin marketplace add jleonceo/skill-adherencia-reglas1GitHub starsstable - 18Active
claudestat
SkillReal-time execution trace and cost intelligence for Claude Code
githubmeasured growthOpen source ↗
Install
/plugin marketplace add DeibyGS/claudestat34GitHub starsstable - 19Active
tulip-agents
AgentThe agent framework where the model never holds the trigger — every consequential action clears your policy first, waits for a human when it matters, and lands on a record you can verify. Build on it, or put it around the agent you already…
githubmeasured growthOpen source ↗
1GitHub starsstable - 20Active
DeepSeek-Infra
OtherLocal-first Agentic AI Infrastructure Platform with LLM Gateway, Agent DAG Runtime, MCP Tool Hub, A2A Agent Mesh, Local RAG, Tool Sandbox and Observability.
githubmeasured growthOpen source ↗
1GitHub starsstable - 21Dormant
🔍 AI observability skill for Claude Code. Debug LangChain/LangGraph agents by fetching execution traces from LangSmith Studio directly in your terminal.
githubmeasured growthOpen source ↗
Install
git clone https://github.com/OthmanAdi/langsmith-fetch-skill ~/.claude/skills/langsmith-fetch-skill27GitHub starsstable - 22Active
agentgateway
MCPAgentGateway — independent third-party profile of a public API surface, by API Evangelist. AgentGateway is an open-source, AI-native proxy and gateway for routing, observing, and governing traffic to and from AI agents, LLM providers, and…
githubmeasured growthOpen source ↗
0GitHub starsstable - 23Dormant
astragraph
AgentPolicy-enforced observability and fail-closed guardrails for MCP/A2A multi-agent systems.
githubmeasured growthOpen source ↗
26GitHub starsstable - 24Active
skill-receipts
SkillClaude Code skills where every entry ships receipts — accuracy-gated benchmarks against baseline and placebo, rejects published
githubmeasured growthOpen source ↗
Install
git clone https://github.com/sjh9714/skill-receipts ~/.claude/skills/skill-receipts2GitHub starsstable - 25Active
Sentry instrumentation skill for system-behavior tracking
githubmeasured growthOpen source ↗
Install
git clone https://github.com/tortastudios/sentry-instrumentation ~/.claude/skills/sentry-instrumentation24GitHub starsstable