LLM observability: tools for developers
Monitoring, traces, evaluation and quality tracking for LLM applications. An entry can appear under several use cases when its topics or description provide several explicit signals.
Use case
Activity
Sort by
121 entries in this view.
Tool ranking
Ranked by measured growth, normalized across sources. Raw values keep their source-specific window, shown under each tool; a change requires at least two snapshots from the last 7 days.
- 1Active
skill-kit
Skilllocal-first analytics for AI agent skills
githubmeasured growthOpen source ↗
Install
git clone https://github.com/crafter-station/skill-kit ~/.claude/skills/skill-kit77GitHub stars+1 (+1.3 %) - 2Active
dsh-plugins
AgentGeneric DeepSeek Harness (dsh) plugins: A2A protocol server, session storage mirror, and Langfuse observability.
githubmeasured growthOpen source ↗
1GitHub starsstable - 3Active
skill-eval-harness
SkillAgent Skill evaluation harness for paired variants, trace artifacts, and runner adapters
githubmeasured growthOpen source ↗
Install
git clone https://github.com/adewale/skill-eval-harness ~/.claude/skills/skill-eval-harness73GitHub stars+4 (+5.8 %) - 4Active
adherencia-reglas
SkillMide si tu agente de IA cumple las reglas que le escribiste. Lee el historial local de Claude Code y devuelve un porcentaje por regla. Sin instalar nada, sin red, biblioteca estándar.
githubmeasured growthOpen source ↗
Install
git clone https://github.com/jleonceo/adherencia-reglas ~/.claude/skills/adherencia-reglas1GitHub starsstable - 5Active
AgentX-Python
OtherAgentX python SDK. Build multi-agent AI workforce. Run evaluation. Trace your agent. Full Observability.
githubmeasured growthOpen source ↗
68GitHub starsstable - 6Active
a2a-otel-kit
MCPVendor-neutral OpenTelemetry tracing for A2A agents and MCP services, with W3C context propagation and privacy-safe telemetry.
githubmeasured growthOpen source ↗
1GitHub starsstable - 7Active
seo-skill-bench
SkillOpen benchmark for Claude Code SEO skills — real headless execution against fixture sites with planted-defect answer keys. Deterministic scoring, pre-registered rubric.
githubmeasured growthOpen source ↗
Install
git clone https://github.com/aleclindz/seo-skill-bench ~/.claude/skills/seo-skill-bench51GitHub stars-5 (-8.9 %) - 8Dormant
An AI-powered multi-agent system that demonstrates clinical triage, OTC medication recommendations, and e-pharmacy integration for respiratory conditions. Built with modular agents that collaborate to provide safe, intelligent healthcare…
githubmeasured growthOpen source ↗
0GitHub starsstable - 9Active
8 MCP SMB products — standalone AI servers for SMBs: Guardrails, FinOps, Observability, Router, Trust Score, Memory, ThinkSecure, A2A Lite. 37 tools, TNC credits billing.
githubmeasured growthOpen source ↗
1GitHub starsstable - 10Active
nora
MCPOpen-source, self-hosted control plane for OpenClaw and Hermes AI-agent fleets on Docker/Kubernetes — REST, CLI, and MCP.
githubmeasured growthOpen source ↗
50GitHub stars+2 (+4.2 %) - 11Active
SkillForge
SkillA skill creator that proves its skills work. Evidence-driven skill creation for Claude Code and Codex: baseline-tested generation, per-skill regression evals, ecosystem doctor, cross-runtime compile, and an opt-in proactive advisor.
githubmeasured growthOpen source ↗
Install
git clone https://github.com/tripleyak/SkillForge ~/.claude/skills/SkillForge884GitHub stars-1 (-0.11 %) - 12Active
Skill instalable para Claude Code que mide si tu agente cumple las reglas establecidas. Lee el historial de sesiones de tu disco y devuelve un porcentaje por regla de cumplimiento. Sin instalar nada, sin red, biblioteca estándar.
githubmeasured growthOpen source ↗
Install
/plugin marketplace add jleonceo/skill-adherencia-reglas1GitHub starsstable - 13Active
claudestat
SkillReal-time execution trace and cost intelligence for Claude Code
githubmeasured growthOpen source ↗
Install
/plugin marketplace add DeibyGS/claudestat34GitHub starsstable - 14Active
tulip-agents
AgentThe agent framework where the model never holds the trigger — every consequential action clears your policy first, waits for a human when it matters, and lands on a record you can verify. Build on it, or put it around the agent you already…
githubmeasured growthOpen source ↗
1GitHub starsstable - 15Active
opus-fable-playbook
SkillMake Claude Opus 4.8 behave like Claude Fable 5 — doctrine output style, drift-catching hooks, and an eval loop against golden Fable transcripts. Claude Code plugin.
githubmeasured growthOpen source ↗
Install
/plugin marketplace add rennf93/opus-fable-playbook34GitHub stars+1 (+3.0 %) - 16Active
DeepSeek-Infra
OtherLocal-first Agentic AI Infrastructure Platform with LLM Gateway, Agent DAG Runtime, MCP Tool Hub, A2A Agent Mesh, Local RAG, Tool Sandbox and Observability.
githubmeasured growthOpen source ↗
1GitHub starsstable - 17Dormant
🔍 AI observability skill for Claude Code. Debug LangChain/LangGraph agents by fetching execution traces from LangSmith Studio directly in your terminal.
githubmeasured growthOpen source ↗
Install
git clone https://github.com/OthmanAdi/langsmith-fetch-skill ~/.claude/skills/langsmith-fetch-skill27GitHub starsstable - 18Active
agentgateway
MCPAgentGateway — independent third-party profile of a public API surface, by API Evangelist. AgentGateway is an open-source, AI-native proxy and gateway for routing, observing, and governing traffic to and from AI agents, LLM providers, and…
githubmeasured growthOpen source ↗
0GitHub starsstable - 19Dormant
gt8004-sdk
OtherOfficial Python SDK for GT8004 — AI agent observability with MCP, A2A, x402 payment tracking. FastAPI, Flask, FastMCP middleware included.
githubmeasured growthOpen source ↗
1GitHub starsstable - 20Dormant
astragraph
AgentPolicy-enforced observability and fail-closed guardrails for MCP/A2A multi-agent systems.
githubmeasured growthOpen source ↗
26GitHub starsstable - 21Active
OpenJudge
SkillOpenJudge: A Unified Framework for Holistic Evaluation and Quality Rewards
githubmeasured growthOpen source ↗
Install
git clone https://github.com/agentscope-ai/OpenJudge ~/.claude/skills/OpenJudge812GitHub stars+7 (+0.87 %) - 22Active
skill-receipts
SkillClaude Code skills where every entry ships receipts — accuracy-gated benchmarks against baseline and placebo, rejects published
githubmeasured growthOpen source ↗
Install
git clone https://github.com/sjh9714/skill-receipts ~/.claude/skills/skill-receipts2GitHub starsstable - 23Active
Sentry instrumentation skill for system-behavior tracking
githubmeasured growthOpen source ↗
Install
git clone https://github.com/tortastudios/sentry-instrumentation ~/.claude/skills/sentry-instrumentation24GitHub starsstable - 24Dormant
CustoFlow
OtherMulti-agent customer support system with Google ADK & Gemini 2.5 Flash Lite. Kaggle capstone demonstrating 11+ concepts. Automates 80%+ queries, <10s response time.
githubmeasured growthOpen source ↗
2GitHub starsstable - 25Active
Deterministic, local-only audit reports for Claude Code AI agent sessions. Rust + SQLite. Zero network calls. MCP server for agents.
githubmeasured growthOpen source ↗
22GitHub stars+2 (+10.0 %)