LLM observability: tools for developers
Monitoring, traces, evaluation and quality tracking for LLM applications. An entry can appear under several use cases when its topics or description provide several explicit signals.
Use case
Activity
Sort by
122 entries in this view.
Tool ranking
Ranked by known GitHub stars, highest first; archived repositories come after maintained repositories. Raw values keep their source-specific window, shown under each tool; a change requires at least two snapshots from the last 7 days.
- 1Dormant
adaptive-harness
SkillA self-improving harness router for Claude Code.
githubmeasured growthOpen source ↗
Install
/plugin marketplace add SeongwoongCho/adaptive-harness8GitHub starsstable - 2Active
anchor
SkillAnchor — the production-grade AGENTS.md template for AI coding agents. 51 sections of battle-tested rules. Keep your agents grounded.
githubmeasured growthOpen source ↗
Install
git clone https://github.com/peva3/anchor ~/.claude/skills/anchor8GitHub starsstable - 3Active
memroos
AgentMemroOS / memroos: memory OS and governance layer for AI agents, agent workflows, dispatch, proof, and context continuity.
githubmeasured growthOpen source ↗
7GitHub starsstable - 4Active
Amazon Bedrock AgentCore enterprise platform accelerator with AWS CDK, Terraform organization guardrails, MCP/A2A agents, security, memory, and observability.
githubmeasured growthOpen source ↗
6GitHub stars+1 (+20.0 %) - 5Dormant
kiboserve
MCPOpen-source Python framework to deploy AI agents via HTTP, A2A, and MCP with built-in observability
githubmeasured growthOpen source ↗
5GitHub starsstable - 6Active
axiom
SkillAxiom is a curated marketplace of shared plugins for Claude Code and Codex.
githubmeasured growthOpen source ↗
Install
git clone https://github.com/netopsengineer/axiom ~/.claude/skills/axiom5GitHub starsstable - 7Active
agent-mmm
SkillMarketing Mix Model expert agent plugin for Claude Code - pymc-marketing v0.18.2+
githubmeasured growthOpen source ↗
Install
/plugin marketplace add Yakoub-ai/agent-mmm4GitHub starsstable - 8Dormant
forge-skills
SkillAn assembly line for AI software development. 35 skills, 11 agent personas, 29 commands. From raw idea to shipped code.
githubmeasured growthOpen source ↗
Install
/plugin marketplace add aneja5/forge-skills3GitHub starsstable - 9Active
sre-on-call
AgentMulti-agent SRE on-call investigator that auto-triages Slack/Discord infrastructure alerts via AWS Bedrock AgentCore, fanning out to specialized agents (CloudWatch, EKS, Slack/Discord scanners) for parallel investigation.
githubmeasured growthOpen source ↗
3GitHub starsstable - 10Dormant
agenttap
AgentReal-time debugging proxy for Agent2Agent (A2A) multi-agent systems
githubmeasured growthOpen source ↗
3GitHub starsstable - 11Active
Production reliability, benchmarking, and evaluation layer for MCP agent runtimes — protocol negotiation, A2A interop, circuit breakers, connection pooling, adaptive streaming, and autoscaling, with measured throughput/latency and a…
githubmeasured growthOpen source ↗
3GitHub starsstable - 12Active
skill-receipts
SkillClaude Code skills where every entry ships receipts — accuracy-gated benchmarks against baseline and placebo, rejects published
githubmeasured growthOpen source ↗
Install
git clone https://github.com/sjh9714/skill-receipts ~/.claude/skills/skill-receipts2GitHub starsstable - 13Dormant
CustoFlow
OtherMulti-agent customer support system with Google ADK & Gemini 2.5 Flash Lite. Kaggle capstone demonstrating 11+ concepts. Automates 80%+ queries, <10s response time.
githubmeasured growthOpen source ↗
2GitHub starsstable - 14Active
Multi-agent quality gate skill for Claude Code that researches, reviews, tests and challenges AI-generated work before the final answer.
githubmeasured growthOpen source ↗
Install
git clone https://github.com/ma-nucho-pro/supervisor-skill-claude ~/.claude/skills/supervisor-skill-claude2GitHub starsstable - 15Dormant
Agents-eval
AgentA Multi-Agent System (MAS) evaluation framework using PydanticAI that generates and evaluates scientific paper reviews through a three-tiered assessment approach: traditional metrics, LLM-as-a-Judge, and graph-based complexity analysis.
githubmeasured growthOpen source ↗
2GitHub starsstable - 16Active
agent-stack
SkillProduction patterns for agent orchestrators, harnesses, evals, MCP/A2A interoperability, memory, provider routing, and LLM usage metering.
githubmeasured growthOpen source ↗
Install
git clone https://github.com/ssheleg/agent-stack ~/.claude/skills/agent-stack2GitHub stars+1 (+100.0 %) - 17Active
nudge
MCPA typed, replayable, budget-aware programming language for LLM agents. Compiles to Python & TypeScript.
githubmeasured growthOpen source ↗
2GitHub starsstable - 18Active
Description Evidence-driven evaluation, benchmarking, and verified repair for Agent Skills across Codex, Claude Code, Gemini CLI, and Antigravity.
githubestimated momentumOpen source ↗
Install
git clone https://github.com/MaxLaurieHutchinson/skill-evaluation-graph ~/.claude/skills/skill-evaluation-graph2GitHub stars— - 19Active
verdict-contract
SkillYour LLM reviewer said APPROVE. Did it? A structured verdict contract: prompt rule + parser + exit-code gate in one stdlib file, with the 42 counterexamples that forced every line. Free, MIT.
githubestimated momentumOpen source ↗
Install
git clone https://github.com/tonydzi/verdict-contract ~/.claude/skills/verdict-contract2GitHub stars— - 20Active
8 MCP SMB products — standalone AI servers for SMBs: Guardrails, FinOps, Observability, Router, Trust Score, Memory, ThinkSecure, A2A Lite. 37 tools, TNC credits billing.
githubmeasured growthOpen source ↗
1GitHub starsstable - 21Dormant
primal-core
AgentThe reliability and interoperability layer for AI agents. A2A-native. MCP-bridged.
githubmeasured growthOpen source ↗
1GitHub starsstable - 22Active
a2a-otel-kit
MCPVendor-neutral OpenTelemetry tracing for A2A agents and MCP services, with W3C context propagation and privacy-safe telemetry.
githubmeasured growthOpen source ↗
1GitHub starsstable - 23Active
tulip-agents
AgentThe agent framework where the model never holds the trigger — every consequential action clears your policy first, waits for a human when it matters, and lands on a record you can verify. Build on it, or put it around the agent you already…
githubmeasured growthOpen source ↗
1GitHub starsstable - 24Active
pyxen
AgentA lightweight Python library that decouples agentic runtime from applications it builds
githubmeasured growthOpen source ↗
1GitHub starsstable
Learning resources
Ranked by known GitHub stars, highest first; archived repositories come after maintained repositories. These resources remain available separately and do not take part in the main tool ranking.
- 1ActiveOpen source ↗
tunelab
SkillClaude Code plugin for LLM fine-tuning, distillation, and evaluation — decide whether you need fine-tuning at all, distill your LLM logs into small local models (MLX/LoRA), evaluate with held-out discipline, and learn the why at every step.
githubInstall
Resourcemeasured growth/plugin marketplace add rchaz/tunelab6GitHub starsstable