LLM observability: tools for developers
Monitoring, traces, evaluation and quality tracking for LLM applications. An entry can appear under several use cases when its topics or description provide several explicit signals.
Use case
Activity
Sort by
122 entries in this view.
Tool ranking
Mixed ranking: measured growth takes priority; entries without two snapshots are estimated. Raw values keep their source-specific window, shown under each tool; a change requires at least two snapshots from the last 7 days.
- 1Active
simple-output-styles
SkillMake Claude write clearly, for everyone.
githubmeasured growthOpen source ↗
Install
/plugin marketplace add stefanobaghino/simple-output-styles16GitHub starsstable - 2Active
agent-mmm
SkillMarketing Mix Model expert agent plugin for Claude Code - pymc-marketing v0.18.2+
githubmeasured growthOpen source ↗
Install
/plugin marketplace add Yakoub-ai/agent-mmm4GitHub starsstable - 3Active
tulip-agents
AgentThe agent framework where the model never holds the trigger — every consequential action clears your policy first, waits for a human when it matters, and lands on a record you can verify. Build on it, or put it around the agent you already…
githubmeasured growthOpen source ↗
1GitHub starsstable - 4Dormant
Agents-eval
AgentA Multi-Agent System (MAS) evaluation framework using PydanticAI that generates and evaluates scientific paper reviews through a three-tiered assessment approach: traditional metrics, LLM-as-a-Judge, and graph-based complexity analysis.
githubmeasured growthOpen source ↗
2GitHub starsstable - 5Dormant
MCP as a Judge: a behavioral MCP that strengthens AI coding assistants via explicit LLM evaluations
mcpmeasured growthOpen source ↗
Install
claude mcp add mcp-as-a-judge -- uvx mcp-as-a-judge17GitHub starsstable - 6Dormant
An AI-powered multi-agent system that demonstrates clinical triage, OTC medication recommendations, and e-pharmacy integration for respiratory conditions. Built with modular agents that collaborate to provide safe, intelligent healthcare…
githubmeasured growthOpen source ↗
0GitHub starsstable - 7Active
langfuse-mcp
MCPA Model Context Protocol (MCP) server for Langfuse, enabling AI agents to query Langfuse trace data for enhanced debugging and observability
githubmeasured growthOpen source ↗
105GitHub starsstable - 8Dormant
CustoFlow
OtherMulti-agent customer support system with Google ADK & Gemini 2.5 Flash Lite. Kaggle capstone demonstrating 11+ concepts. Automates 80%+ queries, <10s response time.
githubmeasured growthOpen source ↗
2GitHub starsstable - 9Active
Architecture-first Python scaffold for an auditable multi-agent corporate credit desk using A2A, MCP, deterministic credit policies, model routing, and OpenTelemetry.
githubmeasured growthOpen source ↗
0GitHub starsstable - 10Active
kubesphere
SkillThe container platform tailored for Kubernetes multi-cloud, datacenter, and edge management ⎈ 🖥 ☁️
githubmeasured growthOpen source ↗
Install
git clone https://github.com/kubesphere/kubesphere ~/.claude/skills/kubesphere17 035GitHub starsstable - 11Dormant
Two Claude Code skills for adversarial code review (single-model PAR and multi-model MMAR with cross-critique to catch hallucinations), plus a fixture-based eval suite.
githubmeasured growthOpen source ↗
Install
/plugin marketplace add prime-radiant-inc/parallel-adversarial-review17GitHub starsstable - 12Active
Build tested agent skills and govern their lifecycle through a user-defined marketplace: evidence, discovery, updates, rollback, quarantine, and 17-platform distribution.
githubmeasured growthOpen source ↗
Install
git clone https://github.com/FrancyJGLisboa/agent-skills-platform ~/.claude/skills/agent-skills-platform2 376GitHub starsstable - 13Active
sc-prism-releases
SkillRun many AIs on one board and keep control of all of it. Deterministic code decides who acts — never a model. A privacy floor keeps sensitive work on your machine, your own tests decide what counts as done, and every action lands on a…
githubmeasured growthOpen source ↗
Install
git clone https://github.com/sandhusukhdeep2/sc-prism-releases ~/.claude/skills/sc-prism-releases1GitHub starsstable - 14Active
headsup
SkillGlanceable Claude Code and Codex state in your terminal tabs: white=idle, blue=working, orange=waiting. Multi-terminal (iTerm2, WezTerm, AI Power Term), a live status bar with Anthropic usage limits, /sfl and /nil window save-and-restore,…
githubmeasured growthOpen source ↗
Install
git clone https://github.com/wasulajr/headsup ~/.claude/skills/headsup1GitHub starsstable - 15Active
rocketplaneIO
OtherSelf-hosted AI SRE for Kubernetes — zero-instrumentation eBPF observability plus a copilot that fixes issues through guardrailed, self-verifying actions. BYO-LLM, air-gapped capable.
githubmeasured growthOpen source ↗
131GitHub stars-4 (-3.0 %) - 16Active
74 open-source Agent Skills for Claude Code and Codex: AI SEO, AEO and GEO, code review with an A-F ship grade, CI gates, AI evals, design systems, conversion copy, Instagram growth, iOS and Android app shipping, creator rights, and…
githubmeasured growthOpen source ↗
129GitHub stars-4 (-3.0 %) - 17Active
superlog
OtherOpen-source observability tool that uses AI agents to self-heal your software
githubestimated momentumOpen source ↗
1 404GitHub stars—
Learning resources
Ranked by measured growth, normalized across sources. These resources remain available separately and do not take part in the main tool ranking.
- 1ActiveOpen source ↗
Agentic_AI_Engineer
AgentMy complete journey to becoming an Agentic AI Engineer through structured learning, projects, experiments, and production-ready implementations of modern AI systems.
githubResourcemeasured growth18GitHub starsstable - 2ActiveOpen source ↗
Curated + scored map of official MCP servers and agents for DevOps, Cloud, SRE, and Platform Engineering — every entry rated on production access, approval gates, and audit evidence.
githubResourcemeasured growth77GitHub starsstable - 3ActiveOpen source ↗
Handbook técnico aberto sobre engenharia de IA em produção: ML tradicional, LLMs, RAG, agentes, segurança, observabilidade, FinOps e deployment.
githubResourcemeasured growth1GitHub starsstable - 4ActiveOpen source ↗
trigger_tree
SkillDocumentation-discovery telemetry for Claude Code — heat/cold maps, health grade, evidence-backed router fixes. 100% local, zero tokens. /tt
githubInstall
Resourcemeasured growth/plugin marketplace add Hedde/trigger_tree14GitHub starsstable - 5ActiveOpen source ↗
tunelab
SkillClaude Code plugin for LLM fine-tuning, distillation, and evaluation — decide whether you need fine-tuning at all, distill your LLM logs into small local models (MLX/LoRA), evaluate with held-out discipline, and learn the why at every step.
githubInstall
Resourcemeasured growth/plugin marketplace add rchaz/tunelab6GitHub starsstable