LLM observability: tools for developers
Monitoring, traces, evaluation and quality tracking for LLM applications. An entry can appear under several use cases when its topics or description provide several explicit signals.
Use case
Activity
Sort by
121 entries in this view.
Tool ranking
Mixed ranking: measured growth takes priority; entries without two snapshots are estimated. Raw values keep their source-specific window, shown under each tool; a change requires at least two snapshots from the last 7 days.
- 1Active
agent-mmm
SkillMarketing Mix Model expert agent plugin for Claude Code - pymc-marketing v0.18.2+
githubmeasured growthOpen source ↗
Install
/plugin marketplace add Yakoub-ai/agent-mmm4GitHub starsstable - 2Active
tulip-agents
AgentThe agent framework where the model never holds the trigger — every consequential action clears your policy first, waits for a human when it matters, and lands on a record you can verify. Build on it, or put it around the agent you already…
githubmeasured growthOpen source ↗
1GitHub starsstable - 3Active
adlc-team-skills
Skill🐙 ADLC Team Skills — Agentic SDLC for Engineering Teams
githubmeasured growthOpen source ↗
Install
git clone https://github.com/tikalk/adlc-team-skills ~/.claude/skills/adlc-team-skills132GitHub stars+1 (+0.76 %) - 4Dormant
Agents-eval
AgentA Multi-Agent System (MAS) evaluation framework using PydanticAI that generates and evaluates scientific paper reviews through a three-tiered assessment approach: traditional metrics, LLM-as-a-Judge, and graph-based complexity analysis.
githubmeasured growthOpen source ↗
2GitHub starsstable - 5Dormant
MCP as a Judge: a behavioral MCP that strengthens AI coding assistants via explicit LLM evaluations
mcpmeasured growthOpen source ↗
Install
claude mcp add mcp-as-a-judge -- uvx mcp-as-a-judge17GitHub starsstable - 6Dormant
An AI-powered multi-agent system that demonstrates clinical triage, OTC medication recommendations, and e-pharmacy integration for respiratory conditions. Built with modular agents that collaborate to provide safe, intelligent healthcare…
githubmeasured growthOpen source ↗
0GitHub starsstable - 7Active
langfuse-mcp
MCPA Model Context Protocol (MCP) server for Langfuse, enabling AI agents to query Langfuse trace data for enhanced debugging and observability
githubmeasured growthOpen source ↗
105GitHub starsstable - 8Dormant
CustoFlow
OtherMulti-agent customer support system with Google ADK & Gemini 2.5 Flash Lite. Kaggle capstone demonstrating 11+ concepts. Automates 80%+ queries, <10s response time.
githubmeasured growthOpen source ↗
2GitHub starsstable - 9Active
Architecture-first Python scaffold for an auditable multi-agent corporate credit desk using A2A, MCP, deterministic credit policies, model routing, and OpenTelemetry.
githubmeasured growthOpen source ↗
0GitHub starsstable - 10Active
kubesphere
SkillThe container platform tailored for Kubernetes multi-cloud, datacenter, and edge management ⎈ 🖥 ☁️
githubmeasured growthOpen source ↗
Install
git clone https://github.com/kubesphere/kubesphere ~/.claude/skills/kubesphere17 035GitHub starsstable - 11Dormant
Two Claude Code skills for adversarial code review (single-model PAR and multi-model MMAR with cross-critique to catch hallucinations), plus a fixture-based eval suite.
githubmeasured growthOpen source ↗
Install
/plugin marketplace add prime-radiant-inc/parallel-adversarial-review17GitHub starsstable - 12Active
Mini Program Engineering Skill Suite is an Agent Skill suite for evidence‑first mini‑program development. It helps agents bring WeChat or other mini‑program projects from vague intent to reliable engineering work: project intake, product…
githubmeasured growthOpen source ↗
Install
git clone https://github.com/NocodeMrLi/mini-program-engineering-skill-suite ~/.claude/skills/mini-program-engineering-skill-suite45GitHub stars+6 (+15.4 %) - 13Active
Build tested agent skills and govern their lifecycle through a user-defined marketplace: evidence, discovery, updates, rollback, quarantine, and 17-platform distribution.
githubmeasured growthOpen source ↗
Install
git clone https://github.com/FrancyJGLisboa/agent-skills-platform ~/.claude/skills/agent-skills-platform2 376GitHub starsstable - 14Active
sc-prism-releases
SkillRun many AIs on one board and keep control of all of it. Deterministic code decides who acts — never a model. A privacy floor keeps sensitive work on your machine, your own tests decide what counts as done, and every action lands on a…
githubmeasured growthOpen source ↗
Install
git clone https://github.com/sandhusukhdeep2/sc-prism-releases ~/.claude/skills/sc-prism-releases1GitHub starsstable - 15Active
superlog
OtherOpen-source observability tool that uses AI agents to self-heal your software
githubestimated momentumOpen source ↗
1 404GitHub stars—
Learning resources
Ranked by measured growth, normalized across sources. These resources remain available separately and do not take part in the main tool ranking.
- 1ActiveOpen source ↗
Agentic_AI_Engineer
AgentMy complete journey to becoming an Agentic AI Engineer through structured learning, projects, experiments, and production-ready implementations of modern AI systems.
githubResourcemeasured growth18GitHub starsstable - 2ActiveOpen source ↗
Curated + scored map of official MCP servers and agents for DevOps, Cloud, SRE, and Platform Engineering — every entry rated on production access, approval gates, and audit evidence.
githubResourcemeasured growth77GitHub starsstable - 3ActiveOpen source ↗
50+ curated LLM observability tools PLUS 26 Agent Skills (several with runnable, unit-tested scripts) to build, evaluate, debug, secure & monitor reliable LLM apps. Tracing, evals, guardrails, LLMOps.
githubInstall
Resourcemeasured growthgit clone https://github.com/ContextJet-ai/awesome-llm-observability ~/.claude/skills/awesome-llm-observability33GitHub stars+2 (+6.5 %) - 4ActiveOpen source ↗
Handbook técnico aberto sobre engenharia de IA em produção: ML tradicional, LLMs, RAG, agentes, segurança, observabilidade, FinOps e deployment.
githubResourcemeasured growth1GitHub starsstable - 5ActiveOpen source ↗
trigger_tree
SkillDocumentation-discovery telemetry for Claude Code — heat/cold maps, health grade, evidence-backed router fixes. 100% local, zero tokens. /tt
githubInstall
Resourcemeasured growth/plugin marketplace add Hedde/trigger_tree14GitHub starsstable - 6ActiveOpen source ↗
tunelab
SkillClaude Code plugin for LLM fine-tuning, distillation, and evaluation — decide whether you need fine-tuning at all, distill your LLM logs into small local models (MLX/LoRA), evaluate with held-out discipline, and learn the why at every step.
githubInstall
Resourcemeasured growth/plugin marketplace add rchaz/tunelab6GitHub starsstable