LLM observability: tools for developers

Monitoring, traces, evaluation and quality tracking for LLM applications. An entry can appear under several use cases when its topics or description provide several explicit signals.

Use case

Activity

Sort by

117 entries in this view.

Tool ranking

Ranked by measured growth, normalized across sources. Raw values keep their source-specific window, shown under each tool; a change requires at least two snapshots from the last 7 days.

  1. 1
    Active

    Query Spanlens LLM observability from Cursor, Claude Desktop, or Continue via MCP.

    mcpmeasured growthOpen source ↗

    Install claude mcp add mcp-server -- npx @spanlens/mcp-server

    12GitHub starsstable
  2. 2

    Other
    Active

    Test your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare performance of GPT, Claude, Gemini, DeepSeek, and more. Simple declarative configs with command line and CI/CD integration. Used by OpenAI…

    githubmeasured growthOpen source ↗

    24 768GitHub stars+142 (+0.58 %)
  3. 3

    Other
    Active

    AI SRE AgenticOps for Kubernetes and cloud infrastructure.

    githubmeasured growthOpen source ↗

    782GitHub stars+1 (+0.13 %)
  4. 4

    Agent
    Active

    The open source AI engineering platform for agents, LLMs, and ML models. MLflow enables teams of all sizes to debug, evaluate, monitor, and optimize production-quality AI applications while controlling costs and managing access to models…

    githubmeasured growthOpen source ↗

    27 783GitHub stars+82 (+0.30 %)
  5. 5

    Library
    Active

    Prefect is a workflow orchestration framework for building resilient data pipelines in Python.

    githubmeasured growthOpen source ↗

    23 766GitHub stars+66 (+0.28 %)
  6. 6

    Other
    Active

    A high-performance observability data pipeline.

    githubmeasured growthOpen source ↗

    22 509GitHub stars+40 (+0.18 %)
  7. 7

    Other
    Active

    Open source observability platform for logs, metrics, traces, frontend monitoring, pipelines and LLM observability. A sophisticated, simple and highly performant alternative to Datadog, Splunk, and Elasticsearch with 140x lower storage…

    githubmeasured growthOpen source ↗

    21 615GitHub stars+96 (+0.45 %)
  8. 8

    Other
    Active

    SigNoz is an open-source, OpenTelemetry-native observability platform for your team and their AI agents. Get logs, metrics, and traces in one tool with features like APM, distributed tracing, log management, infra monitoring, etc. Combined…

    githubmeasured growthOpen source ↗

    32 003GitHub stars+50 (+0.16 %)
  9. 9

    Other
    Active

    Mastra is the modern TypeScript framework for AI-powered applications and agents.

    githubmeasured growthOpen source ↗

    27 658GitHub stars+128 (+0.46 %)
  10. 10

    Other
    Active

    The fastest path to AI-powered full stack observability, even for lean teams.

    githubmeasured growthOpen source ↗

    80 412GitHub stars+85 (+0.11 %)
  11. 11
    Active

    🫖 Status page with uptime monitoring & API monitoring as code 🫖

    githubmeasured growthOpen source ↗

    9 056GitHub stars+31 (+0.34 %)
  12. 12
    Active

    Self-hosted AI SRE for Kubernetes — zero-instrumentation eBPF observability plus a copilot that fixes issues through guardrailed, self-verifying actions. BYO-LLM, air-gapped capable.

    githubmeasured growthOpen source ↗

    131GitHub stars-4 (-3.0 %)
  13. 13

    Other
    Active

    eBPF-based Networking, Security, and Observability

    githubmeasured growthOpen source ↗

    25 047GitHub stars+29 (+0.12 %)
  14. 14

    Agent
    Active

    eBPF-powered network observability for Kubernetes. Indexes L4/L7 traffic with full K8s context, decrypts TLS without keys. Queryable by AI agents via MCP and humans via dashboard.

    githubmeasured growthOpen source ↗

    12 068GitHub stars+8 (+0.07 %)
  15. 15

    Agent
    Active

    DataBuff is an AI-native APM built on Opentelemetry,with multi-agent troubleshooting out of the box.

    githubmeasured growthOpen source ↗

    642GitHub stars+32 (+5.2 %)
  16. 16

    Other
    Active

    Evidence-first infrastructure for reproducible, isolated executable evaluation and online rewards.

    githubmeasured growthOpen source ↗

    140GitHub stars+24 (+20.7 %)
  17. 17

    Skill
    Active

    The Fable Workflow: how Claude Fable 5 worked, distilled into skills any model can run, with the eval that keeps it honest. Think / act / prove.

    githubmeasured growthOpen source ↗

    Install git clone https://github.com/Sahir619/fable-method ~/.claude/skills/fable-method

    2 272GitHub stars+16 (+0.71 %)
  18. 18

    Agent
    Active

    The Operating System for Scalable Enterprise AI Agents - Run, orchestrate, and deploy Compliant Enterprise AI Agents at scale across frameworks, without lock-in, rewrites or fragile glue code. Native support for MCP, A2A. Interface with all…

    githubmeasured growthOpen source ↗

    156GitHub stars+19 (+13.9 %)
  19. 19
    Active

    Production reliability, benchmarking, and evaluation layer for MCP agent runtimes — protocol negotiation, A2A interop, circuit breakers, connection pooling, adaptive streaming, and autoscaling, with measured throughput/latency and a…

    githubmeasured growthOpen source ↗

    3GitHub starsstable
  20. 20

    Skill
    Dormant

    An assembly line for AI software development. 35 skills, 11 agent personas, 29 commands. From raw idea to shipped code.

    githubmeasured growthOpen source ↗

    Install /plugin marketplace add aneja5/forge-skills

    3GitHub starsstable
  21. 21
    Active

    Architecture-first Python scaffold for an auditable multi-agent corporate credit desk using A2A, MCP, deterministic credit policies, model routing, and OpenTelemetry.

    githubmeasured growthOpen source ↗

    0GitHub starsstable
  22. 22

    Agent
    Active

    Multi-agent SRE on-call investigator that auto-triages Slack/Discord infrastructure alerts via AWS Bedrock AgentCore, fanning out to specialized agents (CloudWatch, EKS, Slack/Discord scanners) for parallel investigation.

    githubmeasured growthOpen source ↗

    3GitHub starsstable
  23. 23

    MCP
    Active

    AURA is a production-tested SRE agent platform you can deploy in minutes. AURA handles the guardrails, APIs, state management, streaming, and failure handling required to put AI to work safely on production infrastructure.

    githubmeasured growthOpen source ↗

    254GitHub stars-3 (-1.2 %)
  24. 24

    Agent
    Active

    Generic DeepSeek Harness (dsh) plugins: A2A protocol server, session storage mirror, and Langfuse observability.

    githubmeasured growthOpen source ↗

    1GitHub starsstable
  25. 25

    MCP
    Active

    A typed, replayable, budget-aware programming language for LLM agents. Compiles to Python & TypeScript.

    githubmeasured growthOpen source ↗

    2GitHub starsstable