ModelRefs / LLM Observability — Canonical Workflow
LLM Observability — Canonical Workflow
LLM Observability: provisional AI workflow implementation reference with candidate models, providers, tools, and architecture.
Overview
LLM observability is the end-to-end monitoring stack for a deployed LLM application — traces, evals, cost tracking, drift detection, and incident response so quality or cost regressions surface before users notice them.
Use this page to check what an observability stack needs to cover — tracing, evaluation, cost, drift, and incident workflows — and review the related tool, architecture, and evaluation-pipeline references before implementing one.
This workflow's representative-workload evaluation protocol has been drafted, but ModelRefs does not yet have recorded run evidence confirming detection and attribution across prompt, model, provider, and deployment changes. Treat readiness or completion claims as provisional until validated on your own workload.
Implementation profile
| Category | llms |
|---|---|
| Implementation maturity | enterprise |
| Evidence status | partial |
| Primary use cases | reasoning |
| Deployment options | managed-api, hybrid |
| Architectures | serverless-api, managed-container, self-hosted-cluster |
Candidate models with published references
- BGE-M3
- GPT-5
- GPT-5 Mini
- Claude Opus 4
- Llama 4 Scout
- DeepSeek R1
- Mistral Large 2
- Command R+
- o3
- o4 Mini
- Text Embedding 3 Large
- Claude Sonnet 4
Coverage means the model is a candidate worth evaluating for this workflow, not a ranking or a recommendation. Models whose reference pages are still in review are omitted.
Benchmarks relevant to this workflow
miracl, mkqa, mldr, swe-bench, aider-polyglot, gpqa, aime-2025, tau-bench, browsecomp-long-context, longfact-concepts, terminal-bench, mmmu, mmlu-pro, livecodebench.
Relevance is a coverage signal from the canonical registry. Each benchmark only describes its own protocol and date, so confirm the harness matches your workload before treating a score as evidence.
Continue your research
Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to LLM Observability — Canonical Workflow.