ModelRefs / P99 Latency — AI Glossary
P99 Latency — AI Glossary
The 99th-percentile request latency; the worst-case response time experienced by 1-in-100 requests under load. The average hides what users actually experience.
Overview
P99 (and P95, P50) latency percentiles capture tail behavior that median or average masks. For LLM APIs, P99 spikes during traffic bursts, KV cache evictions, or cold starts. SLOs typically specify P50 < 300 ms TTFT and P99 < 2 s. Profiling with tools like Grafana or Datadog exposes P99 regressions.
Reference details
| Topic | inference |
|---|---|
| Also known as | tail latency, 99th percentile latency |
| Last reviewed | 2026-06-24 |
Related terms
Commonly confused with
The average hides what users actually experience. A one-in-a-hundred worst case sounds rare until a page makes several model calls: at ten calls, the chance of at least one landing in that tail is 1 − 0.99¹⁰ ≈ 9.6%, so roughly one page view in ten is as slow as the P99. Report a percentile and a sample size; an average alone is not a latency claim.
Continue your research
Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to P99 Latency — AI Glossary.
Frequently asked questions
What is P99 Latency?
The 99th-percentile request latency; the worst-case response time experienced by 1-in-100 requests under load.
Is P99 Latency the same as tail latency?
Yes — tail latency, 99th percentile latency are common aliases for P99 Latency.
What concepts are related to P99 Latency?
Closely related concepts include time to first token, tokens per second, throughput.