ModelRefs / Time-to-First-Token P95 Leaderboard — AI Model Scores
Time-to-First-Token P95 Leaderboard — AI Model Scores
95th-percentile first-token latency under 256-token prompt. Current leaders, methodology, and citation sources for Time-to-First-Token P95.
Overview
95th-percentile first-token latency under 256-token prompt.
How it is measured: Aggregated over 1k requests, US-East endpoint.
How this benchmark is scored
| Category | latency |
|---|---|
| Maximum score | 5000 ms |
| Direction | Lower is better |
Primary source: https://artificialanalysis.ai/
Continue your research
Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to Time-to-First-Token P95 Leaderboard — AI Model Scores.