ModelRefs / Throughput @ Batch 32 Leaderboard — AI Model Scores
Throughput @ Batch 32 Leaderboard — AI Model Scores
Sustained tokens/sec under batched server-side inference. Current leaders, methodology, and citation sources for Throughput @ Batch 32.
Overview
Sustained tokens/sec under batched server-side inference.
How it is measured: Mean tok/s with batch=32, 1024-token gen, fp8 where supported.
How this benchmark is scored
| Category | latency |
|---|---|
| Maximum score | 5000 tok/s |
| Direction | Higher is better |
Primary source: https://artificialanalysis.ai/
Continue your research
Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to Throughput @ Batch 32 Leaderboard — AI Model Scores.