ModelRefs / Inference Provider — AI Glossary
Inference Provider — AI Glossary
A cloud service offering on-demand access to hosted LLMs via API, handling GPU infrastructure, scaling, and model serving.
Overview
Inference providers include: OpenAI (proprietary models), Anthropic (Claude API), Google AI (Gemini), Together AI (open models), Groq (LPU hardware), Fireworks AI, Replicate, Modal, and AWS Bedrock. They abstract GPU management, scale, and model versioning. Selection criteria: latency, throughput, model availability, pricing, and rate limits.
Reference details
| Topic | inference |
|---|---|
| Last reviewed | 2026-06-24 |
Related terms
Continue your research
Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to Inference Provider — AI Glossary.
Frequently asked questions
What is Inference Provider?
A cloud service offering on-demand access to hosted LLMs via API, handling GPU infrastructure, scaling, and model serving.
What concepts are related to Inference Provider?
Closely related concepts include serverless inference, openai compatible, model routing.