ModelRefs / Inference Provider — AI Glossary

Inference Provider — AI Glossary

A cloud service offering on-demand access to hosted LLMs via API, handling GPU infrastructure, scaling, and model serving.

Overview

Inference providers include: OpenAI (proprietary models), Anthropic (Claude API), Google AI (Gemini), Together AI (open models), Groq (LPU hardware), Fireworks AI, Replicate, Modal, and AWS Bedrock. They abstract GPU management, scale, and model versioning. Selection criteria: latency, throughput, model availability, pricing, and rate limits.

Reference details

Topicinference
Last reviewed2026-06-24

Continue your research

Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to Inference Provider — AI Glossary.

Frequently asked questions

What is Inference Provider?

A cloud service offering on-demand access to hosted LLMs via API, handling GPU infrastructure, scaling, and model serving.

What concepts are related to Inference Provider?

Closely related concepts include serverless inference, openai compatible, model routing.