OpenAI API
OpenAI's developer platform for calling its GPT and reasoning models, with chat, tools, audio and image features.
Best for: Access to OpenAI's own models.
Visit OpenAI API →LLM API providers, open-model inference hosts, fast-chip clouds, routers and self-hosted serving tools, plus a side-by-side comparison of the vLLM, SGLang, TensorRT-LLM, llama.cpp and Ollama inference engines.
OpenAI's developer platform for calling its GPT and reasoning models, with chat, tools, audio and image features.
Best for: Access to OpenAI's own models.
Visit OpenAI API →Anthropic's API for the Claude model family, with tool use, long context and vision.
Best for: Claude models with strong coding and agent support.
Visit Anthropic Claude API →Google's API for the Gemini models, with multimodal input, long context and tool use.
Best for: Multimodal and long-context workloads.
Visit Google Gemini API →Mistral's API for its proprietary and open-weight models, with function calling and embeddings.
Best for: European provider with open and hosted models.
Visit Mistral AI API →Cohere's API for its Command, embedding and rerank models, aimed at enterprise search and generation.
Best for: Enterprise RAG and retrieval.
Visit Cohere API →API for xAI's Grok models.
Best for: Access to Grok models.
Visit xAI API →DeepSeek's API for its chat and reasoning models, using an OpenAI-compatible format.
Best for: Low-cost open-lab models behind a simple API.
Visit DeepSeek API →Moonshot AI's developer platform for the Kimi models.
Best for: Long-context and agentic Kimi models.
Visit Kimi Open Platform →Alibaba Cloud's platform for using Qwen and other models through APIs.
Best for: Qwen models on Alibaba Cloud.
Visit Alibaba Cloud Model Studio →Perplexity's API for search-grounded answers and its models.
Best for: Answers that cite web sources.
Visit Perplexity API →DigitalOcean's single endpoint, compatible with OpenAI and Anthropic APIs, for running production LLM workloads.
Best for: Teams already on DigitalOcean.
Visit DigitalOcean serverless inference →Runs inference for open-weight models and supports LoRA adapters, full fine-tuning and DPO tuning.
Best for: Open models plus customisation.
Visit Together AI →Serves open-weight models and supports full-parameter fine-tuning on very large models.
Best for: Fast serving and fine-tuning of open models.
Visit Fireworks AI →Managed inference API with a catalogue of dozens of models and compliance certifications.
Best for: Open models with enterprise compliance needs.
Visit Nebius →Platform with OpenAI-compatible endpoints for running a large catalogue of language and multimodal models.
Best for: Wide model choice in one place.
Visit SiliconFlow →Inference platform that serves open-source, custom and fine-tuned models.
Best for: Custom models in production.
Visit Baseten →Pay-per-use API for running many open models without managing servers.
Best for: Simple pay-as-you-go access.
Visit DeepInfra →Hugging Face service that gives one interface to several inference providers for models on the Hub.
Best for: Hub models via a unified API.
Visit Hugging Face Inference Providers →NVIDIA's catalogue for trying and calling optimised models, built on NVIDIA microservices.
Best for: Testing NVIDIA-optimised models.
Visit NVIDIA API catalog →Runs inference on its own custom LPU chips, aimed at low latency.
Best for: Latency-sensitive chat and agents.
Visit Groq →Runs inference on its wafer-scale processors and offers it through a cloud API.
Best for: Very high tokens per second.
Visit Cerebras →Cloud inference built on its own dataflow chips, also available as on-premises systems.
Best for: Enterprise inference on dedicated hardware.
Visit SambaNova →AWS service for using foundation models from several providers through one API, with security and governance features.
Best for: Models inside an AWS environment.
Visit Amazon Bedrock →Serverless inference on Cloudflare's network, with a catalogue of open models.
Best for: Edge-served inference close to users.
Visit Cloudflare Workers AI →Serverless compute platform for running GPU workloads, deployed from the command line.
Best for: Self-managed models without servers.
Visit Modal →Unified API layer over hundreds of models from many providers.
Best for: One account, many providers.
Visit OpenRouter →Open-source gateway and SDK that exposes many model providers through an OpenAI-style interface.
Best for: Provider switching and spend tracking.
Visit LiteLLM →High-throughput open-source engine for serving LLMs, known for PagedAttention and continuous batching.
Best for: The common default for self-hosted serving.
Visit vLLM →Open-source serving framework with a fast runtime for language and multimodal models.
Best for: High-performance serving and structured output.
Visit SGLang →NVIDIA's open-source library for optimised LLM inference on NVIDIA GPUs.
Best for: Peak performance on NVIDIA hardware.
Visit TensorRT-LLM →C/C++ engine that runs quantised models on CPUs and many GPUs.
Best for: Local and edge inference.
Visit llama.cpp →Tool for downloading and running open models locally through a simple command line and API.
Best for: Easy local use.
Visit Ollama →Kubernetes-native project for distributed LLM inference with cache-aware routing and disaggregated serving.
Best for: Cluster-scale serving on Kubernetes.
Visit llm-d →NVIDIA's open framework for distributed, disaggregated inference of generative models.
Best for: Multi-node GPU serving.
Visit NVIDIA Dynamo →Independent comparison of API providers on price, output speed and latency for the same models.
Best for: Comparing providers before you commit.
Visit Artificial Analysis provider benchmarks →vLLM, SGLang, TensorRT-LLM, llama.cpp and Ollama all turn model weights into tokens, but they optimise for different things.
Core mechanism: PagedAttention manages the KV cache like virtual memory — paging it into non-contiguous blocks so memory is used almost fully instead of being wasted on padding. That, plus continuous batching, is what lets it serve many concurrent requests efficiently.
Core mechanism: RadixAttention organizes the KV cache in a radix tree, so requests that share a long common prefix — a system prompt, a RAG context block, a multi-turn chat history — reuse cached computation instead of recomputing it.
Core mechanism: Compiles models into optimized NVIDIA runtime engines — kernel fusion, in-flight batching, and quantization tuned specifically to the target GPU — extracting close to peak hardware FLOPS.
Core mechanism: A C/C++ inference runtime built around GGUF quantized model files, single-stream focused, and portable across an unusually wide range of hardware backends.
Core mechanism: A friendly wrapper around llama.cpp that handles model pulling, automatic quantization selection, and a simple CLI/API — optimized for a smooth "one command and it runs" experience.
| Dimension | vLLM | SGLang | TensorRT-LLM | llama.cpp | Ollama |
|---|---|---|---|---|---|
| Core mechanism | PagedAttention | RadixAttention | Compiled kernels | GGUF runtime | llama.cpp wrapper |
| Primary hardware | NVIDIA (+ multi-backend) | NVIDIA, AMD, TPU, NPU | NVIDIA only | Almost anything | NVIDIA, AMD, Apple, CPU |
| Best for | Multi-user production serving | Prefix-heavy RAG / chat | Max throughput at scale | Edge / non-NVIDIA / portability | Local dev & prototyping |
| Concurrency model | Continuous batching | Continuous batching + cache reuse | In-flight batching | Single-stream focused | Single-stream focused |
| Setup complexity | Moderate | Moderate | High | Low | Very low |
| App-facing API | Built-in OpenAI-compatible server | Built-in OpenAI-compatible server | Needs Triton wrapper | Built-in server | Built-in CLI/API |
These aren't five competitors picking from the same menu — they cluster into families. llama.cpp is the portable C/C++ foundation that Ollama (and several other local-inference tools) wrap for a friendlier developer experience. vLLM and SGLang are both Python-native, cluster-oriented servers built for concurrent production traffic, differentiated mainly by their KV-cache strategy — pure paging versus prefix-aware reuse. TensorRT-LLM sits apart as NVIDIA's own compiled runtime, trading portability entirely for peak throughput on its own hardware.
There's no single "fastest" inference engine — only the fastest engine for your hardware, your concurrency pattern, and how much of your traffic shares a prefix. — The recurring lesson of every inference-engine benchmark
In practice, many teams end up running more than one: llama.cpp or Ollama on a developer's laptop during prototyping, vLLM or SGLang for the production API once the model and traffic pattern are settled, and TensorRT-LLM as the final throughput optimization once cost-per-token on NVIDIA infrastructure starts to matter.
Match your requirement to a starting point, then compare vendors.
| Use a specific frontier model | The model maker's own API, such as OpenAI, Anthropic or Google |
|---|---|
| Run open models without servers | An open-model host such as Together, Fireworks, Nebius or DeepInfra |
| Lowest latency | A fast-chip cloud such as Groq, Cerebras or SambaNova |
| Stay inside a cloud account | Amazon Bedrock or your cloud's own model service |
| Avoid lock-in to one provider | A router or gateway such as OpenRouter or LiteLLM |
| Full control of data and cost | Self-hosted serving with vLLM, SGLang or llm-d |
Neighbouring directories and hubs.
Quick answers.
It is a service or tool that runs a large language model so applications can send prompts and receive responses. It can be a model maker's API, a host for open models, a cloud service or software you run yourself.
Most providers charge per million input and output tokens, sometimes with discounts for cached input or batch jobs. Dedicated endpoints are usually billed by GPU time instead.
The provider accepts the same request format as the OpenAI API, so you can switch by changing the base URL and key. Check support for tools, streaming and structured output, which can differ.
Time to first token, output tokens per second, price per million tokens, context length, rate limits, uptime and data-handling terms. Test with your own prompts.
Use Ollama for laptop prototyping, llama.cpp for non-NVIDIA or edge hardware, vLLM for a self-hosted production API, SGLang when many requests share long prefixes such as RAG, and TensorRT-LLM for peak throughput on NVIDIA data-centre GPUs. See the engines comparison on this page.
It depends on the provider and plan. Read the data-retention and training terms, and look for zero-retention options if you handle sensitive data.