LLM Inference Platforms and Tools logopeterindia.netLLM Inference Platforms
Home / LLM Hub / LLM Inference Platforms and Tools
LLM infrastructure · 2026

LLM Inference Platforms & Tools

LLM API providers, open-model inference hosts, fast-chip clouds, routers and self-hosted serving tools, plus a side-by-side comparison of the vLLM, SGLang, TensorRT-LLM, llama.cpp and Ollama inference engines.

35platforms and tools
7Categories
6FAQs answered
01Model-maker APIs

OpenAI API

OpenAI's developer platform for calling its GPT and reasoning models, with chat, tools, audio and image features.

Best for: Access to OpenAI's own models.

Visit OpenAI API →
02Model-maker APIs

Anthropic Claude API

Anthropic's API for the Claude model family, with tool use, long context and vision.

Best for: Claude models with strong coding and agent support.

Visit Anthropic Claude API →
03Model-maker APIs

Google Gemini API

Google's API for the Gemini models, with multimodal input, long context and tool use.

Best for: Multimodal and long-context workloads.

Visit Google Gemini API →
04Model-maker APIs

Mistral AI API

Mistral's API for its proprietary and open-weight models, with function calling and embeddings.

Best for: European provider with open and hosted models.

Visit Mistral AI API →
05Model-maker APIs

Cohere API

Cohere's API for its Command, embedding and rerank models, aimed at enterprise search and generation.

Best for: Enterprise RAG and retrieval.

Visit Cohere API →
06Model-maker APIs

xAI API

API for xAI's Grok models.

Best for: Access to Grok models.

Visit xAI API →
07Model-maker APIs

DeepSeek API

DeepSeek's API for its chat and reasoning models, using an OpenAI-compatible format.

Best for: Low-cost open-lab models behind a simple API.

Visit DeepSeek API →
08Model-maker APIs

Kimi Open Platform

Moonshot AI's developer platform for the Kimi models.

Best for: Long-context and agentic Kimi models.

Visit Kimi Open Platform →
09Model-maker APIs

Alibaba Cloud Model Studio

Alibaba Cloud's platform for using Qwen and other models through APIs.

Best for: Qwen models on Alibaba Cloud.

Visit Alibaba Cloud Model Studio →
10Model-maker APIs

Perplexity API

Perplexity's API for search-grounded answers and its models.

Best for: Answers that cite web sources.

Visit Perplexity API →
11Open-model inference hosts

DigitalOcean serverless inference

DigitalOcean's single endpoint, compatible with OpenAI and Anthropic APIs, for running production LLM workloads.

Best for: Teams already on DigitalOcean.

Visit DigitalOcean serverless inference →
12Open-model inference hosts

Together AI

Runs inference for open-weight models and supports LoRA adapters, full fine-tuning and DPO tuning.

Best for: Open models plus customisation.

Visit Together AI →
13Open-model inference hosts

Fireworks AI

Serves open-weight models and supports full-parameter fine-tuning on very large models.

Best for: Fast serving and fine-tuning of open models.

Visit Fireworks AI →
14Open-model inference hosts

Nebius

Managed inference API with a catalogue of dozens of models and compliance certifications.

Best for: Open models with enterprise compliance needs.

Visit Nebius →
15Open-model inference hosts

SiliconFlow

Platform with OpenAI-compatible endpoints for running a large catalogue of language and multimodal models.

Best for: Wide model choice in one place.

Visit SiliconFlow →
16Open-model inference hosts

Baseten

Inference platform that serves open-source, custom and fine-tuned models.

Best for: Custom models in production.

Visit Baseten →
17Open-model inference hosts

DeepInfra

Pay-per-use API for running many open models without managing servers.

Best for: Simple pay-as-you-go access.

Visit DeepInfra →
18Open-model inference hosts

Hugging Face Inference Providers

Hugging Face service that gives one interface to several inference providers for models on the Hub.

Best for: Hub models via a unified API.

Visit Hugging Face Inference Providers →
19Open-model inference hosts

NVIDIA API catalog

NVIDIA's catalogue for trying and calling optimised models, built on NVIDIA microservices.

Best for: Testing NVIDIA-optimised models.

Visit NVIDIA API catalog →
20Fast-chip inference clouds

Groq

Runs inference on its own custom LPU chips, aimed at low latency.

Best for: Latency-sensitive chat and agents.

Visit Groq →
21Fast-chip inference clouds

Cerebras

Runs inference on its wafer-scale processors and offers it through a cloud API.

Best for: Very high tokens per second.

Visit Cerebras →
22Fast-chip inference clouds

SambaNova

Cloud inference built on its own dataflow chips, also available as on-premises systems.

Best for: Enterprise inference on dedicated hardware.

Visit SambaNova →
23Cloud & serverless platforms

Amazon Bedrock

AWS service for using foundation models from several providers through one API, with security and governance features.

Best for: Models inside an AWS environment.

Visit Amazon Bedrock →
24Cloud & serverless platforms

Cloudflare Workers AI

Serverless inference on Cloudflare's network, with a catalogue of open models.

Best for: Edge-served inference close to users.

Visit Cloudflare Workers AI →
25Cloud & serverless platforms

Modal

Serverless compute platform for running GPU workloads, deployed from the command line.

Best for: Self-managed models without servers.

Visit Modal →
26Routers & aggregators

OpenRouter

Unified API layer over hundreds of models from many providers.

Best for: One account, many providers.

Visit OpenRouter →
27Routers & aggregators

LiteLLM

Open-source gateway and SDK that exposes many model providers through an OpenAI-style interface.

Best for: Provider switching and spend tracking.

Visit LiteLLM →
28Self-hosted serving

vLLM

High-throughput open-source engine for serving LLMs, known for PagedAttention and continuous batching.

Best for: The common default for self-hosted serving.

Visit vLLM →
29Self-hosted serving

SGLang

Open-source serving framework with a fast runtime for language and multimodal models.

Best for: High-performance serving and structured output.

Visit SGLang →
30Self-hosted serving

TensorRT-LLM

NVIDIA's open-source library for optimised LLM inference on NVIDIA GPUs.

Best for: Peak performance on NVIDIA hardware.

Visit TensorRT-LLM →
31Self-hosted serving

llama.cpp

C/C++ engine that runs quantised models on CPUs and many GPUs.

Best for: Local and edge inference.

Visit llama.cpp →
32Self-hosted serving

Ollama

Tool for downloading and running open models locally through a simple command line and API.

Best for: Easy local use.

Visit Ollama →
33Self-hosted serving

llm-d

Kubernetes-native project for distributed LLM inference with cache-aware routing and disaggregated serving.

Best for: Cluster-scale serving on Kubernetes.

Visit llm-d →
34Self-hosted serving

NVIDIA Dynamo

NVIDIA's open framework for distributed, disaggregated inference of generative models.

Best for: Multi-node GPU serving.

Visit NVIDIA Dynamo →
35Comparison

Artificial Analysis provider benchmarks

Independent comparison of API providers on price, output speed and latency for the same models.

Best for: Comparing providers before you commit.

Visit Artificial Analysis provider benchmarks →

Inference engines compared

vLLM, SGLang, TensorRT-LLM, llama.cpp and Ollama all turn model weights into tokens, but they optimise for different things.

vLLM

Production multi-user serving

Core mechanism: PagedAttention manages the KV cache like virtual memory — paging it into non-contiguous blocks so memory is used almost fully instead of being wasted on padding. That, plus continuous batching, is what lets it serve many concurrent requests efficiently.

Hardware

  • NVIDIA (primary)
  • AMD ROCm, CPU, TPU, Gaudi, Ascend (secondary)

Ideal use case

  • Small-team to mid-scale production serving
  • The default choice for a self-hosted LLM API

Strengths

  • Broad backend support
  • Tensor/pipeline/data/expert parallelism

Watch out for

  • Dependency drift — needs precise CUDA/PyTorch pinning

SGLang

Prefix-heavy, structured generation

Core mechanism: RadixAttention organizes the KV cache in a radix tree, so requests that share a long common prefix — a system prompt, a RAG context block, a multi-turn chat history — reuse cached computation instead of recomputing it.

Hardware

  • NVIDIA, AMD MI300/MI355
  • Intel Xeon, Google TPU, Ascend NPU

Ideal use case

  • RAG pipelines and multi-turn chat with repeated context
  • Structured/constrained output generation

Strengths

  • Large speedups specifically in high prefix-reuse workloads

Watch out for

  • The advantage is workload-specific, not a universal speedup

TensorRT-LLM

Peak throughput on NVIDIA silicon

Core mechanism: Compiles models into optimized NVIDIA runtime engines — kernel fusion, in-flight batching, and quantization tuned specifically to the target GPU — extracting close to peak hardware FLOPS.

Hardware

  • NVIDIA-exclusive: H100, H200, L4, RTX, Jetson AGX Orin

Ideal use case

  • Large-scale production on NVIDIA data-center GPUs
  • Typically paired with Triton Inference Server

Strengths

  • Maximum throughput per GPU on NVIDIA hardware

Watch out for

  • NVIDIA-only; no native app-facing API — needs a Triton wrapper; steeper MLOps overhead

llama.cpp

Maximum portability

Core mechanism: A C/C++ inference runtime built around GGUF quantized model files, single-stream focused, and portable across an unusually wide range of hardware backends.

Hardware

  • CUDA, HIP/ROCm, Metal, Vulkan, OpenCL
  • AVX/AVX-512 CPU, RISC-V, Intel GPU

Ideal use case

  • Single-user local inference; edge and non-NVIDIA hardware
  • CPU-plus-GPU hybrid workloads

Strengths

  • Runs almost anywhere; very CPU-efficient

Watch out for

  • Single-stream focus — not built for concurrent multi-user serving

Ollama

Local developer experience

Core mechanism: A friendly wrapper around llama.cpp that handles model pulling, automatic quantization selection, and a simple CLI/API — optimized for a smooth "one command and it runs" experience.

Hardware

  • NVIDIA CUDA, AMD ROCm, Apple Silicon Metal, CPU fallback

Ideal use case

  • Local development, prototyping, single-user laptops/workstations

Strengths

  • Deliberately short install path; broad hardware coverage out of the box

Watch out for

  • Limited multi-GPU support; not designed for concurrent batched serving

Side by side

DimensionvLLMSGLangTensorRT-LLMllama.cppOllama
Core mechanismPagedAttentionRadixAttentionCompiled kernelsGGUF runtimellama.cpp wrapper
Primary hardwareNVIDIA (+ multi-backend)NVIDIA, AMD, TPU, NPUNVIDIA onlyAlmost anythingNVIDIA, AMD, Apple, CPU
Best forMulti-user production servingPrefix-heavy RAG / chatMax throughput at scaleEdge / non-NVIDIA / portabilityLocal dev & prototyping
Concurrency modelContinuous batchingContinuous batching + cache reuseIn-flight batchingSingle-stream focusedSingle-stream focused
Setup complexityModerateModerateHighLowVery low
App-facing APIBuilt-in OpenAI-compatible serverBuilt-in OpenAI-compatible serverNeeds Triton wrapperBuilt-in serverBuilt-in CLI/API

Which engine should you use?

Prototyping on a laptopOllama. Fastest path from "I have a model name" to a running local endpoint, with sane defaults.
Non-NVIDIA or edge hardwarellama.cpp. Broadest hardware coverage of any engine here, and efficient on CPU-only boxes.
Self-hosted production APIvLLM. The default serving choice for teams running their own OpenAI-compatible endpoint at moderate concurrency.
RAG or long multi-turn chat at scaleSGLang. RadixAttention's prefix-sharing pays off specifically when many requests reuse the same context.
Max throughput on data-center NVIDIA GPUsTensorRT-LLM. Worth the setup overhead when you're optimizing cost-per-token at real scale.
Not sure yetStart with vLLM or Ollama depending on whether you're prototyping or serving — both have the gentlest path to "it works," and you can graduate to SGLang or TensorRT-LLM once your workload pattern is clear.

How they relate

These aren't five competitors picking from the same menu — they cluster into families. llama.cpp is the portable C/C++ foundation that Ollama (and several other local-inference tools) wrap for a friendlier developer experience. vLLM and SGLang are both Python-native, cluster-oriented servers built for concurrent production traffic, differentiated mainly by their KV-cache strategy — pure paging versus prefix-aware reuse. TensorRT-LLM sits apart as NVIDIA's own compiled runtime, trading portability entirely for peak throughput on its own hardware.

There's no single "fastest" inference engine — only the fastest engine for your hardware, your concurrency pattern, and how much of your traffic shares a prefix. — The recurring lesson of every inference-engine benchmark

In practice, many teams end up running more than one: llama.cpp or Ollama on a developer's laptop during prototyping, vLLM or SGLang for the production API once the model and traffic pattern are settled, and TensorRT-LLM as the final throughput optimization once cost-per-token on NVIDIA infrastructure starts to matter.

Choosing the right approach

Match your requirement to a starting point, then compare vendors.

Use a specific frontier modelThe model maker's own API, such as OpenAI, Anthropic or Google
Run open models without serversAn open-model host such as Together, Fireworks, Nebius or DeepInfra
Lowest latencyA fast-chip cloud such as Groq, Cerebras or SambaNova
Stay inside a cloud accountAmazon Bedrock or your cloud's own model service
Avoid lock-in to one providerA router or gateway such as OpenRouter or LiteLLM
Full control of data and costSelf-hosted serving with vLLM, SGLang or llm-d

Related pages on PeterIndia.net

Neighbouring directories and hubs.

Frequently asked questions

Quick answers.

What is an LLM inference platform?

It is a service or tool that runs a large language model so applications can send prompts and receive responses. It can be a model maker's API, a host for open models, a cloud service or software you run yourself.

How is LLM API pricing usually charged?

Most providers charge per million input and output tokens, sometimes with discounts for cached input or batch jobs. Dedicated endpoints are usually billed by GPU time instead.

What does OpenAI-compatible mean?

The provider accepts the same request format as the OpenAI API, so you can switch by changing the base URL and key. Check support for tools, streaming and structured output, which can differ.

What should I measure when comparing providers?

Time to first token, output tokens per second, price per million tokens, context length, rate limits, uptime and data-handling terms. Test with your own prompts.

Which inference engine should I choose?

Use Ollama for laptop prototyping, llama.cpp for non-NVIDIA or edge hardware, vLLM for a self-hosted production API, SGLang when many requests share long prefixes such as RAG, and TensorRT-LLM for peak throughput on NVIDIA data-centre GPUs. See the engines comparison on this page.

Is my data used for training?

It depends on the provider and plan. Read the data-retention and training terms, and look for zero-retention options if you handle sensitive data.