Together AI
Cloud platform for running open-source models through serverless and dedicated endpoints, with fine-tuning and GPU clusters.
Best for: Open models behind a simple API.
Visit Together AI →Managed inference APIs, custom-silicon clouds, hyperscaler platforms, GPU clouds, inference engines and serving frameworks for running AI models in production.
Cloud platform for running open-source models through serverless and dedicated endpoints, with fine-tuning and GPU clusters.
Best for: Open models behind a simple API.
Visit Together AI →Inference platform for open and custom models, focused on speed, throughput and production reliability.
Best for: Fast serving of open models.
Visit Fireworks AI →Inference platform for deploying and scaling models in production, including custom and open-source models.
Best for: Production model serving with autoscaling.
Visit Baseten →Pay-per-use API for running many open models without managing servers.
Best for: Low-friction access to open models.
Visit DeepInfra →API and model library for running open-source and community models, with simple deployment of custom models.
Best for: Trying and shipping models quickly.
Visit Replicate →Generative media platform for fast inference of image, video and audio models.
Best for: Generative media workloads.
Visit fal →Managed service for deploying models from the Hugging Face Hub onto dedicated infrastructure.
Best for: Deploying Hub models without running servers.
Visit Hugging Face Inference Endpoints →Unified API and marketplace that routes requests to many model providers with one account.
Best for: Switching providers behind one API.
Visit OpenRouter →Inference cloud built on Groq's own LPU chips, aimed at very fast token generation.
Best for: Latency-sensitive applications.
Visit Groq →AI company running inference on its wafer-scale processors, offered through a cloud API.
Best for: Very high output speed.
Visit Cerebras →AI platform built on its own reconfigurable dataflow chips, offering cloud inference and on-premises systems.
Best for: Enterprise inference on dedicated hardware.
Visit SambaNova →AWS service for accessing foundation models from several providers through one API, with security and governance features.
Best for: Managed models inside AWS.
Visit Amazon Bedrock →Microsoft's catalogue and hosting for foundation models within Azure's AI platform.
Best for: Models within Azure.
Visit Microsoft Foundry models →Google Cloud's AI platform for building and deploying models and agents, formerly known as Vertex AI.
Best for: Models and agents on Google Cloud.
Visit Google Gemini Enterprise Agent Platform →Serverless inference on Cloudflare's global network, with a catalogue of open models.
Best for: Low-latency inference close to users.
Visit Cloudflare Workers AI →Serverless compute platform for running code and models on GPUs with fast cold starts.
Best for: Python-first serverless GPU workloads.
Visit Modal →GPU cloud with pods and serverless endpoints for running and scaling AI workloads.
Best for: Affordable GPU endpoints.
Visit RunPod →Specialised AI cloud providing large-scale GPU infrastructure for training and inference.
Best for: Large GPU capacity.
Visit CoreWeave →Lambda's inference API and GPU cloud for serving open models.
Best for: GPU cloud plus managed inference.
Visit Lambda Inference →AI cloud with GPU infrastructure and managed inference services.
Best for: European AI cloud capacity.
Visit Nebius →High-throughput open-source engine for serving large language models, known for PagedAttention and continuous batching.
Best for: The common default for self-hosted LLM serving.
Visit vLLM →Open-source serving framework with a fast runtime for large language and multimodal models.
Best for: High-performance serving and structured generation.
Visit SGLang →NVIDIA's open-source library for optimised LLM inference on NVIDIA GPUs.
Best for: Maximum performance on NVIDIA hardware.
Visit TensorRT-LLM →C/C++ inference engine that runs quantised models on CPUs and a wide range of GPUs.
Best for: Local and edge inference.
Visit llama.cpp →Tool for downloading and running open models locally with a simple command line and API.
Best for: Easy local model use.
Visit Ollama →Hugging Face's inference server for large language models. The project is now in maintenance mode, so check alternatives for new work.
Best for: Existing deployments that already use it.
Visit Text Generation Inference →Cross-platform runtime for running ONNX models on many hardware targets.
Best for: Portable inference across devices.
Visit ONNX Runtime →NVIDIA's framework for distributed, disaggregated inference of generative models across many GPUs.
Best for: Large-scale multi-node serving.
Visit NVIDIA Dynamo →Kubernetes-native project for distributed LLM inference with cache-aware routing and disaggregated serving.
Best for: Cluster-level LLM serving on Kubernetes.
Visit llm-d →NVIDIA's inference server, formerly Triton Inference Server, for serving models from many frameworks.
Best for: Multi-framework model serving.
Visit NVIDIA Dynamo-Triton →Kubernetes-based standard for serving machine learning and generative models, with autoscaling.
Best for: Model serving on Kubernetes.
Visit KServe →Scalable model-serving library built on Ray, for composing models and business logic in Python.
Best for: Python model pipelines at scale.
Visit Ray Serve →Framework and platform for packaging models as APIs and deploying them to production.
Best for: Packaging and shipping model services.
Visit BentoML →NVIDIA's prebuilt inference microservices that package optimised models for deployment on NVIDIA GPUs.
Best for: Quick, optimised deployments on NVIDIA GPUs.
Visit NVIDIA NIM →Open-source KV-cache layer that reuses and shares computed attention state across requests and servers.
Best for: Cutting repeated prefill work.
Visit LMCache →Match your requirement to a starting point, then compare vendors.
| Try open models fast, pay per use | A managed platform such as Together AI, Fireworks, DeepInfra or Baseten |
|---|---|
| Lowest latency per token | A custom-silicon cloud such as Groq, Cerebras or SambaNova |
| Stay within your cloud and governance | Bedrock, Microsoft Foundry or Google Cloud's platform |
| Own your GPUs and costs | A GPU cloud plus an engine such as vLLM or SGLang |
| Run on a laptop or edge device | llama.cpp, Ollama or ONNX Runtime |
| Scale across a GPU cluster | Dynamo, llm-d, KServe or Ray Serve |
Neighbouring directories and hubs.
Quick answers.
Inference is running a trained model to produce outputs, such as answering a prompt or classifying an image. It is the production phase that follows training, and it is where most ongoing cost and latency arise.
Managed platforms are faster to start and handle scaling. Self-hosting with an engine such as vLLM gives control over data, cost and tuning but you operate the stack. Many teams start managed and move selected workloads in-house.
An inference engine runs the model efficiently on hardware. A platform packages engines with GPUs, autoscaling, APIs and billing so you do not run the infrastructure yourself.
Model size, quantisation, batching, caching, context length, hardware and traffic pattern all matter. Measure time to first token, tokens per second and cost per million tokens on your own prompts.
Prefer OpenAI-compatible APIs, keep prompts and evaluation sets portable, and consider a gateway or router so you can switch providers.