Technical Deep Dive · AI Infrastructure

Inside the Production AI Inference Stack

Serving one model on one GPU is a solved problem. Serving hundreds of models across thousands of GPUs, with predictable latency, is the infrastructure discipline now forming around vLLM, llm-d, and Kubernetes — and it changes what "infrastructure" even means.

For a long time, the infrastructure question under any AI workload was simple: how many GPUs do we have? That question made sense when one model fit on one node. It stops making sense the moment a fleet is running dozens of models, thousands of concurrent requests, and a latency budget measured in milliseconds — because at that point, the GPU count tells you almost nothing about what the fleet can actually deliver.

A new stack has formed to answer the better question — how much useful inference can we extract from every GPU — and it has its own components, its own vocabulary, and its own failure modes. Here's how it fits together, from the application down to the silicon.

In this article
  1. The stack — eight layers between an application and a GPU.
  2. Why vLLM — the engine every layer above it depends on.
  3. When one node isn't enough — where distributed inference begins.
  4. llm-d — LLM-aware distributed serving for Kubernetes.
  5. Prefill vs. decode — the split that changes everything.
  6. New metrics — what replaces CPU and RAM.
  7. The next stack — where this fits in infrastructure history.
  8. The takeaway

01 · Architecture

The stack

A single inference request passes through eight distinct layers before a token comes back — each one solving a problem the layer below it can't solve on its own.

AI Applications agents, chat UIs, batch pipelines Inference Gateway auth, rate limits, model routing entry point llm-d Router prefix-cache-aware, latency-aware request routing Prefill / Decode Workers disaggregated compute-bound vs. memory-bound pools vLLM continuous batching, PagedAttention, tensor parallelism GPU Pools heterogeneous accelerator capacity High-Speed Fabric NVLink / InfiniBand / RDMA interconnect Kubernetes scheduling, autoscaling, lifecycle management
Highlighted layers (llm-d Router, Prefill/Decode Workers) are where distributed-inference-specific logic lives; everything below vLLM is shared with ordinary cloud-native workloads, and everything above it is where an application's request actually enters the stack. Generated tokens stream back up this same path.

02 · Single-Node Engine

Why vLLM

Before a fleet exists, there's a single serving engine, and vLLM has become the default one for a reason: it turns a GPU's raw FLOPs into throughput that a production API can actually rely on. Originating at UC Berkeley's Sky Computing Lab and now maintained by a community of thousands of contributors under an Apache-2.0 license, it improves model serving through six capabilities that matter well beyond a single node:

  • 01Continuous batching — new requests join a running batch as soon as a GPU slot frees up, instead of waiting for the whole batch to finish, which is what keeps GPU utilization high under bursty traffic.
  • 02PagedAttention — manages the KV cache in fixed-size pages, the way an OS manages virtual memory, instead of one contiguous block per request — which is what makes memory-efficient, high-concurrency serving possible at all.
  • 03KV-cache management — reuses and evicts cached attention state intelligently, so a shared system prompt or a multi-turn conversation doesn't recompute what it already computed.
  • 04Tensor parallelism — splits a single model's weights and computation across multiple GPUs, for models too large to fit in one accelerator's memory.
  • 05OpenAI-compatible APIs — serves an endpoint that drops into existing client code and tooling built against that API shape, with no rewrite required.
  • 06High-throughput inference — the combined effect of the five techniques above: more tokens per second, per GPU, than a naive serving loop achieves.

03 · Where It Breaks

When one node isn't enough

Everything in that list is a single-node, or single-model-across-a-few-GPUs, concern. It's a different problem once an organization is operating tens, hundreds, or thousands of GPUs at once: which of dozens of models should handle this request, which replica has this request's system prompt already cached, what happens when a node fails mid-generation, and how does the fleet scale a specific model up in the ninety seconds before a traffic spike turns into dropped requests.

None of that is a serving-engine problem anymore. It's a distributed-systems problem — and it's exactly the problem Kubernetes was built to solve for stateless microservices, except LLM inference is neither stateless (the KV cache is state, and it's expensive to rebuild) nor uniform (a 200-token chat reply and a 4,000-token document summary cost wildly different amounts of compute). That mismatch is what llm-d exists to close.

04 · Fleet Coordination

llm-d: LLM-aware distributed serving for Kubernetes

llm-d is an open-source, Kubernetes-native project — a CNCF Sandbox project founded by Red Hat, Google, and IBM — built on the observation that the disaggregation and KV-cache techniques behind the fastest inference systems (the kind DeepSeek and similar teams built in-house) were out of reach for most ML platform teams, simply because building and operating them from scratch is its own multi-year engineering project. llm-d packages those techniques as Kubernetes-native primitives instead, running vLLM (and other engines) across the cluster rather than replacing it. It brings LLM-aware distributed serving concepts to Kubernetes, including:

  • 🔐Intelligent request routing — sends each request to the replica best positioned to serve it, not just the least-busy one.
  • 🔐Prefix-cache-aware scheduling — routes a request to a replica that already has its shared prompt prefix cached, avoiding redundant computation across the fleet.
  • 🔐Prefill/Decode disaggregation — runs the two phases of generation on independently sized, independently optimized worker pools (more on this below).
  • 🔐Expert parallelism for MoE — distributes a mixture-of-experts model's experts across the fleet, so only the parameters a token actually needs are activated and transferred.
  • 🔐Autoscaling — grows and shrinks each worker pool independently, based on the signals that actually predict inference load, not generic CPU utilization.
  • 🔐Distributed KV-cache strategies — extends cache reuse beyond a single node, across a tiered cache that spans GPU memory, host memory, and fast local storage.

05 · The Core Split

Prefill vs. decode: the split that changes everything

Generating a response happens in two phases with almost opposite hardware profiles, and treating them as one workload is where a lot of wasted GPU capacity hides.

Prefill

Process the prompt→compute intensive→build KV cache

Decode

Generate output tokens→memory / bandwidth sensitive

Prefill reads the whole prompt at once and can saturate a GPU's compute with large matrix multiplies — it's a job that benefits from batching several prompts together. Decode generates one token at a time, and each step is bottlenecked by how fast the GPU can read the growing KV cache from memory, not by how many FLOPs it has — it benefits from more replicas and more memory bandwidth, not more raw compute.

Run both phases on the same worker pool, and a long prefill (a big document, a long system prompt) will stall the in-flight decode steps of every other request sharing that GPU — a form of head-of-line blocking that shows up to users as a sudden latency spike with no obvious cause. Disaggregating the two phases onto separate, independently scaled pools — a prefill pool sized for burst compute, a decode pool sized for sustained memory bandwidth — removes that interference and lets each phase scale against the metric that actually constrains it.

06 · A New Vocabulary

New metrics for a new kind of infrastructure

Traditional infrastructure capacity planning revolved around four resources everyone already had intuitions about. AI infrastructure keeps those and adds a second set that most infrastructure teams have never had to reason about before:

Traditional infrastructure

  • CPU
  • RAM
  • Storage
  • Network

AI infrastructure adds

  • GPU
  • Model Weights
  • KV Cache
  • Tokens / sec
  • TTFT
  • Inter-token Latency

TTFT (time-to-first-token) is the delay between a request arriving and the first generated token coming back — dominated by the prefill phase, and the number a user actually feels as "how fast did it start responding." Inter-token latency is the gap between each subsequent token — dominated by decode, and the number that determines how fast the rest of the response streams in. Neither shows up in a CPU/RAM dashboard, and neither correlates cleanly with raw GPU count.

That's the actual shift: the question stops being "how many GPUs do we have?" and becomes "how much useful inference can we extract from every GPU?" — a question about goodput (throughput that actually meets a latency target), not just throughput.

07 · Where This Fits

The stack that's replacing the last decade's stack

This isn't the first time infrastructure has re-organized itself around a new abstraction layer. The last decade's progression abstracted compute, then packaging, then orchestration, then the application boundary itself:

Last decade
VMs→ Containers→ Kubernetes→ Microservices
Emerging now
GPUs→ Model Servers→ Distributed Inference→ AI Agents

Kubernetes was the abstraction that let a microservices architecture exist without every team reimplementing scheduling and failover by hand. llm-d is aiming at the same role one layer up: the abstraction that lets an agent architecture exist without every team reimplementing prefix-cache routing and prefill/decode disaggregation by hand.

08 · Summary

The takeaway

GPU runs the model.

vLLM serves the model.

llm-d coordinates the fleet.

Kubernetes operates the platform.

AI isn't eliminating infrastructure engineering. It's creating a new infrastructure discipline — one measured in tokens per second and milliseconds to first token, not just GPU count.