Peter's Weblogs · AI Infrastructure & Cost Engineering

Techniques for LLM Inference Optimization

"Which model is the right model for this request?" That one question, asked consistently at every layer of the stack, is worth more to your AI budget than any single pricing negotiation. Cutting LLM inference cost is rarely about finding a cheaper model — it's about designing the inference layer itself so that routing, caching, batching, and compression all pull in the same direction.

Most enterprise AI budgets don't blow up because the model is expensive per token. They blow up because every request pays the full price of a frontier model, a cold context window, and an unbatched GPU — regardless of whether the request actually needed any of that. The fix isn't a cheaper model; it's an inference layer that only spends what a given request genuinely requires. That layer is built from a dozen or so techniques that work at different points in the request path — before the model is even called, inside the serving engine, inside the model itself, and after the response ships. This piece walks through twelve of them, grouped by where in the stack they operate, with the mechanics and the real numbers behind each one.

INCOMING REQUEST Context Window Audit trim to relevant context only Semantic + Prefix Cache Check embedding match · exact KV prefix match Cache Hit → Return skips model call entirely miss Complexity-Aware Router classify task; estimate cost/quality trade-off Small / On-Prem Model routine · sensitive/local data Mid-Tier Model moderate reasoning, most traffic Frontier Model complex · ~10-15% of traffic Quantized Weights (AWQ / GPTQ / FP8) + Fine-Tuned / Distilled Variants smaller memory footprint, same or narrower task scope Serving Engine Continuous batching + PagedAttention Speculative decoding · Compressed KV cache (GQA/MLA/FP8) Async queue for latency-tolerant workloads Streamed, Length-Capped Response max_tokens / stop sequences enforced Governance & Observability quotas · budgets · rate limits spend tagged by team / app / model feeds router thresholds back Observability watches every stage and continuously retunes routing and quota thresholds — it never serves a token itself request path feedback loop — tunes, doesn't serve
A request first checks the cache layer, then a router picks the cheapest model tier capable of handling it, that model's quantized or distilled weights run through a serving engine doing continuous batching and speculative decoding, and the response streams back under a token cap. Observability watches spend at every stage and continuously retunes the router and the governance guardrails — it sits outside the request path entirely.

01Model Right-Sizing and Intelligent Routing

The single highest-leverage decision in the whole pipeline happens before a single token is generated: which model handles this request at all. Frontier models are priced for the hardest 10-15% of tasks a system will ever see, but most production traffic is routine — classification, extraction, short summarization, templated drafting — and never needed frontier-level reasoning in the first place. A router that estimates request complexity and dispatches accordingly turns that observation into savings without touching output quality on the requests that actually matter.

This isn't theoretical. RouteLLM, a widely cited open routing framework, reported 85% cost savings while retaining 95% of GPT-4-level quality on MT-Bench, sending only 14% of queries to the stronger model. Field reports from teams running tuned routers in production land in a broad 40-85% bill-reduction range, and the math is intuitive once you model it as a traffic split: routing 70% of volume to a cheap model and 30% to a frontier model (say, Haiku-class to Opus-class) nets roughly 67% savings versus sending everything to the frontier tier; an 80/20 split nets closer to 79%. The savings compound sharply once the cheap-model share crosses 50% of total volume — which is exactly where most real workloads sit once you actually measure them.

Hybrid cloud/on-prem routing is the same idea applied to a different axis: instead of routing by model size, route by data sensitivity and predictability. Workloads with stable, well-understood shapes — and anything where sending data off-premises is itself the cost (compliance risk, egress fees, latency to a regulated data residency zone) — stay on local or on-prem inference. Genuinely open-ended reasoning still goes to a cloud frontier model when the task warrants it. Both forms of routing rest on the same principle: complexity and sensitivity, not habit, should decide where a request lands.

02Serving-Layer Efficiency: Continuous Batching and PagedAttention

Even after routing picks the right model, the serving engine underneath it can waste most of a GPU's capacity if it's naive about batching. Static batching — the older approach — locks in a fixed batch and holds every finished sequence idle until the slowest sequence in that batch completes, which is brutal on GPU utilization when output lengths vary (and they always do).

Continuous batching fixes this with iteration-level scheduling: batch composition is re-evaluated every forward pass, so the instant one sequence finishes, a new one is inserted in its place. The GPU never idles waiting for a straggler. The measured gains are large: vLLM's implementation, combined with its memory-management layer, has been benchmarked at roughly 23x throughput improvement over naive static batching under high output-length variance, with text-generation-inference and Ray Serve's continuous-batching-only implementations landing around 8x, and NVIDIA FasterTransformer's optimized static batching around 4x.

PagedAttention is the memory layer that makes aggressive continuous batching possible in the first place. It treats KV-cache memory the way an operating system treats virtual memory — allocating it in fixed-size blocks ("pages") rather than one large contiguous reservation per sequence — which eliminates the fragmentation and over-provisioning that static allocation causes. The result is 95%+ effective KV-cache memory utilization versus 50-65% without it, for only a 2-5% compute overhead, and it's the foundation vLLM, SGLang, and most modern serving stacks are built on. Chunked prefill — splitting a long prompt's initial processing into smaller pieces interleaved with ongoing decode steps — is the complementary technique that keeps long-context requests from stalling the batch entirely while they're being ingested.

03Caching Strategy: Prefix, Semantic, and KV Cache

Caching shows up at three different layers of the stack, and conflating them is the most common mistake in a cost-optimization plan. Each one catches a different kind of repetition.

Prefix (prompt) caching

Prefix caching operates inside the model's attention mechanism: it stores KV-cache entries for a prompt's beginning and reuses them when a new request's prefix matches exactly, token for token. It only reduces time-to-first-token — output generation still runs, and still costs, at full price. But the TTFT effect scales with prefix length: roughly 7% improvement at a 1,024-token shared prefix, climbing to 67% at 150,000+ tokens, and around 79% TTFT reduction with a 90% cost cut on the cached portion at a 100,000-token cached prefix in benchmark testing. This is the mechanism behind the "prompt caching" discounts offered by Anthropic, OpenAI, Google, and AWS Bedrock.

Semantic caching

Semantic caching sits a layer above the model entirely. Incoming queries are converted to embeddings and compared against cached queries by cosine similarity against a configurable threshold; on a hit, the LLM is never called at all — both input and output tokens are fully saved, not just discounted. Because it matches on meaning rather than exact text, it catches paraphrased and reworded repeats that prefix caching can't touch, which is why teams report meaningfully higher hit rates and total savings in the range of 30-40%+ of LLM spend when it's tuned well. The trade-off is a small risk of a false-positive match returning a stale or subtly wrong cached answer, which is why threshold tuning and cache invalidation policy matter more here than in exact-match caching.

The two are complementary, not competing — a layered strategy runs exact-match caching first, semantic caching second, and prefix caching for whatever's left that still has to reach the model.

KV cache compression

Independent of caching strategy, the KV cache itself can be shrunk architecturally. Grouped-Query Attention (GQA) delivers 4-8x compression and is the default in Llama 3.x and Mistral; Multi-Query Attention (MQA) pushes to 32x compression at a noticeable 1-3 point quality cost; Multi-head Latent Attention (MLA), used in DeepSeek's V2 through V4 models, achieves 7-14x compression with under 0.2 points of regression. Layered on top, FP8 KV-cache quantization halves memory again for a modest 0.3-0.7 point hit. Stacked together, these compound multiplicatively — a 70B model at 1M-token context can go from 135GB of KV cache at FP16 baseline down to roughly 8GB with MLA plus FP8, a documented 4-40x reduction on long-context inference depending on which techniques are combined.

Cache layerWhat it matchesSavesTypical impact
Exact-match cacheIdentical repeated queriesInput + output tokens100% on hits, narrow coverage
Semantic cacheParaphrased / similar queries (embeddings)Input + output tokens~30-40%+ of eligible spend
Prefix / prompt cacheExact shared prompt prefixInput tokens (TTFT) onlyUp to ~90% cost on cached portion
KV cache compressionN/A — architectural, not query-levelGPU memory for any request4-40x combined (GQA/MLA + FP8)

04Compressing the Model: Quantization and Distillation

Routing and caching reduce how often and how much of the model runs. Quantization and distillation reduce the cost of running the model itself.

Quantization lowers the numeric precision of model weights (and optionally activations and the KV cache) from 16-bit down to 8-bit or 4-bit representations. The two dominant post-training methods differ in how they decide what to protect: GPTQ relies on layer-wise second-order weight optimization, which is computationally expensive and less targeted; AWQ (Activation-Aware Weight Quantization) instead profiles activation magnitudes during calibration to identify the roughly 1% of "salient" weight channels that matter most, scales those channels to preserve precision, and quantizes the rest aggressively — and it consistently outperforms GPTQ by 1-2 quality points at the same bit-width. The economics are substantial: a 70B model drops from roughly 140GB in BF16 to 35-40GB in AWQ INT4, a ~71% memory reduction, which in one benchmark comparison took serving cost from $0.757 per million tokens on dual A100s down to $0.253 per million tokens on a single A100 — a ~67% cost reduction — with throughput improving roughly 1.5x on top of that.

Distillation and task-specific fine-tuning attack the same problem from a different angle: instead of shrinking a general-purpose model's weights, train a much smaller student model to reproduce a larger teacher's behavior on a narrower task. This is the right call precisely when prompts are long, repetitive, and narrow in scope — sending the same lengthy context and instructions on every call is often more expensive over time than the one-time cost of fine-tuning a small model to do that one job without the context at all. The trade-off is scope: a distilled or fine-tuned model is deliberately worse at everything outside its training distribution, which is why it belongs behind a router, not in front of general traffic.

Compression needs an eval gate, not just a benchmark run

Quality regressions from quantization and distillation are measured in fractions of a point on public benchmarks, but those benchmarks rarely match a specific production task's distribution. Before shipping a quantized or distilled model into the routing tier, run it against a held-out sample of your own real traffic — not just MMLU or MT-Bench — and set an explicit tolerance for acceptable degradation before it goes live.

05Accelerating Decode: Speculative Decoding, Async, and Streaming

Everything above reduces how much compute a request costs. This section is about latency — and, indirectly, cost, since lower latency means fewer retries and less GPU time held open per request.

Speculative decoding pairs a small, fast "draft" model with the full target model. The draft model proposes several candidate tokens ahead (typically 3-12) in the time it would take the target model to produce one; the target model then verifies all of them in a single forward pass, using rejection sampling to keep only the tokens it would have generated on its own and discard the rest. Because rejected tokens are simply regenerated by the target model itself, the final output is provably identical in distribution to standard autoregressive decoding — there is no quality trade-off, only a latency one. A representative example: generating three tokens autoregressively at 200ms each takes 600ms; with speculative decoding verifying a draft batch in one pass, the same three tokens land in around 250ms. Reported end-to-end speedups across production deployments generally fall in the 2-3x range.

Asynchronous inference is the simpler lever: not every workload needs a synchronous, sub-second response. Batch scoring, overnight report generation, and background enrichment jobs can be queued and processed during off-peak capacity, which both improves GPU utilization and often qualifies for lower off-peak or batch API pricing where providers offer it.

Streaming responses address perceived latency rather than actual compute cost: delivering tokens as they're generated, rather than waiting for the full response, keeps a user or downstream system engaged instead of timing out or retrying — and retries are pure wasted spend, since a timed-out request that gets resent pays for the discarded generation too.

06Context Engineering: Auditing What You Actually Send

Every token in the context window is compute you're paying for whether or not it helps the answer — and, on current transformer architectures, attention cost scales with context length, so a bloated context window is a compounding cost, not a flat one. Context window auditing means treating retrieval and prompt assembly as a relevance-filtering problem: pull in what a request needs, not everything a retrieval system could plausibly return. This isn't only a cost argument — models are also demonstrably worse at using information buried in the middle of an very long context than information near the start or end, so an aggressively padded prompt can simultaneously cost more and perform worse.

Output token controls close the other end of the same problem. An unbounded max_tokens setting, combined with a model that tends to over-explain, restate, or hedge, can quietly become one of the largest line items in an inference bill — and it's one of the cheapest to fix, since explicit output limits and stop sequences cost nothing to implement and directly cap the most expensive part of a request (output tokens are typically priced several times higher than input tokens across most providers).

07Governance and Observability: Guardrails Before the Bill Arrives

Every technique above is a lever inside the request path. This last pair operates outside it, and it's the layer that keeps the other eleven from silently drifting back toward waste.

Usage governance — quotas, per-team and per-application budgets, rate limits, and hard guardrails on which model tiers a given caller can even reach — has to exist before a runaway integration or an unbounded agent loop turns into a five-figure surprise, not after. A router that's technically capable of sending everything to a frontier model will do exactly that if nothing stops it, whether the cause is a misconfigured default or a legitimate but unbounded batch job.

Cost and usage observability is what makes every other technique in this piece measurable rather than assumed. Tracking spend by application, feature, team, workflow, model, and user segment turns "our AI costs went up" into "feature X's fallback path is calling the frontier model on cache misses at 3x the volume we expected" — an actionable finding instead of a vague budget alarm. This is also the feedback loop that closes back into routing and caching: a router's complexity thresholds and a semantic cache's similarity threshold are not "set once" parameters — they need retuning as traffic mix shifts, and that retuning is only possible with granular, current cost data to look at.

Optimization without measurement is guesswork

It's tempting to treat this list as a one-time implementation checklist. In practice, routing thresholds, cache hit rates, and quantization quality all drift as traffic composition and model versions change. Observability isn't the twelfth technique on the list — it's the mechanism that tells you whether the other eleven are still working six months from now.

08Putting It Together

None of these twelve techniques is a substitute for the others, and the diagram earlier in this piece is really the argument in full: a request should be checked against a cache before it ever reaches a model, routed to the cheapest tier actually capable of handling it, served by an engine that keeps the GPU saturated and the KV cache compressed, generated by weights that are no larger or slower than the task requires, and capped and streamed back efficiently — all while an observability layer watches spend closely enough to keep tightening every threshold in that chain. Skip any one layer and the others compensate for it by burning more of the layers around them; a brilliant router sending every request to an unquantized model on a statically-batched server still wastes most of its GPU capacity, just as the best serving engine in the world can't fix a context window stuffed with irrelevant retrieval results.

The framing this piece opened with is worth repeating as the close, because the twelve techniques above are really just its implementation: LLM cost optimization is an architecture problem, not a model-pricing problem. The goal was never simply cheaper inference — it's intelligent routing, efficient context, appropriately sized models, strong observability, and predictable economics, working together so that AI systems can scale without their infrastructure cost scaling uncontrollably alongside them.