Peter's Weblogs · LLM Deployment & Quantization

GGUF vs GPTQ vs AWQ vs EXL2: How LLM Model Formats Actually Differ

Download a quantized model and you'll find a bewildering pile of filenames — Q4_K_M, 4bit-128g, 4.65bpw — that all claim to shrink the same checkpoint. They're not interchangeable, and they're not even solving the same problem: some are containers that package a whole model, others are methods that decide which bits to keep. Mixing the two up is the single most common reason people pick the wrong format for their hardware.

Technical Guide LLM inference & quantization ~10 min read

Every quantized LLM you'll run locally or in production owes its file format to one of a handful of projects, each solving a different piece of the same problem: a 70-billion-parameter model in 16-bit precision needs roughly 140GB of memory, which is out of reach for almost every consumer GPU and most single-node servers. Quantization closes that gap by storing weights in fewer bits — but how a project decides which weights can lose precision without wrecking output quality, and where that quantized model is allowed to run, is what actually separates GGUF, GPTQ, AWQ, and EXL2 from one another. This piece walks through each one, what it optimizes for, and which format fits which hardware.

01The Distinction That Decides Everything: Container vs. Method

The four names in this piece aren't really four versions of the same thing. GGUF is primarily a container format — a single self-describing file that bundles quantized weights, tokenizer, chat template, and metadata together, built for llama.cpp's CPU-first, portable execution model. GPTQ and AWQ are quantization methods — algorithms that decide how to round 16-bit weights down to 3-4 bits with minimal accuracy loss — and the resulting weights are typically shipped as ordinary safetensors checkpoints, meant to be loaded straight into a GPU serving engine. EXL2 (and its successor EXL3) is both at once: a quantization method and a native file layout, built specifically for ExLlamaV2/V3's own runtime and no one else's.

That distinction is what should actually drive your format choice, more than any single benchmark number: are you memory-constrained on a CPU, an Apple Silicon Mac, or a GPU with too little VRAM for the full model (reach for GGUF)? Are you serving GPU traffic at scale through a framework like vLLM or SGLang (reach for AWQ or GPTQ)? Or are you a single user on a consumer NVIDIA card trying to squeeze out every token per second at a bitrate that fits your exact VRAM budget (reach for EXL2/EXL3)?

FP16 / BF16 Checkpoint ~2 bytes per weight, full precision CONTAINER PATH GGUF K-quants / I-quants, embeds tokenizer QUANTIZATION-METHOD PATH GPTQ / AWQ / EXL2 / EXL3 safetensors or native layout, 2-4 bit llama.cpp, Ollama, LM Studio CPU · Apple Silicon · split GPU/CPU vLLM, SGLang, ExLlamaV2/V3 cloud GPU serving · consumer NVIDIA GPU Memory-constrained inference Throughput / tokens-per-second
The same FP16 checkpoint forks into two philosophies: GGUF packages everything a CPU-class runtime needs into one portable container, while GPTQ, AWQ, and EXL2/EXL3 quantize weights for GPU serving engines built for throughput.

02GGUF: The Portable Container for CPU and Edge

GGUF (GPT-Generated Unified Format) is llama.cpp creator Georgi Gerganov's successor to the earlier GGML format, and it solves a problem the other three formats don't even attempt: running a model somewhere that isn't a data-center GPU. A GGUF file is self-contained — tokenizer, special tokens, chat template, and quantized tensors all live in one file with typed key-value metadata, so a runtime can load it without hunting down a separate config or tokenizer repository.

What makes GGUF distinctive is its menu of quantization types rather than a single method: legacy schemes like Q4_0 and Q8_0, the newer "K-quants" (Q2_K through Q6_K, spanning roughly 2.6 to 6.6 bits per weight), and "I-quants" (IQ1_S through IQ4_XS) that push down toward 1.5-4.25 bits for extreme memory savings. An optional importance-matrix calibration step, run through llama-imatrix, measurably improves quality at the lowest bit-widths by weighting which activations matter most during rounding. On Llama-2-7B, the commonly cited quality/size tradeoffs land around Q4_K_M at 4.1GB with roughly a 1.7% perplexity increase over FP16, improving to under 0.15% at Q6_K for about 5.5GB — a useful reference point for picking a quant level rather than guessing.

GGUF's real advantage is where it runs: llama.cpp, Ollama, LM Studio, and GPT4All all treat it as a first-class format, and it's the only one of the four that comfortably supports splitting a model across CPU RAM and whatever GPU VRAM is actually available — which matters enormously for anyone running a model larger than their card. vLLM support exists but is still described by its own maintainers as experimental, so GGUF is not the format to reach for if high-throughput GPU serving is the actual goal.

03GPTQ: The Original Post-Training GPU Quantization Method

GPTQ predates the other three by nearly a year — published in October 2022 by researchers at IST Austria and ETH Zurich, and accepted at ICLR 2023 — and it was the method that first made 3-4 bit post-training quantization practical at scale, reportedly quantizing a 175-billion-parameter model in around four GPU-hours. Its core idea is layer-by-layer second-order (Hessian-based) error correction: after rounding one weight, GPTQ adjusts the remaining weights in that layer to compensate for the rounding error just introduced, propagating the correction across the row instead of letting error accumulate unchecked. No retraining is involved — it's entirely a post-training, calibration-data-driven process, typically taking on the order of twenty minutes for an 8B model on a single A100.

Two configuration choices matter in practice: group size (how many weights share one quantization scale — smaller groups such as 128g trade a little extra metadata for better accuracy) and act-order (reordering columns by importance before quantizing, which generally improves results at negligible cost). Reported throughput gains from GPTQ-quantized inference run around 3.25x on an A100 and 4.5x on an A6000 versus unquantized serving. The tooling around GPTQ has shifted since its original release — the once-standard AutoGPTQ library is no longer actively maintained, with GPTQModel and llm-compressor now the more current paths — but the output format itself remains widely supported across vLLM, SGLang, and Hugging Face Transformers.

04AWQ: Protecting the 1% of Weights That Matter Most

AWQ (Activation-Aware Weight Quantization) came out of MIT's Song Han lab in mid-2023 and won the MLSys 2024 Best Paper Award, and its central insight is a departure from GPTQ's approach: instead of deciding which weights to protect by inspecting the weights themselves, AWQ watches activations during a short calibration pass to find the roughly 1% of "salient" weight channels that have an outsized effect on output quality. Those channels are protected through a mathematical scaling transformation — not by giving them extra bits, which would break the uniform, hardware-friendly format the rest of the tensor uses — while everything else is quantized aggressively. There's no backpropagation involved, which is part of why AWQ's calibration is markedly faster: roughly ten minutes for an 8B model on one A100, about half of GPTQ's calibration time.

In its original benchmarks, AWQ reported serving speedups of over 3x versus the Hugging Face FP16 baseline across both desktop and mobile GPU targets, and in head-to-head comparisons at matched bit-widths it tends to edge out GPTQ on output quality. Like GPTQ, its original reference library (AutoAWQ) is now deprecated in favor of vLLM's llm-compressor pipeline, though AWQ checkpoints remain directly loadable in vLLM and SGLang — and, notably, AWQ also has support in Apple's MLX-LM, making it one of the few methods that spans both cloud GPU serving and Apple Silicon.

05EXL2 and EXL3: Bits-Per-Weight Dialed to the Decimal

EXL2 is ExLlamaV2's own quantization format, built by developer turboderp specifically for single-user, consumer-NVIDIA-GPU inference where every gigabyte of VRAM is precious. Technically, it reuses GPTQ's same layer-wise optimization approach under the hood, but adds a distinctive twist: instead of picking one discrete bit-width for the whole model, EXL2 measures quantization error against calibration data per column and mixes bit-widths within a single tensor, spending more bits on columns that need it. The result is a file named by its exact average bitrate — 4.65bpw, for instance — rather than a rounded bucket like "4-bit," which lets a user dial a quantized model to fit an exact VRAM budget rather than jumping between coarse presets. The tradeoff is portability: EXL2's tensor layout is specific enough to ExLlamaV2 that it's genuinely difficult to load anywhere else, so it lives almost entirely inside the ExLlamaV2 / TabbyAPI ecosystem.

EXL3, the newer successor, builds on QTIP (trellis-coded quantization with incoherence processing) from Cornell RelaxML rather than GPTQ's method, computing Hessians on the fly during conversion instead of requiring a separate calibration run. The headline result is extreme low-bit coherence: a Llama-3.1-70B model reportedly stays coherent down to around 1.6 bits per weight with a 3-bit output layer, fitting under 16GB of VRAM — territory where most other quantization methods produce unusable output. EXL3 also adds 2-8 bit KV-cache quantization, tensor and expert parallelism, speculative decoding, and CPU offloading for mixture-of-experts models, though it currently requires CUDA 12.4+ and lists ROCm support as a work in progress.

FormatTypeCalibrationBest hardwarePrimary runtime
GGUFContainer + quant typesOptional (imatrix)CPU, Apple Silicon, split VRAMllama.cpp, Ollama, LM Studio
GPTQQuantization methodRequired (~20 min / 8B)Cloud / data-center GPUsvLLM, SGLang, Transformers
AWQQuantization methodRequired (~10 min / 8B)Cloud GPUs, Apple Silicon (MLX)vLLM, SGLang, MLX-LM
EXL2Method + native layoutRequiredConsumer NVIDIA GPUsExLlamaV2, TabbyAPI
EXL3Method + native layoutBuilt-in (on-the-fly)Consumer NVIDIA GPUs, extreme compressionExLlamaV3, TabbyAPI

Calibration cost is a real deployment variable, not a footnote

GPTQ, AWQ, and EXL2/EXL3 all need a calibration pass over representative data before the quantized weights exist at all — GGUF is the outlier that can skip it. If you're re-quantizing frequently (fine-tuning a model weekly, say), AWQ's roughly ten-minute calibration versus GPTQ's twenty adds up fast across a release cadence. If you're downloading a pre-quantized checkpoint someone else calibrated, this cost is invisible to you — but it's worth knowing which format made that tradeoff, especially if you ever need to quantize your own fine-tune.

06How to Choose: A Hardware-First Decision Path

Every recommendation in this space collapses to the same first question: what hardware is actually running inference?

  • CPU-only, Apple Silicon, or a model bigger than your VRAM: start with GGUF at Q4_K_M, and step up to Q5_K_M or Q6_K if you have memory headroom to spare — the perplexity gains taper quickly past that point.
  • High-throughput cloud or data-center serving: AWQ or GPTQ through vLLM or SGLang are the well-trodden path; on Hopper- or Blackwell-class cards, FP8 is increasingly the more attractive option ahead of either.
  • Single-user, consumer NVIDIA GPU, maximizing tokens per second: EXL2 or EXL3 via TabbyAPI, dialing the bits-per-weight to exactly what your VRAM allows.
  • Budget fine-tuning rather than pure inference: bitsandbytes NF4 (from the QLoRA paper) needs no calibration step at all and is the de facto standard for parameter-efficient fine-tuning workflows.
  • Apple-native Python workflows: MLX-LM's own quantization path, built specifically for Apple Silicon's unified memory architecture.

None of these choices are permanent or mutually exclusive — it's common to keep a GGUF copy for local experimentation and an AWQ or GPTQ copy of the same fine-tune for production serving. What matters is not defaulting to whichever format you happened to download first, since the wrong one can leave real throughput or memory headroom on the table for no benefit.

A quant level is a quality tradeoff, not a free lunch

Every format above trades some accuracy for size and speed — the difference is only in how gracefully. Below roughly 3 bits per weight, quality degradation accelerates sharply for most methods unless they were specifically engineered for that regime (as EXL3's trellis coding was). Before shipping a quantized model into production, evaluate it against your own held-out task data, not just a public perplexity number — the number that looks fine on a benchmark can still be the wrong call for your specific workload.

07Putting It Together

GGUF, GPTQ, AWQ, and EXL2/EXL3 aren't competing for the same job — they're answers to different constraints. GGUF answers "how do I run this somewhere without a data-center GPU." GPTQ and AWQ answer "how do I serve this efficiently at scale on GPUs I already have," with AWQ generally winning on calibration speed and quality at matched bit-widths. EXL2 and EXL3 answer "how do I get the absolute most tokens per second out of one consumer GPU," at the cost of being locked to their own runtime. Picking correctly starts with naming your actual hardware constraint out loud, not with copying whichever quant someone else recommended on a forum thread for a completely different setup.

Further reading and grounding for the details referenced above: