Every technique for reducing LLM hallucination does exactly one of four things to the system: it changes what the model has access to, it changes what shape the model is allowed to answer in, it changes the model's own parameters, or it checks the model's output after the fact and intervenes if something looks wrong. Once you sort by what a mechanism changes rather than when you'd reach for it, eight seemingly separate techniques collapse into four families — and the family a technique belongs to tells you its real tradeoffs far better than its name does.
This is a survey of that taxonomy: four mechanism classes, what each one costs, and which of eight common techniques falls into which.
Grounding
Gives the model real, external information to work from.
Constraint
Narrows what shape or content the model's output is allowed to take.
Verification
Checks a draft answer after it's written, before it reaches the user.
Adaptation
Changes the model's own parameters, offline, ahead of any request.
Grounding — feeding the model real information
A model hallucinates most confidently in the gap between what it memorized during training and what a specific question actually needs. Grounding mechanisms close that gap by handing the model verifiable, external material at the moment it's needed — documents, passages, or live tool results — so the model is reporting on evidence in front of it rather than reconstructing an answer from parametric memory.
Retrieval-Augmented Generation
Ground responses in trusted documents, databases, and enterprise knowledge instead of model memory alone, by retrieving relevant passages at query time via embedding search and inserting them into the prompt as context.
Mechanism of action: substitutes retrieval for recall, so an answer is reconstructed from a passage rather than pattern-matched from training data.
Chunking + Reranking
Retrieve smaller passages, then prioritize the context most relevant to the actual question — retrieving wide with embedding search, then reordering the top candidates with a cross-encoder reranker before the model ever sees them.
Mechanism of action: raises the quality of what grounding actually supplies, since retrieval alone tends to surface passages that are topically similar but not always the ones that answer the specific question.
Tool Calling
Let the model query search engines, databases, calculators, or APIs when facts need external validation — arithmetic, live prices, current status — instead of generating a plausible-looking number from memory.
Mechanism of action: grounding extended into the generation step itself, for facts that don't exist in any retrievable document because they're computed or live.
Constraint — narrowing what the model is allowed to say
Grounding gives a model better material; constraint mechanisms restrict what it's permitted to do with any material at all. An ambiguous instruction, or an unconstrained output format, leaves room the model will fill with something plausible rather than something true. These mechanisms close that room off before generation even completes.
Prompt Grounding
Use clear instructions, constraints, and few-shot examples to reduce ambiguity and unsupported claims — explicitly naming the source of truth, giving the model permission to say "not found," and showing the exact answer shape wanted.
Mechanism of action: removes the invitation to guess, by making silence or refusal an explicitly acceptable output rather than an implicit failure.
Structured Output
Use JSON schemas, templates, and validators to make responses predictable and easier to verify — for example, requiring a citation field that must reference an actual retrieved chunk ID, so a fabricated one fails validation instead of shipping.
Mechanism of action: turns a claim's shape into a checkpoint, so an invented or malformed answer is mechanically rejectable rather than only stylistically odd.
Verification — catching mistakes after they're written
Grounding and constraint both act before or during generation; verification mechanisms accept that some hallucinations will still get through and catch them afterward, before a person ever sees the result. They cost real latency and inference budget, which is why they tend to be applied selectively rather than to every request.
Advanced RAG
Add fact-checking, verification, or correction steps before the final answer reaches the user — a second model pass that checks each claim in a draft against the specific passages it cited, flagging or rewriting anything unsupported.
Mechanism of action: catches drift introduced during generation itself — a merged detail, a softened qualifier — that grounding alone doesn't prevent, because the model had the right source and still stated it slightly wrong.
Confidence Gates + Human Review
Escalate uncertain, sensitive, or high-impact outputs for human approval — using a confidence signal such as self-consistency across sampled generations, token-level log-probability, or a lightweight classifier, to sort the routine cases from the ones that need a closer look.
Mechanism of action: bounds the worst case rather than preventing every error, by routing the outputs most likely to be wrong to a reviewer instead of the end user.
Adaptation — changing the model itself
Every mechanism above operates around a fixed model. Adaptation is the one mechanism that changes the model's own parameters — which makes it powerful for some problems and a poor fit for others.
Fine-Tuning on Trusted Data
Improve domain behavior and terminology using curated data, while still grounding changing facts externally — fine-tuning teaches vocabulary, format, and domain reasoning patterns well, but is a poor way to keep a model current, since every update requires a retraining cycle and the model has no way to "forget" a stale fact it was trained on.
Mechanism of action: shifts the model's default behavior for a stable class of inputs; it does not substitute for grounding on facts that change faster than the retraining cycle.
Choosing mechanisms, not techniques
Once a technique is filed under its real mechanism, the tradeoff question becomes simple: which of the four things do you need to change, and what does changing it cost you?
| Mechanism | Changes | Primary cost | Fails when… |
|---|---|---|---|
| Grounding | What the model has access to | Corpus & retrieval infrastructure | the corpus is stale, wrong, or the right passage isn't retrieved |
| Constraint | What the model is allowed to output | Low — mostly prompt/schema design | the constraint doesn't cover the failure mode it wasn't designed for |
| Verification | Whether a draft answer ships as-is | Latency & inference cost, applied per-request | it's skipped on cost grounds for the request that needed it most |
| Adaptation | The model's own parameters | Training time, data curation, redeploy cycles | it's asked to keep up with facts that change faster than retraining |
None of the four mechanisms is sufficient alone, and none of them is free. A production system that's serious about hallucination usually ends up running at least one mechanism from each family — not because more is always better, but because each family catches a category of failure the other three structurally cannot.