What LLMOps is
LLMOps is what it takes to keep an application built on a large language model working once real users arrive. It borrows the loop of DevOps and MLOps — build, test, release, monitor, feed back — and adapts it to a different kind of artifact.
In most LLM applications you do not train the core model. What you ship and maintain is a pipeline: a chosen model, prompts, retrieved context, tools and guardrails. That shifts the work from training runs towards prompt and context quality, evaluation of open-ended text, token cost and latency, and new failure modes such as hallucination and prompt injection.
The artifact changes
Instead of a trained model you own, you manage a prompt-and-context pipeline around a foundation model.
Quality is harder to score
Free-form answers need purpose-built evaluations, not just accuracy against a label.
Cost scales with usage
Tokens, context length and model choice drive spend, so cost control is a first-class concern.
The LLM application lifecycle
Nine stages cover most of what teams operate. You rarely need all of them on day one, but each one is a place where production systems tend to fail when it is skipped.
Model selection & access
Choose between frontier, open-weight and small models against your own tasks, latency budget, data-residency needs and licence terms. Keep the choice swappable so a better or cheaper model is a configuration change.
Prompt & context management
Treat prompts as versioned artifacts: templates, variables, review, rollback and A/B comparison. Context assembly — what you put in the window and in what order — matters as much as the wording.
Retrieval & data (RAG)
Ground answers in your own content: chunk documents, embed them, retrieve, rerank and keep the index fresh. Most quality problems in RAG systems trace back to retrieval, not the model.
Adaptation & fine-tuning
When prompting and retrieval are not enough, adapt the model with supervised fine-tuning or preference methods. Track datasets, hyperparameters and the resulting model versions like any other release.
Evaluation & testing
Build a golden set of real cases, score outputs with rules, reference answers and LLM judges, and run it on every change like a regression suite. Add adversarial tests for jailbreaks and unsafe output before release.
Serving & cost control
Pick an inference engine, quantize where quality allows, batch and cache aggressively, and route easy requests to cheaper models. Watch time-to-first-token and throughput, not just average latency.
Gateway & routing
Put a single endpoint in front of providers for key management, budgets, rate limits, fallbacks and a uniform API. It is also the natural place to log traffic and enforce policy.
Observability & feedback
Trace every request through retrieval, prompts and model calls, and record tokens, cost, latency and user feedback. Quality drifts silently as models, data and prompts change, so monitor it continuously.
Safety, security & governance
Validate inputs and outputs, filter PII, defend against prompt injection, and keep audit logs. Map your controls to recognised guidance rather than inventing a taxonomy.
LLMOps compared with MLOps
| Dimension | MLOps | LLMOps |
|---|---|---|
| Primary artifact | A model you train and register | A prompt-and-context pipeline around a foundation model |
| Core loop | Data → train → validate → deploy → monitor | Prompt/context → evaluate → deploy → trace → refine |
| Quality check | Metrics against labelled data | Rubrics, reference answers and LLM judges on open-ended output |
| Main cost driver | Training and serving compute | Tokens, context length and model choice per request |
| Typical failures | Data drift, model drift | Hallucination, prompt injection, retrieval misses, cost blow-ups |
| Change cadence | Retrain on a schedule or trigger | Prompt, model or index changes can ship daily |
A pragmatic starting checklist
- Keep prompts and configuration in version control, separate from application code where possible.
- Build a golden evaluation set from real traffic and run it before every prompt, model or retrieval change.
- Route all model calls through one gateway or client layer so provider and model are swappable.
- Trace every request end to end, with token counts, cost and latency attached.
- Set per-user and per-feature budgets and alerts before launch, not after the first bill.
- Add input and output guardrails, and test them with adversarial prompts.
- Collect user feedback and feed failing cases back into the evaluation set.
- Record which model, prompt and index version produced each answer, for audit and rollback.
Risks to plan for
Hallucination
Confident but wrong output; mitigate with grounding, verification and confidence gating.
Prompt injection
Untrusted text in context can redirect the model; treat retrieved and user content as untrusted.
Cost overruns
Long contexts, retries and agent loops multiply spend; cap and monitor them.
Silent quality drift
Provider model updates and changing data can degrade results without any deploy.
Data exposure
Prompts and logs can contain personal or confidential data; control retention and access.