AI Infra Notes Field notes on AI systems
Data Infrastructure

Vector Databases for Enterprise AI

Embeddings turned "search" into "meaning." Vector databases are the infrastructure that makes that meaning searchable at production scale — and in 2026, getting that layer right has become one of the harder engineering decisions in enterprise AI.

Read time ~12 min Level Practitioner Topics Embeddings · ANN indexing · Hybrid search · RAG infrastructure

Most teams meet the term "vector database" for the first time somewhere around their second retrieval-augmented generation (RAG) prototype, right after the first one breaks. The pattern is familiar by now: a proof of concept works beautifully on a few thousand documents using a Python list and cosine similarity, then someone points it at the real corpus — millions of chunks, dozens of data sources, sub-second latency requirements — and the whole thing falls over. That gap between "it works in a notebook" and "it works in production" is exactly the gap vector databases exist to close.

This piece works through the mechanics first — what embeddings are, how similarity search works, and how approximate nearest-neighbor (ANN) indexing makes it fast at scale — and then turns to the part enterprise teams actually lose sleep over: choosing, operating, and governing a vector layer once agentic systems start hammering it with concurrent, high-precision queries instead of the occasional chatbot lookup.

Foundations

What a Vector Database Actually Stores

A vector database is built around a simple idea: instead of indexing data by exact values — an ID, a price, a date — it indexes data by meaning. It stores each data object (a document chunk, an image, a product description) alongside a vector embedding: a long list of floating-point numbers produced by a machine learning model that has learned to represent the object's semantic content as a position in high-dimensional space.

The reason this matters is that roughly four-fifths of enterprise data is unstructured — support tickets, contracts, transcripts, images, sensor logs — and unstructured data has never fit neatly into rows and columns. Embeddings sidestep that problem entirely: instead of trying to structure the data, you represent its meaning numerically and let geometry do the work. Objects with similar meaning end up close together in the embedding space, and objects with different meaning end up far apart — which turns "find things like this" into a nearest-neighbor problem that computers are very good at solving.

A useful mental model, and the one most introductions to the topic reach for, is color. The RGB system represents a color as three numbers — how red, how green, how blue — and two similar shades of green end up with numerically close vectors. Word and sentence embeddings work the same way, just in hundreds or thousands of dimensions instead of three: "wolf" and "dog" land close together because a language model has learned they're related, "cat" sits nearby because it shares the "pet" context, and "banana" ends up on the other side of the space entirely. The model doing the embedding matters — different embedding models carve up the space differently, so mixing embeddings from two different models in one index produces geometry that means nothing.

Distance Metrics: How "Close" Gets Measured

Once data lives as vectors, similarity becomes a distance calculation, and there are several standard ways to compute it. Picking the right one isn't a style choice — it should match whatever metric the embedding model was trained against, or the resulting rankings will be subtly wrong.

Cosine similarity
range: 0 – 2 (as distance)

Measures the angle between two vectors, ignoring their length. The default for most modern text-embedding models, because it's insensitive to magnitude — only direction, which is where the semantic signal lives, is compared.

Dot product
range: −∞ – ∞

Cosine similarity without the normalization step. When every vector is already unit length — a common indexing-time preprocessing choice — dot product gives identical rankings to cosine similarity while being cheaper to compute at query time.

Euclidean (L2)
range: 0 – ∞

Straight-line distance between two points. Intuitive, but sensitive to vector magnitude — two vectors pointing the same direction with different lengths can register as dissimilar.

Manhattan (L1)
range: 0 – ∞

Sums the absolute differences across each dimension rather than the straight-line distance. Less sensitive to large outliers in any single dimension than L2.

Hamming
count of differing dimensions

Counts how many dimensions differ outright. Used almost exclusively with binary or hashed embeddings rather than dense float vectors.

The Indexing Problem

Why Brute Force Breaks, and What ANN Indexing Does About It

The naive way to find the closest vectors to a query is exhaustive: compute the distance between the query and every single vector in the database, then sort. This is the k-Nearest Neighbors (kNN) approach, and it's exact — it always finds the true nearest neighbors. It's also unusable at scale. A modern embedding model typically produces 768- to 1,536-dimensional vectors, and comparing a query against a million of them means well over a billion floating-point operations for a single search. At tens or hundreds of millions of vectors — not an unusual size for an enterprise document corpus — brute-force search simply cannot hit interactive latency.

Approximate nearest-neighbor (ANN) search is the fix, and the trade it makes is deliberate: give up the guarantee of finding the mathematically exact top-k results in exchange for a large speedup, on the observation that "the 9th and 11th closest matches" are rarely distinguishable to whatever is consuming the results — an LLM doing retrieval augmentation, a recommender, a search UI. Instead of scanning everything, ANN algorithms pre-organize the vector space — into graphs, clusters, or compressed codes — so a query only has to examine a small fraction of the data to find results that are almost certainly correct. Three families of ANN algorithm show up in essentially every production vector database.

Graph-based

HNSW

Hierarchical Navigable Small World. Builds a multi-layer graph — a sparse top layer for long jumps, denser lower layers for fine-grained navigation. A query greedily walks from the top layer down, narrowing in on the neighborhood that contains its true nearest neighbors.

ComplexityO(log N)
Trade-offSpeed & recall, at memory cost
Cluster-based

IVF

Inverted File Index. Runs a clustering pass (typically k-means) over the vector space ahead of time. At query time, only the handful of clusters nearest the query vector get searched — a typical configuration scans roughly 1% of the data for well over 90% recall.

Complexity~O(√N)
Trade-offLower memory, needs cluster tuning
Compression-based

Product Quantization

Splits each vector into subvectors and clusters each subspace independently, so a vector is stored as a short sequence of cluster IDs rather than raw floats — often hundreds of times smaller. Distance calculations then use precomputed lookup tables.

Compression~100–1000×
Trade-offAccuracy loss; pairs well with IVF

These aren't mutually exclusive. Most production systems layer them: IVF narrows the search to a handful of promising clusters, PQ compresses the candidates within those clusters so more of them fit in memory and get scored cheaply, and graph-based indexes like HNSW are reached for when recall and latency matter more than memory footprint. Other index families — tree-based (ANNOY), hash-based (locality-sensitive hashing) — exist too, but see far less production use in enterprise settings than the three above.

Indexing enables fast retrieval at query time, but building the index in the first place can take real time and memory — that build cost is part of the capacity plan, not an afterthought.

The Recall–Latency Trade-off, in Practice

Every ANN index exposes some knob — HNSW's ef_search, IVF's nprobe, or an equivalent — that trades search breadth for speed. Recall measures what fraction of the true nearest neighbors a query actually returns; 90% recall means 9 of the top 10 real matches came back. Push recall higher and latency climbs, because the index has to examine more candidates to be more confident it hasn't missed anything.

In production, the right target is almost never 99–100%. Most teams land in the 90–95% band: pushing from 95% to 99% recall can triple query latency while the downstream application — an LLM synthesizing an answer from the top-k chunks, a recommender blending several signals — often can't tell the difference. The only reliable way to find your own sweet spot is to build a small ground-truth set with exhaustive search and measure how recall actually moves your application-level metrics, rather than tuning to a recall number in the abstract.

Enterprise Architecture

How Production Vector Layers Are Actually Built

A vector database is rarely just an index. Because it stores the original object alongside its embedding — not the embedding alone — its architecture typically layers several kinds of storage: object storage for the source data, an inverted index for keyword and metadata filtering, and the vector index itself, often organized into shards that can be distributed and replicated independently across a cluster.

  • Sharding splits the vector index across machines so indexing and search both parallelize; each shard runs its own ANN search and results are merged, typically with a heap, at query time.
  • Metadata filtering — "only chunks from documents tagged region:EU and status:published" — has to run without destroying index efficiency. The two common approaches are a separate metadata index intersected with vector results, or partitioned indexes that pre-split data by common filter values.
  • Hybrid search merges dense vector similarity with traditional sparse keyword scoring (BM25), usually through a weighted combination or reciprocal rank fusion. It exists because pure vector search can miss exact-match signals — SKUs, error codes, legal clause numbers — that keyword search handles natively, and pure keyword search misses the paraphrases and synonyms that vector search is good at.
  • Dynamic updates are the sore spot for graph-based indexes like HNSW, which are optimized for reads. Most systems either queue writes and periodically rebuild, or use index variants built for incremental updates at some ongoing performance cost.
Decision Framework

Do You Actually Need a Vector Database?

Not every embedding-backed feature needs dedicated infrastructure. The honest answer depends on scale, update frequency, and how much accuracy loss the application can tolerate.

Probably skip it

  • Fewer than roughly 100K vectors — brute-force search with NumPy or a plain in-memory library is fast enough.
  • Vectors that change constantly — indexing overhead can exceed what you save on search.
  • Applications that need exact, guaranteed-correct nearest neighbors, not approximate ones — exact search with a library like FAISS is more appropriate than an ANN store.

Worth the investment

  • Millions of vectors with a real low-latency requirement.
  • Production semantic search, RAG, or recommendation systems operating at scale.
  • A need to filter by metadata without giving up index performance.
  • A need for the operational plumbing — sharding, replication, rolling updates — that a hand-rolled index doesn't give you for free.

Many teams start with the simplest option available and graduate to a dedicated vector database only once volume or latency requirements force the issue. That's usually the right order of operations — premature infrastructure is its own kind of technical debt.

2026 Snapshot

The Enterprise Landscape in 2026: Past the "Scale Wall"

A lot of the standard vector-database advice was written for chatbot-era RAG: one user, one query, a handful of retrieved chunks, forgiving latency. Agentic AI systems break that assumption. An agent working through a multi-step task can fire off dozens of parallel retrieval calls with much higher precision requirements than a single conversational lookup — and 2025-era RAG architectures, built and tuned for the simpler case, have been hitting what the industry is now calling the scale wall.

Enterprise survey data from early 2026 makes the shift concrete. Intent to adopt hybrid retrieval — combining dense vector search with sparse keyword matching and reranking, rather than relying on vectors alone — roughly tripled in a single quarter, and the top reason cited for needing a dedicated vector layer moved decisively toward raw operational reliability at scale rather than search quality in isolation.

10.3% → 33.3%Enterprise intent to adopt hybrid retrieval, Q1 2026 — nearly tripling in one quarter
35.6%Share of enterprises now running custom retrieval stacks rather than a single off-the-shelf vector database
31.1%Enterprises citing operational reliability at scale as the top driver for a dedicated vector layer, up sharply from January

Two things follow from that data. First, standalone vector databases as a category have lost some adoption share to custom, composed retrieval stacks — teams stitching together a vector index, a keyword engine, and a reranker rather than treating any single product as the whole answer. Second, general-purpose databases are pushing further into vector search rather than ceding the space to specialists: Postgres via pgvector, MongoDB Atlas Vector Search, and Oracle's unified vector search in Database 23ai/26ai all now let teams add a vector index to data that already lives in their operational database, instead of running a parallel system. For many enterprises, that consolidation — one fewer moving part, one fewer data-sync problem — outweighs the raw performance edge of a purpose-built vector database.

A Working Comparison

SystemTypePractical scaleBest fit
PineconeManaged, proprietaryBillionsTeams that want zero infrastructure ops
Milvus / Zilliz CloudOpen-source / managed100B+Very large-scale, multi-index deployments
WeaviateOpen-source, self-host or managedBillionsNative hybrid search + structured schema
QdrantOpen-source (Rust), self-host or managedTens of millionsHigh-performance filtered search
pgvector (Postgres)Open-source extensionMillionsTeams already standardized on Postgres
MongoDB Atlas Vector SearchManaged, integratedMillionsUnifying operational data and vectors in one collection
ChromaOpen-source, embedded or serverSmall–mediumPrototyping and lightweight production apps
LanceDBOpen-source, object-storage nativeSmall–largeRunning directly on cheap object storage
FAISSLibrary, not a databaseCustom / anyExact or research-grade ANN, embedded in your own service

Scale and pricing details shift quickly in this market — treat the table above as a starting shortlist, not a substitute for benchmarking against your own data and query patterns.

Enterprise reality check

The rise of custom retrieval stacks isn't really an argument against vector databases — it's an argument against treating any single component as sufficient. The teams reporting the best outcomes in 2026 are the ones that stopped asking "which vector database should we buy" and started asking "what does our retrieval pipeline need to guarantee end-to-end" — recall, latency, freshness, and access control together, not any one of them in isolation.

Governance

Governing the Vector Layer

Enterprise adoption brings requirements that a weekend prototype never has to face, and a surprising number of production incidents trace back to teams treating the vector store as "just another cache" rather than a system that needs the same governance as any other data store.

  • Access control at the chunk level. If source documents carry row-level or document-level permissions, the vector index needs to enforce them too — a RAG pipeline that retrieves a chunk the requesting user shouldn't see is a data leak, not a bug to fix later.
  • Data residency and multi-tenancy. Embeddings still encode the semantic content of the source data; storing or processing them outside a required jurisdiction can trigger the same compliance obligations as the source documents themselves.
  • Embedding-model versioning. Swapping embedding models silently breaks an index, because vectors from two different models don't share a meaningful distance space. Re-embedding and re-indexing needs to be a tracked, deliberate operation, not something that happens as a side effect of an unrelated model upgrade.
  • Freshness and staleness. A vector index answering questions from yesterday's data can be worse than no retrieval at all, especially in agentic workflows that act on what they retrieve. Incremental update strategy is a governance decision, not just a performance one.
  • Cost visibility. Vector storage and query cost scale with dimensionality, replica count, and query volume in ways that are easy to lose track of once multiple teams share one deployment — cost allocation per use case is worth setting up early.

Closing Thought

The core algorithmic ideas behind vector databases — embed, index approximately, trade a little accuracy for a lot of speed — have been stable for a few years now. What changed in 2026 is the workload: agentic systems query harder, want more precision, and have far less patience for a slow or unreliable retrieval layer than the conversational RAG use cases that popularized vector databases in the first place. Choosing a vector layer today is less about picking "the best vector database" and more about designing a retrieval architecture — vector index, keyword search, reranking, access control, and freshness — that can actually hold up under that pressure.

Further Reading