Skip to content
VG/Tech

EngineeringBy Alexander KhomenkoSep 202628 min read

Do you need a separate vector database?

Vector search is now available across database types. In theory you can add search to your existing database, or select a database for your data architecture without thinking about semantic search. How feasible is that?

Two years ago, adding semantic search usually meant adding a vector database and a synchronisation pipeline. Today, vector search is available across the data stack. Should you go with it?

A team copies its product catalogue into a vector service, then reimplements customer eligibility rules in application code. The vector service finds similar products, but the relational database still decides which ones a customer can buy.

Most databases now support vector search. The choice comes down to what each one can do with the nearest neighbours in the same query.

This guide is organised by database type because each type contributes a different operation to retrieval. The product determines whether that combination remains correct and fast under your real filters.

The complement test

Before jumping to your database type, answer three questions.

  1. Is the source of truth already here? If an embedding sits beside the source record it describes, one system owns both writes. A copy elsewhere creates a pipeline, and a pipeline creates another freshness boundary.
  2. Does the query need this type's native operator in the same request? Name the operator. If the answer is “no, we just need top-k”, using the current database may still be simpler, but it adds nothing specific to the query.
  3. Does this product's filtered vector search survive your tightest filter? This is not a type-level property. Two products with the same data model can apply filters at different points in an approximate-nearest-neighbour search and return different results.

If the source of truth is already here, the query uses a native operation, and filtered recall holds up, stay with the current database. If the query only needs top-k results, decide based on operational simplicity.

How eleven database types combine vector search with their native query operations.How eleven database types combine vector search with their native query operations.

Operational databases

Relational: similarity filtered by live data

Relational databases can keep embeddings in the same database as orders, stock, and permissions. A similarity query can then filter on the current data, with no copy to keep in sync. For product recommendations, similarity proposes candidates while SQL determines which ones are available to the current customer.

In practice: PostgreSQL with pgvector. The extension brings vector search into PostgreSQL, so similarity can run alongside joins and filters. Other major relational databases expose similar operations, but their filter behaviour is not interchangeable.

In practice, the query joins similarity with live stock and tenant filters:

sql
-- pgvector 0.8.6 syntax; inspect the real plan with EXPLAIN (ANALYZE, BUFFERS)
SELECT p.id, p.name, p.embedding <=> $1 AS distance
FROM products p
JOIN inventory i ON i.product_id = p.id
WHERE i.region = $2 AND i.in_stock > 0
  AND p.tenant_id = $3
ORDER BY p.embedding <=> $1
LIMIT 10;

What to test: With an approximate index, filtering can leave fewer than k rows. Iterative scans can continue until enough results are found or a configured bound is reached. Use EXPLAIN to compare strict and relaxed scan modes against exact results. Test RLS-constrained queries separately because security semantics do not guarantee retrieval quality.

Where it stops: move vector reads to a replica if scans compete with transactions. If filtered recall still fails after sensible partitioning, or the workload depends on lexical ranking the SQL stack cannot provide, use a separate search system.

Document: the embedding lives on the document

Document databases keep the embedding with the document it describes. They fit when varying fields are already the unit of read and write.

Agent memory is a good example because session documents can remain scoped to the current user.

In practice: MongoDB. $vectorSearch supports approximate and exact nearest-neighbour queries inside the aggregation pipeline. Fields used to filter results must be declared in the vector-search index. Azure Cosmos DB for NoSQL offers similar document-based retrieval with hybrid keyword and vector search.

javascript
// MongoDB 8.3.4+ / mongot 1.70.1+ supported self-managed pairing
db.products.aggregate([
  {$vectorSearch: {index: "product_vectors", path: "embedding",
    queryVector: queryVector, filter: {tenantId: tenantId},
    numCandidates: 200, limit: 10}},
  {$lookup: {from: "inventory", localField: "sku",
    foreignField: "sku", as: "stock"}},
  {$match: {"stock.available": {$gt: 0}}}
])

What to test: Filter fields are an index-time schema decision even in a schema-flexible database. Use MongoDB's exact: true path to establish recall for the most selective declared filter, then repeat across numCandidates values. Measure write-to-search visibility separately because the search process creates an indexing boundary.

Where it stops: a document database loses its advantage when search requires frequent index changes or the separate search process becomes a system of its own. Move out if its indexing delay cannot meet freshness requirements.

Key-value: similarity at the speed and lifetime of a cache

Key-value stores make vector search part of hot, short-lived state. The typical use is temporary data: when a key expires or is evicted, its vector drops out of search with it.

A semantic cache for LLM responses fits this model. Similarity finds a reusable answer, while the key's lifecycle keeps the cache bounded.

In practice: Redis. Redis has two vector surfaces. Vector sets store and search similar elements under one key. The Query Engine adds schema-based search over existing HASH or JSON values, including hybrid text and vector retrieval. Valkey provides a BSD-licensed alternative.

text
# Redis Open Source 8.4 command syntax; the indexed HASH keys may also expire
FT.HYBRID cache-idx
  SEARCH "@prompt:(refund policy)"
  VSIM @embedding $query KNN 2 K 20
  FILTER 1 "@tenant:{acme}"
  COMBINE RRF 2 WINDOW 20
  LIMIT 0 5
  PARAMS 2 query <query_vector_blob>
EXPIRE cache:answer:42 3600

Choose vector sets when you want to store and search a collection of similar elements under one key. Choose the Query Engine when vectors must participate in a broader search schema. The APIs remain distinct even though both live in Redis.

What to test: Compare filtered results and latency at the tightest filter in your workload. Redis can change its search strategy based on the size of the filtered set, so test with a realistic data distribution.

Where it stops: Redis becomes a poor fit when a large, mostly cold dataset makes memory costly or when the data requires durable history rather than a cache-like lifecycle. It should not be the only copy of embeddings you cannot regenerate.

Wide-column: similarity inside a partition

Wide-column databases make vector search follow the partition model. Similarity is strongest when the same key that scopes ordinary reads also bounds the candidate set.

For example, the partition key can narrow a search for similar sessions to one user before distance is calculated.

In practice: ScyllaDB Cloud. Vector search stays inside a declared partition, which makes partition design part of retrieval. Apache Cassandra offers the same broad wide-column pattern with vector search, but its index and filtering rules differ.

sql
-- ScyllaDB Cloud 2026.3 service syntax
CREATE CUSTOM INDEX session_ann
ON app.sessions((tenant_id), embedding)
USING 'vector_index'
WITH OPTIONS = {'similarity_function': 'COSINE'};

SELECT session_id, summary FROM app.sessions
WHERE tenant_id = ?
ORDER BY embedding ANN OF ? LIMIT 10;

What to test: Check that the partition key and filters match the real access pattern. Queries across partitions expand the work and weaken the reason to keep retrieval in a wide-column database.

Where it stops: move out when similarity must cross partitions or depends on predicates that do not fit the declared index shape. A search engine is a better match when lexical retrieval becomes central, while cloud-only availability rules out self-managed estates.

Graph: from nearest neighbour to what it connects to

Graph databases add relationship traversal. Vector search supplies an entry point; the graph supplies the context around it.

GraphRAG is one example. Vector search finds a relevant entry node, then traversal expands it into connected context.

In practice: Neo4j. Neo4j can find similar nodes or relationships, then follow their connections through the graph. Properties declared with the index can filter the vector search. Memgraph and FalkorDB also combine vector search with graph traversal, but their filter semantics differ.

cypher
// Neo4j 2026.08 / Cypher 25 syntax
MATCH (chunk:Chunk)
  SEARCH chunk IN (
    VECTOR INDEX chunk_embeddings FOR $query_vector
    WHERE chunk.tenantId = $tenant_id
    LIMIT 20
  ) SCORE AS score
MATCH (chunk)-[:MENTIONS]->(entity)-[:PART_OF]->(topic)
RETURN chunk.text, entity.name, topic.name, score LIMIT 8

What to test: Property predicates inside SEARCH are in-index filters. Conditions applied after SEARCH may reduce the result below k. Because traversal is the reason to use a graph, compare the final traversed result with exact similarity over the same authorised subgraph.

Where it stops: a graph database adds little when the collection has no meaningful relationships. Move out when post-traversal predicates still undermine recall after measured oversampling, or when the workload is ordinary document retrieval.

Multi-model: one query across several models

Multi-model databases let vector search cross data models. Their advantage is composition over one copy of the data, not that any single model is necessarily deeper than a specialist engine.

For example, retrieval can begin with a filtered vector search and continue through connected records without copying data between systems.

In practice: SurrealDB. SurrealDB combines nearest-neighbour search, filters, and graph traversal in SurrealQL. ArangoDB is another multi-model option, with vector functions in AQL. Check each product's filter behaviour before adopting it.

sql
-- SurrealDB 3.1+ syntax; confirm the current plan with EXPLAIN FULL
SELECT id, title,
  ->mentions->entity.{ id, name } AS entities
FROM article
WHERE tenant_id = $tenant
  AND embedding <|10,100|> $query_vector
LIMIT 10;
-- Add EXPLAIN FULL and look for predicate on KnnScan.

What to test: Confirm that every intended scalar predicate is visible on the vector scan, then test relationship conditions separately: a traversal after KNN is still a post-filter. Repeat this check after upgrades because vector optimisation in a fast-moving multi-model product changes more quickly than the data-model label does.

Where it stops: a multi-model system loses its case when one model dominates and a specialist engine offers more operational depth. Do not adopt a new database merely to avoid one vector service, especially when its production ecosystem is immature.

Search and analytics

Search engine: semantics added to a search product

Search engines add semantic retrieval to a mature search stack. If search already powers part of your product, vector retrieval goes into the same query as keyword search, alongside its filters and facets.

Catalogue search shows the value. Semantic similarity joins an established ranking and navigation experience without replacing it.

In practice: OpenSearch. A knn_vector field adds vector retrieval to the existing search stack. Hybrid queries combine lexical and vector results through a search pipeline. Elasticsearch serves the same broad use cases, but its APIs and licensing remain separate decisions.

json
GET /catalog/_search?search_pipeline=rrf-pipeline
{"size":10,"query":{"hybrid":{"queries":[
  {"match":{"description":"quiet travel headphones"}},
  {"knn":{"embedding":{"vector":<query_vector>,"k":50,
    "filter":{"term":{"region":"sg"}}}}}
]}}}

rrf-pipeline contains the score-ranker-processor with the rrf technique.

What to test: Put the filter inside the kNN clause when you need efficient filtering. A top-level post-filter can return far fewer than k. Check the configured engine and method, and fuse by rank when lexical and vector scores are not comparable.

Where it stops: operating a search engine is hard to justify when the dataset is small or search is not a product. Move the vector workload elsewhere when cluster memory and plugin complexity outweigh the search features being used.

Analytical warehouse: similarity as a batch join

Analytical warehouses add set-based batch SQL and joins across whole governed tables. In this type, similarity is often a many-query-to-many-record operation rather than one user-facing lookup.

Batch deduplication is a natural use case. A query table of new records can be matched against governed warehouse data in one similarity join.

In practice: BigQuery. Managed-only VECTOR_SEARCH accepts both a base table and a query table, so one statement runs a batch of nearest-neighbour searches. It supports approximate and exact results. Other warehouses expose similar operations, but their production maturity and serving models vary.

sql
-- BigQuery managed service syntax
SELECT query.new_id, base.existing_id, distance
FROM VECTOR_SEARCH(
  TABLE `prod.catalogue`, 'embedding',
  TABLE `staging.new_items`, 'embedding',
  top_k => 5, distance_type => 'COSINE'
)
WHERE distance < @duplicate_threshold;

What to test: Measure cost and runtime for a representative batch, then compare approximate and exact results under the same filters. Confirm that the index stores the fields used to narrow the search.

Where it stops: a warehouse is a poor fit for interactive serving or data that must become searchable immediately. Use it to compute batch results, then publish compact outputs to an operational store.

Time-series: similarity within a time window

Time-series databases make time part of candidate selection. They fit when “similar” is meaningless without “during which period?”

For incident analysis, a time window can bound the historical candidates before similarity ranks them.

In practice: TimescaleDB with pgvectorscale. A hypertable divides time into chunks, while pgvectorscale adds vector indexing and filtering to pgvector. KDB-X with KDB.AI provides a time-series-first alternative in the q ecosystem.

sql
-- TimescaleDB 2.x with current pgvectorscale syntax
CREATE INDEX ticket_embedding_idx ON tickets
USING diskann (embedding vector_cosine_ops);

SELECT id, opened_at, summary
FROM tickets
WHERE opened_at > now() - interval '30 days'
ORDER BY embedding <=> $1
LIMIT 10;

What to test: Chunk exclusion applies to the column that partitions the hypertable, not any column that happens to contain a timestamp. Confirm excluded chunks and vector-index use with EXPLAIN, then compare with a filter on a secondary timestamp to see the difference. If label filters matter, include the labels in the DiskANN index rather than assuming every SQL predicate is pushed into it.

Where it stops: a time-series database adds little when queries span all history or the records are not genuinely time-shaped. Move out when the task is pattern search over raw windows rather than nearest-neighbour retrieval over embeddings.

Storage and embedded

Object storage: vectors where cold data rests

Object storage makes vector search a serverless storage operation. You accept a latency floor in exchange for not provisioning a search cluster.

It fits occasional search over an archive, where the data can remain in durable storage.

In practice: Amazon S3 Vectors. This managed-only service stores vectors in dedicated buckets with strongly consistent writes and filterable metadata. Check its service limits against your planned collection and ingest rate. It integrates with the wider AWS retrieval stack. The scale guide covers other object-storage-native systems.

python
# Amazon S3 Vectors managed API
result = s3vectors.query_vectors(
    vectorBucketName="knowledge-cold",
    indexName="documents",
    queryVector={"float32": query_vector},
    topK=20,
    filter={"tenant": {"$eq": tenant_id}},
    returnDistance=True,
    returnMetadata=True,
)

What to test: Measure end-to-end latency on your expected access pattern, including after idle periods. Exercise real filters and ingestion volume against the service limits. Very small matching sets can still return fewer than topK.

Where it stops: move out when measured latency cannot support interactive serving or updates exceed service limits. Object storage is also a poor fit when hybrid retrieval must happen inside the same system or an AWS-only primitive conflicts with the deployment model.

Embedded and on-device: vectors in the application's file

Embedded databases keep vector search inside the application process, using a local file that works without a network connection.

Offline semantic search in a mobile app fits well. The index lives with local data and requires no server round trip.

In practice: SQLite with sqlite-vec. The dependency-free extension runs wherever SQLite does, including WASM. A vec0 virtual table stores vectors beside fields used to scope or return results. Other embedded databases offer vector indexes, but maturity and licensing vary.

sql
-- sqlite-vec 0.1.x syntax
CREATE VIRTUAL TABLE memories USING vec0(
  user_id TEXT PARTITION KEY,
  embedding FLOAT[768],
  +content TEXT
);
SELECT rowid, content, distance FROM memories
WHERE embedding MATCH ? AND user_id = ? AND k = 10
ORDER BY distance;

What to test: Put the scoping field in the vec0 table so it participates in vector search rather than filtering afterwards. Benchmark on the lowest-end supported device, where resource limits and update costs differ from desktop development.

Where it stops: move out when the data does not fit the device or the workload requires shared server-side writes. A local file also loses its advantage when cross-device synchronisation becomes the harder problem. Pre-v1 extension churn may rule it out for applications with strict release constraints.

When none of these fit: a dedicated vector database

A dedicated vector database is the right answer when the complement test fails everywhere you already store the data.

That usually means vector retrieval is the workload. No other data model contributes an operation important enough to make it a natural home. It can also mean the current database cannot maintain recall under required filters, or that vector traffic must operate independently of the transactional and analytical systems.

Scale can force the decision too. When retrieval architecture determines unit economics, a purpose-built system may expose controls the current system lacks. The architecture guide by data size covers the relevant tiers and product families; this article does not rank them because no benchmark was run here.

The decision should still follow evidence. “Vector database” is not a quality tier. It is an operating model optimised around retrieval. Adopt it when that workload deserves an independent system, not because an architecture diagram looks more complete with another cylinder.

The same three problems in every type

Filtered recall is a product property. The data model tells you why a combination is useful; the implementation tells you whether it returns the right neighbours. Test each product with its most selective real filter and confirm whether filtering occurs inside or after the vector scan. A post-filter can return fewer than k even when the underlying nearest-neighbour search works as designed.

Filters are a schema decision. Every database type has its own mechanism for deciding which fields can narrow vector search efficiently. The choice is often fixed before query time, so derive it from a query log rather than a workshop guess.

Scores do not compare across retrievers, and embedding models do not migrate. A BM25 score and a cosine score have different meanings; even vector scores from different result sets may not be comparable. Fuse by rank unless an evaluation set justifies calibrated weights. Changing the embedding model means re-embedding and re-indexing the content in every database type; the scale guide's migration checklist applies here too.

Two databases need two jobs

An operational store can own content while a key-value store caches semantically equivalent requests. Both hold some of the same data, but each serves a distinct workload.

The anti-pattern is copying one dataset into three vector-capable databases “so we can compare later”, with three pipelines. Every change to the data now has to reach three places, and without an evaluation set the comparison never happens.

Summary table

Type (example)Native operator the vector step gainsBest fit
Relational (PostgreSQL)Joins, constraints, transactions, RLSConstrained recommendations, entity resolution
Document (MongoDB)Document model, aggregation pipelineCatalogues, scoped memory, one-pipeline retrieval
Key-value (Redis)In-memory keys, TTL, evictionSemantic cache, working memory, routing
Wide-column (ScyllaDB Cloud)Partition access, write volume, replicationPer-user or per-device similarity
Graph (Neo4j)TraversalGraphRAG, entity resolution, connected recommendations
Multi-model (SurrealDB)One query across modelsVector plus document filter plus graph hop
Search engine (OpenSearch)BM25, analyzers, facets, fusionExisting search products and identifier-heavy search
Warehouse (BigQuery)Batch SQL, similarity joinsDeduplication, enrichment, exact results to test recall against
Time-series (TimescaleDB)Time pruning and windowsRecent semantic search and similar episodes
Object storage (S3 Vectors)Serverless, durable, pay per useArchives and low-traffic knowledge bases
Embedded (SQLite + sqlite-vec)In-process, offline fileLocal-first and per-user search

Test your current database before you replace it

Take 1,000 production queries with their filters. For each one, compute exact nearest neighbours over the filtered rows using the product's exact or index-bypassing path. If none exists, use a separate reference implementation.

Then run the same queries through the approximate index and compare recall at the final, user-visible result, not merely before joins or traversals remove candidates. Record latency and cost, but do not let a fast wrong result pass. If recall holds at the tightest real filter, the database you already run is probably the right home. If it does not, you now have a measured failure case to take to the next product evaluation.

Read more

Official implementation documentation for the products used in the examples:

Ready to put this into practice?