Skip to content
Shrijayan
Go back

Vector Databases: A Survey of the Landscape

Vector search has quietly become one of the most important pieces of modern AI infrastructure. Every time you ask an LLM about your own documents, get a product recommendation that actually makes sense, or find a song “similar” to another — there’s a vector database doing the heavy lifting behind the scenes.

I went through the entire landscape — algorithms, libraries, databases, and vendor services — and put together this survey. Think of it as a map: what exists, what each layer does, and how to pick the right tool for your stage.

Table of Contents


What is a Vector Database?

Vector databases are fully featured databases to help you deploy and scale applications that require vector search. They are based on underlying vector indexing/querying libraries/services, which provide the core vector search functionality, but provide MLOps affordances required for productioning applications.

These features include things like:

Vector indexing and querying tools like FAISS and Annoy themselves implement one or more approximate nearest neighbor search algorithms (ANNs). The reason for adopting approximate-matching algorithms is that finding the single best match for a query ultimately requires an exhaustive search of a vector collection, which results in unacceptable latencies for large vector collections associated with many applications.

The Layers of Functionality

We can think of vector databases as being composed of these layers:

Vector Database Layers


Do I Need a Vector Database?

If you’re just producing a POC or experimenting, then you may find just using a standalone ANN tool like FAISS or Annoy (see table below) will be sufficient. But as soon as you start to think of the MLOps challenges associated with deploying and maintaining this type of application, you’ll likely want to consider a vector database, or else build the needed functionality yourself.

Scenario: You build a document search demo with FAISS in a notebook — done in an afternoon. Then production asks: “Can we filter results by department?”, “Can we add new docs without rebuilding the index?”, “What happens when the pod restarts?” Each question is a feature you’d hand-roll. That’s exactly what vector databases sell you.


Vector Search Applications


OSS Vector Indexing/Querying Libraries

These are the engines at the bottom of the stack. They implement the ANN algorithms but leave production concerns to you.

ProjectStarsDescriptionLicenseClientsANN Algorithms
FAISS35kA library for efficient similarity search and clustering of dense vectorsMITC++, PythonHNSW, Exact match (L2 and inner product), IVF_FLAT, LSH (Locality sensitive hashing), SQ (Scalar Quantizer), PQ (Product Quantizer), IVF+SQ, IVFADC, IVFADC+R
gensim16kMainly about topic modeling but does word vector indexing and queryingLGPL v2.1PythonAnnoy, NMSLib, (other similarity methods)
Annoy14kApproximate Nearest Neighbors in C++/Python optimized for memory usage and loading/saving to diskApache 2.0C++, PythonTree-based random projections
SPTAG5kA distributed approximate nearest neighborhood search (ANN) library which provides a high quality vector index build, search and distributed online serving toolkits for large scale vector search scenarioMITC++, PythonSPTAG-KDT (kd-tree and relative neighborhood graph), SPTAG-BKT (balanced k-means tree and relative neighborhood graph)
nmslib4kNon-Metric Space Library (NMSLIB): An efficient similarity search library and a toolkit for evaluation of k-NN methods for generic non-metric spacesApache 2.0C++, PythonHNSW, sw-graph (Small World Graph), vptree (Vantage-Point tree with pruning rule adaptable to non-metric distances), napp (Neighborhood APProximation index), imple_invindx (vanilla uncompressed inverted index), brute_force
hnswlib5kHeader-only C++/Python library for fast approximate nearest neighborsApache 2.0C++, PythonHNSW
NGT1kNearest Neighbor Search with Neighborhood Graph and Tree for High-dimensional DataApache 2.0Python, Ruby, PHP, Rust, JavaScript, Go, C, C++NGT (Graph and tree-based method), QG (Quantized graph-based method), QBG (Quantized blob graph-based method)
PyNNDescent1kA Python nearest neighbor descent for approximate nearest neighborsBSD 2-ClausePythonNNDescent: neighbor graph based searching with random projection trees for initialisation

Library Highlights

FAISS — Optimized for memory usage and speed. Billion-scale similarity search with GPUs. The go-to library for raw performance.

Annoy — Uses static files as indexes; enables sharing index across processes. Used at Spotify for music recommendations. Can load data quickly from disk using nmap.

SPTAG — From Microsoft Research. Associated papers: SPANN (Highly-efficient Billion-scale Approximate Nearest Neighbor Search), Query-driven iterated neighborhood graph search for large scale indexing, Scalable k-NN graph construction for visual descriptors, Trinary-Projection Trees for Approximate Nearest Neighbor Search.

hnswlib — Header-only HNSW from the nmslib project. Lightweight, header-only, no dependencies other than C++11.

NGT — Research papers: Optimization of Indexing Based on k-Nearest Neighbor Graph for Proximity, Pruned Bi-directed K-nearest Neighbor Graph for Proximity Search, On Approximately Searching for Similar Word Embeddings.

PyNNDescent — Integrates well with Scikit-learn, including providing support for the KNeighborTransformer as a drop-in replacement for algorithms that make use of nearest neighbor computations. Paper: Efficient K-Nearest Neighbor Graph Construction for Generic Similarity Measures.


OSS Vector Databases

Also included are search engine tools that support vector-based indexes and querying.

ProjectStarsDescriptionLicenseClientsANN Algorithms
Elasticsearch73kFree and Open, Distributed, RESTful Search EngineApache 2.0, Elastic License 2.0, Server Side Public LicenseJava, JavaScript, Ruby, Go, .NET, PHP, Perl, Python, Eland Python, RustHNSW (via Lucene), Exact match
Redis69kRedis is an in-memory database that persists on disk. The data model is key-value, but many different kind of values are supported: Strings, Lists, Sets, Sorted Sets, Hashes, Streams, HyperLogLogs, BitmapsCustom License——
Milvus35kA cloud-native vector database, storage for next generation AI applicationsApache 2.0Python, Java, Go, Node.jsFLAT, IVF_FLAT, IVF_SQ8, IVF_PQ, HNSW (via faiss), ANNOY
CockroachDB31kCockroachDB — the cloud native, distributed SQL database designed for high availability, effortless scale, and control over data placementCustom License——
MongoDB27k—Server Side Public License——
Qdrant24kVector Database for the next generation of AI applications. Also available in the cloudApache 2.0Go, Rust, Python, JavaScript, Elixir, PHP, Ruby, JavaCustom HNSW in Rust, extended to support categorical
Valkey22kValkey is a specialized vector database designed to manage and search high-dimensional vector data efficientlyBSD 3-Clause License——
Chroma20kThe AI-native open-source embedding databaseApache 2.0Python, JavaScriptHNSW (via hnswlib)
pgvector16kOpen-source vector similarity search for PostgresBespoke licenseAny language with Postgres clientIVF_FLAT
scipy.spatial.kdtree14kSimple “naked” vector queriesBSD 3 ClausePythonkd-tree
Weaviate13kOpen source vector database that stores both objects and vectors, allowing for combining vector search with structured filtering with the fault-tolerance and scalability of a cloud-native databaseBSD 3-ClausePython, JavaScript, Go, JavaCustom HNSW tuned to support scale and full CRUD, plug-in ANN algorithms that support full CRUD
OpenSearch11kOpen source distributed and RESTful search engineApache 2.0Python, Java, JavaScript, Go, Ruby, PHP, .NET, RustHNSW (nmslib, faiss, Lucene), IVF (faiss)
txtai11kSemantic search and workflows powered by language modelsApache 2.0REST API, Python, JavaScript, Java, Rust, GoMultiple index backends: Faiss, ANNOY, hnswlib
Cassandra9kApache Cassandra is a highly-scalable partitioned row store. Rows are organized into tables with a required primary keyApache 2.0——
Deep Lake9kDatabase for AI. Store Vectors, Images, Texts, Videos, etc. Use with LLMs/LangChainMPL 2.0Python—
Vespa6kThe open big data serving engineApache 2.0Python, JavaHNSW (modified for realtime CRUD and metadata filtering)
LanceDB7kDeveloper-friendly, embedded retrieval engine for multimodal AIApache 2.0Python, RustSOTA ANN-index algorithms
Marqo5kUnified embedding generation and search engineApache 2.0REST API, Python ClientHNSW
Infinity4kThe AI-native database built for LLM applications, providing incredibly fast hybrid search of dense vector, sparse vector, tensor (multi-vector), and full-textApache 2.0——
USearch3kFast Open-Source Search and Clustering engine for Vectors and Strings in C++, C, Python, JavaScript, Rust, Java, Objective-C, Swift, C#, GoLang, and WolframApache 2.0——
Featureform2kThe Virtual Feature Store. Turn your existing data infrastructure into a feature storeMozilla Public License 2.0PythonHNSW (via hnswlib)
Vald2kA Highly Scalable Distributed Vector Search EngineApache 2.0RPCNGT (via NGT)

Database Highlights

Elasticsearch — ANN search introduced in Elasticsearch 8.0. The k-NN plugin provides core vector database functionality, plus learning to rank capabilities.

Milvus — Knowhere is the vector search execution engine of Milvus. Dependencies: faiss, hnswlib, NGT, annoy. Server is installed on K8 in Standalone or Cluster mode. Focus on scalability of end-to-end search engine — scalable queries, index and reindex vector data efficiently. Used by Towhee (flexible application-oriented framework for computing embedding vectors over unstructured data), Haystack (open source NLP framework leveraging Transformer models), LangChain (building applications with LLMs through composability), GPTCache (library for creating semantic cache to store responses from LLM queries).

Qdrant — Local and client-server modes. Built entirely in Rust. Extended payload filtering support for queries enabling custom business logic. Good performance and well designed APIs and docs. Multiple vector types supported in a collection (not supported by other options). Integrations:

Chroma — Can BYO vectors, or use built in embeddings with Sentence Transformers. Integrations: LangChain, LlamaIndex.

Weaviate — Expressive query syntax supported with GraphQL-like interface. Impressive question answering component. Integrations: Auto-GPT (memory backend), Cohere, DocArray, Haystack, Hugging Face, LangChain, LlamaIndex, OpenAI ChatGPT retrieval plugin, OpenAI embeddings. Vectorizer modules enable the engine to manage creation of vectors from input objects: text2vec-openai, text2vec-cohere, text2vec-huggingface, text2vec-transformers, text2vec-contextionary, img2vec-neural, multi2vec-clip, ref2vec-centroid.

Vespa — Can be deployed in a range of ways, including multi-node deployments. Focus on low-latency computation over large data sets. Includes BM25 implementation to support hybrid search (BM25 exact matching + vector search). Includes re-ranking support with different models like XGBoost.

txtai — Backed by NeuML. BYO embeddings or use Hugging Face models to generate embeddings. Pipelines powered by language models that run question-answering, labeling, transcription, translation, summarization, LLM prompts and more. Workflows to join pipelines together and aggregate business logic.

LanceDB — Serverless, low-latency vector database for AI applications. Backed by Lance, a modern columnar data format that is an alternative to Parquet. Tagline: “Search More; Manage Less.”

Featureform — Does not replace your existing infrastructure. Rather, Featureform transforms your existing infrastructure into a feature store. Designed for both single data scientists and large enterprise teams. Run locally or on Kubernetes. Embeddinghub provides vector search component. Deployment: Standalone on Kubernetes, Standalone on Docker.

Vald — Deployment: Standalone on Kubernetes, Standalone on Docker.

scipy.spatial.kdtree — Haven’t tested for N=100s of dimensions but apparently doesn’t scale that well.


Vendor Vector Databases

Also included are search engine tools that support vector-based indexes and querying.

VendorDescriptionANN Algorithms
Algolia——
ActiveloopMakers of the OSS Deep Lake AI datastore—
Amazon Elasticsearch——
Amazon OpenSearch——
Chroma——
DeepsetNLP service built on Haystack, that offers many vector DB integrations—
Elastic——
JinaJina is an MLOps framework to build multimodal AI microservice-based applications written in Python that can communicate via gRPC, HTTP and WebSocket protocols. Built on FastAPI/Pydantic, DocArray—
LanceDB——
Marqo——
PineconeFully managed vector database to support unstructured search engine needsExact KNN (via faiss), Proprietary ANN algorithm
Qdrant——
SingleStore——
Vespa——
Weaviate Cloud Services——
Zilliz——

Amazon OpenSearch Offerings

Amazon OpenSearch provides several vector search options:

AWS Blog Posts


Other Libraries

ProjectStarsDescriptionLicenseClientsComments
Jina22kBuild multimodal AI services via cloud native technologiesApache 2.0PythonMLOps framework to build multimodal AI microservice-based applications written in Python that can communicate via gRPC, HTTP and WebSocket protocols. Built on FastAPI/Pydantic and DocArray
Haystack21kAI orchestration framework to build customizable, production-ready LLM applications. Connect components (models, vector DBs, file converters) to pipelines or agents that can interact with your dataApache 2.0PythonCore Concepts: Pipelines, Nodes, Agents, Tools, DocumentStores. Best suited for building RAG, question answering, semantic search or conversational agent chatbots
DocArray3kRepresent, send, and store multimodal data. Neural Search. Vector Search. Document StoreApache 2.0PythonYou can best think of it as a new kind of ORM for vector databases. DocArray’s job is to take multimodal, nested and domain-specific data and to map it to a vector database, store it there, and thus make it searchable. Integrations: FastAPI, Jina, Annlite

Recordings


Articles and Posts


ANN Algorithms and Comparisons


Scaling Considerations

ANN algorithms and metadata filtering can help to reduce the latency for a single user query. However production systems may have e.g. 10k users querying concurrently, so scaling considerations are essential for low-latency performance. Vector databases use replication and sharding to scale horizontally to meet that load.


How to Choose

When evaluating any vector database, compare along these axes:

  1. Value proposition — what makes this engine stand out?
  2. Type — vector database vs big-data platform; managed vs self-hosted
  3. Architecture — sharding, plugins, scalability, hardware details
  4. Algorithm — which ANN approach, and what unique capabilities?
  5. Code — open or closed source?

It boils down to two choices:

  1. Which vector algorithms support your indexing/querying needs?
  2. Which framework tackles your “deployment-at-scale” concerns? (horizontal scaling, replication, fault tolerance, HA, security, selection of ANN algorithms)

Wrapping Up

The vector database landscape has matured into clear layers:

Start simple: FAISS/pgvector for a prototype, then graduate to a dedicated vector database when metadata filtering, online updates, and scaling become real problems. And when you choose — benchmark with your data, your embedding model, and your operational constraints, exactly like Stablecog did.

The landscape keeps evolving, but the fundamentals hold: approximate nearest neighbor search is the engine, and the database around it is what turns a demo into a product.


Share this post on:

Next Post
Trained a Transaction Foundation Model from Scratch on 2×RTX 4090