Vector search has quietly become one of the most important pieces of modern AI infrastructure. Every time you ask an LLM about your own documents, get a product recommendation that actually makes sense, or find a song “similar” to another — there’s a vector database doing the heavy lifting behind the scenes.
I went through the entire landscape — algorithms, libraries, databases, and vendor services — and put together this survey. Think of it as a map: what exists, what each layer does, and how to pick the right tool for your stage.
Table of Contents
- What is a Vector Database?
- Do I Need a Vector Database?
- Vector Search Applications
- OSS Vector Indexing/Querying Libraries
- OSS Vector Databases
- Vendor Vector Databases
- Amazon OpenSearch Offerings
- Other Libraries
- Recordings
- Articles and Posts
- ANN Algorithms and Comparisons
What is a Vector Database?
Vector databases are fully featured databases to help you deploy and scale applications that require vector search. They are based on underlying vector indexing/querying libraries/services, which provide the core vector search functionality, but provide MLOps affordances required for productioning applications.
These features include things like:
- Data management
- Metadata storage and filtering
- Scalability
- Online updates to stored vectors
- Backups and collections
- Ecosystem integration
Vector indexing and querying tools like FAISS and Annoy themselves implement one or more approximate nearest neighbor search algorithms (ANNs). The reason for adopting approximate-matching algorithms is that finding the single best match for a query ultimately requires an exhaustive search of a vector collection, which results in unacceptable latencies for large vector collections associated with many applications.
The Layers of Functionality
We can think of vector databases as being composed of these layers:

Do I Need a Vector Database?
If you’re just producing a POC or experimenting, then you may find just using a standalone ANN tool like FAISS or Annoy (see table below) will be sufficient. But as soon as you start to think of the MLOps challenges associated with deploying and maintaining this type of application, you’ll likely want to consider a vector database, or else build the needed functionality yourself.
Scenario: You build a document search demo with FAISS in a notebook — done in an afternoon. Then production asks: “Can we filter results by department?”, “Can we add new docs without rebuilding the index?”, “What happens when the pod restarts?” Each question is a feature you’d hand-roll. That’s exactly what vector databases sell you.
Vector Search Applications
-
Multi-modality search
- Semantic search — search by what you mean, keyword match not required
- Image similarity search
- Audio similarity search
-
Recommender and personalisation systems
- User and product embeddings
- Vector distances can be combined with business-logic constraints to achieve multiple goals
- e.g. rank product recommendations by satisfaction scores and revenue potential
-
Question and answering
- Augmenting LLM apps with knowledge and media from unstructured collections
-
Other possibilities
- Browse unstructured data
- Filter on metadata
- Re-rank search results
- Hybrid scoring (e.g. combining search engine rank score like BM25F with vector similarity measures across multiple modalities)
OSS Vector Indexing/Querying Libraries
These are the engines at the bottom of the stack. They implement the ANN algorithms but leave production concerns to you.
| Project | Stars | Description | License | Clients | ANN Algorithms |
|---|---|---|---|---|---|
| FAISS | 35k | A library for efficient similarity search and clustering of dense vectors | MIT | C++, Python | HNSW, Exact match (L2 and inner product), IVF_FLAT, LSH (Locality sensitive hashing), SQ (Scalar Quantizer), PQ (Product Quantizer), IVF+SQ, IVFADC, IVFADC+R |
| gensim | 16k | Mainly about topic modeling but does word vector indexing and querying | LGPL v2.1 | Python | Annoy, NMSLib, (other similarity methods) |
| Annoy | 14k | Approximate Nearest Neighbors in C++/Python optimized for memory usage and loading/saving to disk | Apache 2.0 | C++, Python | Tree-based random projections |
| SPTAG | 5k | A distributed approximate nearest neighborhood search (ANN) library which provides a high quality vector index build, search and distributed online serving toolkits for large scale vector search scenario | MIT | C++, Python | SPTAG-KDT (kd-tree and relative neighborhood graph), SPTAG-BKT (balanced k-means tree and relative neighborhood graph) |
| nmslib | 4k | Non-Metric Space Library (NMSLIB): An efficient similarity search library and a toolkit for evaluation of k-NN methods for generic non-metric spaces | Apache 2.0 | C++, Python | HNSW, sw-graph (Small World Graph), vptree (Vantage-Point tree with pruning rule adaptable to non-metric distances), napp (Neighborhood APProximation index), imple_invindx (vanilla uncompressed inverted index), brute_force |
| hnswlib | 5k | Header-only C++/Python library for fast approximate nearest neighbors | Apache 2.0 | C++, Python | HNSW |
| NGT | 1k | Nearest Neighbor Search with Neighborhood Graph and Tree for High-dimensional Data | Apache 2.0 | Python, Ruby, PHP, Rust, JavaScript, Go, C, C++ | NGT (Graph and tree-based method), QG (Quantized graph-based method), QBG (Quantized blob graph-based method) |
| PyNNDescent | 1k | A Python nearest neighbor descent for approximate nearest neighbors | BSD 2-Clause | Python | NNDescent: neighbor graph based searching with random projection trees for initialisation |
Library Highlights
FAISS — Optimized for memory usage and speed. Billion-scale similarity search with GPUs. The go-to library for raw performance.
Annoy — Uses static files as indexes; enables sharing index across processes. Used at Spotify for music recommendations. Can load data quickly from disk using nmap.
SPTAG — From Microsoft Research. Associated papers: SPANN (Highly-efficient Billion-scale Approximate Nearest Neighbor Search), Query-driven iterated neighborhood graph search for large scale indexing, Scalable k-NN graph construction for visual descriptors, Trinary-Projection Trees for Approximate Nearest Neighbor Search.
hnswlib — Header-only HNSW from the nmslib project. Lightweight, header-only, no dependencies other than C++11.
NGT — Research papers: Optimization of Indexing Based on k-Nearest Neighbor Graph for Proximity, Pruned Bi-directed K-nearest Neighbor Graph for Proximity Search, On Approximately Searching for Similar Word Embeddings.
PyNNDescent — Integrates well with Scikit-learn, including providing support for the KNeighborTransformer as a drop-in replacement for algorithms that make use of nearest neighbor computations. Paper: Efficient K-Nearest Neighbor Graph Construction for Generic Similarity Measures.
OSS Vector Databases
Also included are search engine tools that support vector-based indexes and querying.
| Project | Stars | Description | License | Clients | ANN Algorithms |
|---|---|---|---|---|---|
| Elasticsearch | 73k | Free and Open, Distributed, RESTful Search Engine | Apache 2.0, Elastic License 2.0, Server Side Public License | Java, JavaScript, Ruby, Go, .NET, PHP, Perl, Python, Eland Python, Rust | HNSW (via Lucene), Exact match |
| Redis | 69k | Redis is an in-memory database that persists on disk. The data model is key-value, but many different kind of values are supported: Strings, Lists, Sets, Sorted Sets, Hashes, Streams, HyperLogLogs, Bitmaps | Custom License | — | — |
| Milvus | 35k | A cloud-native vector database, storage for next generation AI applications | Apache 2.0 | Python, Java, Go, Node.js | FLAT, IVF_FLAT, IVF_SQ8, IVF_PQ, HNSW (via faiss), ANNOY |
| CockroachDB | 31k | CockroachDB — the cloud native, distributed SQL database designed for high availability, effortless scale, and control over data placement | Custom License | — | — |
| MongoDB | 27k | — | Server Side Public License | — | — |
| Qdrant | 24k | Vector Database for the next generation of AI applications. Also available in the cloud | Apache 2.0 | Go, Rust, Python, JavaScript, Elixir, PHP, Ruby, Java | Custom HNSW in Rust, extended to support categorical |
| Valkey | 22k | Valkey is a specialized vector database designed to manage and search high-dimensional vector data efficiently | BSD 3-Clause License | — | — |
| Chroma | 20k | The AI-native open-source embedding database | Apache 2.0 | Python, JavaScript | HNSW (via hnswlib) |
| pgvector | 16k | Open-source vector similarity search for Postgres | Bespoke license | Any language with Postgres client | IVF_FLAT |
| scipy.spatial.kdtree | 14k | Simple “naked” vector queries | BSD 3 Clause | Python | kd-tree |
| Weaviate | 13k | Open source vector database that stores both objects and vectors, allowing for combining vector search with structured filtering with the fault-tolerance and scalability of a cloud-native database | BSD 3-Clause | Python, JavaScript, Go, Java | Custom HNSW tuned to support scale and full CRUD, plug-in ANN algorithms that support full CRUD |
| OpenSearch | 11k | Open source distributed and RESTful search engine | Apache 2.0 | Python, Java, JavaScript, Go, Ruby, PHP, .NET, Rust | HNSW (nmslib, faiss, Lucene), IVF (faiss) |
| txtai | 11k | Semantic search and workflows powered by language models | Apache 2.0 | REST API, Python, JavaScript, Java, Rust, Go | Multiple index backends: Faiss, ANNOY, hnswlib |
| Cassandra | 9k | Apache Cassandra is a highly-scalable partitioned row store. Rows are organized into tables with a required primary key | Apache 2.0 | — | — |
| Deep Lake | 9k | Database for AI. Store Vectors, Images, Texts, Videos, etc. Use with LLMs/LangChain | MPL 2.0 | Python | — |
| Vespa | 6k | The open big data serving engine | Apache 2.0 | Python, Java | HNSW (modified for realtime CRUD and metadata filtering) |
| LanceDB | 7k | Developer-friendly, embedded retrieval engine for multimodal AI | Apache 2.0 | Python, Rust | SOTA ANN-index algorithms |
| Marqo | 5k | Unified embedding generation and search engine | Apache 2.0 | REST API, Python Client | HNSW |
| Infinity | 4k | The AI-native database built for LLM applications, providing incredibly fast hybrid search of dense vector, sparse vector, tensor (multi-vector), and full-text | Apache 2.0 | — | — |
| USearch | 3k | Fast Open-Source Search and Clustering engine for Vectors and Strings in C++, C, Python, JavaScript, Rust, Java, Objective-C, Swift, C#, GoLang, and Wolfram | Apache 2.0 | — | — |
| Featureform | 2k | The Virtual Feature Store. Turn your existing data infrastructure into a feature store | Mozilla Public License 2.0 | Python | HNSW (via hnswlib) |
| Vald | 2k | A Highly Scalable Distributed Vector Search Engine | Apache 2.0 | RPC | NGT (via NGT) |
Database Highlights
Elasticsearch — ANN search introduced in Elasticsearch 8.0. The k-NN plugin provides core vector database functionality, plus learning to rank capabilities.
Milvus — Knowhere is the vector search execution engine of Milvus. Dependencies: faiss, hnswlib, NGT, annoy. Server is installed on K8 in Standalone or Cluster mode. Focus on scalability of end-to-end search engine — scalable queries, index and reindex vector data efficiently. Used by Towhee (flexible application-oriented framework for computing embedding vectors over unstructured data), Haystack (open source NLP framework leveraging Transformer models), LangChain (building applications with LLMs through composability), GPTCache (library for creating semantic cache to store responses from LLM queries).
Qdrant — Local and client-server modes. Built entirely in Rust. Extended payload filtering support for queries enabling custom business logic. Good performance and well designed APIs and docs. Multiple vector types supported in a collection (not supported by other options). Integrations:
- Cohere — Use Cohere embeddings with Qdrant
- DocArray — Use Qdrant as a document store in DocArray
- LangChain — Use Qdrant as a memory backend for LangChain
- LlamaIndex — Use Qdrant as a Vector Store with LlamaIndex
- OpenAI — ChatGPT retrieval plugin — Use Qdrant as a memory backend for ChatGPT
- Microsoft Semantic Kernel — Use Qdrant as persistent memory with Semantic Kernel
Chroma — Can BYO vectors, or use built in embeddings with Sentence Transformers. Integrations: LangChain, LlamaIndex.
Weaviate — Expressive query syntax supported with GraphQL-like interface. Impressive question answering component. Integrations: Auto-GPT (memory backend), Cohere, DocArray, Haystack, Hugging Face, LangChain, LlamaIndex, OpenAI ChatGPT retrieval plugin, OpenAI embeddings. Vectorizer modules enable the engine to manage creation of vectors from input objects: text2vec-openai, text2vec-cohere, text2vec-huggingface, text2vec-transformers, text2vec-contextionary, img2vec-neural, multi2vec-clip, ref2vec-centroid.
Vespa — Can be deployed in a range of ways, including multi-node deployments. Focus on low-latency computation over large data sets. Includes BM25 implementation to support hybrid search (BM25 exact matching + vector search). Includes re-ranking support with different models like XGBoost.
txtai — Backed by NeuML. BYO embeddings or use Hugging Face models to generate embeddings. Pipelines powered by language models that run question-answering, labeling, transcription, translation, summarization, LLM prompts and more. Workflows to join pipelines together and aggregate business logic.
LanceDB — Serverless, low-latency vector database for AI applications. Backed by Lance, a modern columnar data format that is an alternative to Parquet. Tagline: “Search More; Manage Less.”
Featureform — Does not replace your existing infrastructure. Rather, Featureform transforms your existing infrastructure into a feature store. Designed for both single data scientists and large enterprise teams. Run locally or on Kubernetes. Embeddinghub provides vector search component. Deployment: Standalone on Kubernetes, Standalone on Docker.
Vald — Deployment: Standalone on Kubernetes, Standalone on Docker.
scipy.spatial.kdtree — Haven’t tested for N=100s of dimensions but apparently doesn’t scale that well.
Vendor Vector Databases
Also included are search engine tools that support vector-based indexes and querying.
| Vendor | Description | ANN Algorithms |
|---|---|---|
| Algolia | — | — |
| Activeloop | Makers of the OSS Deep Lake AI datastore | — |
| Amazon Elasticsearch | — | — |
| Amazon OpenSearch | — | — |
| Chroma | — | — |
| Deepset | NLP service built on Haystack, that offers many vector DB integrations | — |
| Elastic | — | — |
| Jina | Jina is an MLOps framework to build multimodal AI microservice-based applications written in Python that can communicate via gRPC, HTTP and WebSocket protocols. Built on FastAPI/Pydantic, DocArray | — |
| LanceDB | — | — |
| Marqo | — | — |
| Pinecone | Fully managed vector database to support unstructured search engine needs | Exact KNN (via faiss), Proprietary ANN algorithm |
| Qdrant | — | — |
| SingleStore | — | — |
| Vespa | — | — |
| Weaviate Cloud Services | — | — |
| Zilliz | — | — |
Amazon OpenSearch Offerings
Amazon OpenSearch provides several vector search options:
- Amazon OpenSearch
- Amazon OpenSearch Serverless
- Built on OpenSearch Serverless
- Automatically adjusts resources by scaling and adapting to changing workload patterns and demand
- No need for re-indexing/re-loading data
- Separate compute for indexing and search
- Vector Engine for Amazon OpenSearch Serverless (Preview)
- Interesting features:
AWS Blog Posts
- Introducing the vector engine for Amazon OpenSearch Serverless, now in preview, 2023-07-26 — Many customers today are using OpenSearch kNN search in managed clusters for offering semantic search and personalization in their applications. With the vector engine, you can get the same functionality with the simplicity of a serverless environment.
- Amazon OpenSearch Service’s vector database capabilities explained, 2023-06-21 — Very good overview of the technical domain and capabilities of the service and also use cases.
- Building an NLU-powered search application with Amazon SageMaker and the Amazon OpenSearch Service KNN feature
Other Libraries
| Project | Stars | Description | License | Clients | Comments |
|---|---|---|---|---|---|
| Jina | 22k | Build multimodal AI services via cloud native technologies | Apache 2.0 | Python | MLOps framework to build multimodal AI microservice-based applications written in Python that can communicate via gRPC, HTTP and WebSocket protocols. Built on FastAPI/Pydantic and DocArray |
| Haystack | 21k | AI orchestration framework to build customizable, production-ready LLM applications. Connect components (models, vector DBs, file converters) to pipelines or agents that can interact with your data | Apache 2.0 | Python | Core Concepts: Pipelines, Nodes, Agents, Tools, DocumentStores. Best suited for building RAG, question answering, semantic search or conversational agent chatbots |
| DocArray | 3k | Represent, send, and store multimodal data. Neural Search. Vector Search. Document Store | Apache 2.0 | Python | You can best think of it as a new kind of ORM for vector databases. DocArray’s job is to take multimodal, nested and domain-specific data and to map it to a vector database, store it there, and thus make it searchable. Integrations: FastAPI, Jina, Annlite |
Recordings
- Open NLP meetup #3: Landscape of Vector Databases and Overcoming DPR’s Input Limit
- Players in Vector Search: algorithms, software and use cases - Dmitry Kan
Articles and Posts
-
Not All Vector Databases Are Made Equal, Oct 2, 2021 — Detailed comparison of Milvus, Pinecone, Vespa, Weaviate, Vald, GSI and Qdrant. Dimensions of comparison: Value proposition (what is the unique feature that makes the whole vector search engine stand out from the crowd?), Type (general type: vector database, big data platform. Managed / Self-hosted), Architecture (high-level system architecture, including aspects of sharding, plugins, scalability, hardware details), Algorithm (what algorithm approach to similarity / vector search was taken), Code (is it open or close source?)
-
Prithivi Da’s LinkedIn posts on how to choose a vector database
- Part 1 — Boils down to two choices: (1) Which vector algorithms support your indexing/querying needs? (2) Choosing a framework that tackles your “deployment-at-scale” concerns. Framework needs to consider: scaling (horizontal), replication, fault tolerance, HA, security, selection of ANN algorithms
- Part 2
- Part 3
-
Navigating the Vector Database Landscape, March 1st 2022 — Tools surveyed: Featureform/Embeddinghub, Milvus, Pinecone, Weaviate, Vald. Dimensions to consider: Usability, Flexibility (e.g. which ANN algorithms supported), Scalability, Cost, Organizational Features
-
Vector search in search engine tools:
-
Open Source Vector Database Comparison: comparing Weaviate to other vector DBs
-
Understanding Semantic Search — Part 7: The Rise of Vector Databases in the World of Semantic Search, Dec 7, 2022 — Why vector DBs? Traditional tabular DBs unsuitable for storing unstructured data. NoSQL DBs not optimized for efficient indexing and retrieval of vector data. Therefore vector-optimized DBs needed. ANN algorithms: Google’s ScaNN, Meta’s FAISS, Spotify’s ANNOY, HNSW. Scaling: ANN algorithms and metadata filtering can help to reduce the latency for a single user query. However prod systems may have e.g. 10k users querying concurrently, so scaling considerations essential for low-latency. Vector DBs use replication and sharding to scale horizontally.
-
Benchmark Vector Search Databases with One Million Data, 16 Nov 2022
-
The Best Vector Database for Stablecog’s Semantic Search, Apr 12, 2023 — Real-world evaluation comparing pgvector, Milvus, Weaviate, and Qdrant. Chose OpenCLIP for embeddings (open source version of OpenAI’s CLIP, trained on public and free data). Chose OpenCLIP ViT-h/14 - LAION 2B which preserves a lot of details in its vector embeddings.
pgvector: Pros — Easy to implement, satisfactory search result relevancy. Cons — Building indexes very slow, search performance unsatisfactory with many vectors stored.
Milvus: Pros — Open source and easily self-hostable, UI component that makes browsing the database easy, search results were satisfactorily relevant. Cons — Can only store 1 vector in a schema/collection (had to choose image or prompt, not both), resource usage was very higher than others, API responses felt less well-designed requiring post-processing, quantization index (compressed vector) (IVF_PQ) seemed to hurt search relevancy.
Weaviate: Pros — Open source and easily self-hostable, relatively good performance, has modules which can be installed to automatically compute the embeddings from raw objects (positions it as a more convenient all-in-one solution). Cons — Documentation is harder to navigate, product quantisation took over 12 hours to re-index (experimental feature), resource consumption was quite high compared to Qdrant, SDKs and API is not as nice to use as Milvus or Qdrant, can only store 1 vector in a schema/collection.
Qdrant (final selection): Pros — Multiple vectors in a collection (can store both prompt embeddings and image embeddings), Scalar Quantization has a noticeable reduction in memory usage while barely having any impact on search relevancy, resource usage is lower than the other options quite significantly, documentation is very easy to digest and straight forward, API is straight forward and just works how you’d expect it to “common sense”, search is very fast the highest performing of all options tried. Cons — Lack of built-in authentication for self-hosted option. We created sc-qdrant as a result, which makes it easy to deploy a cluster behind an nginx proxy with basic auth.
ANN Algorithms and Comparisons
- Comprehensive Guide To Approximate Nearest Neighbors Algorithms, Feb 15, 2020
- http://ann-benchmarks.com — New approximate nearest neighbor benchmarks, GitHub, Paper
- A Data Scientist’s Guide to Picking an Optimal Approximate Nearest-Neighbor Algorithm, GitHub benchmark results
Scaling Considerations
ANN algorithms and metadata filtering can help to reduce the latency for a single user query. However production systems may have e.g. 10k users querying concurrently, so scaling considerations are essential for low-latency performance. Vector databases use replication and sharding to scale horizontally to meet that load.
How to Choose
When evaluating any vector database, compare along these axes:
- Value proposition — what makes this engine stand out?
- Type — vector database vs big-data platform; managed vs self-hosted
- Architecture — sharding, plugins, scalability, hardware details
- Algorithm — which ANN approach, and what unique capabilities?
- Code — open or closed source?
It boils down to two choices:
- Which vector algorithms support your indexing/querying needs?
- Which framework tackles your “deployment-at-scale” concerns? (horizontal scaling, replication, fault tolerance, HA, security, selection of ANN algorithms)
Wrapping Up
The vector database landscape has matured into clear layers:
- Algorithms (HNSW, IVF, PQ) are implemented in
- Libraries (FAISS, hnswlib, Annoy) which are wrapped by
- Databases (Milvus, Qdrant, Weaviate, pgvector) which are offered as
- Managed services (Pinecone, Zilliz, AWS)
Start simple: FAISS/pgvector for a prototype, then graduate to a dedicated vector database when metadata filtering, online updates, and scaling become real problems. And when you choose — benchmark with your data, your embedding model, and your operational constraints, exactly like Stablecog did.
The landscape keeps evolving, but the fundamentals hold: approximate nearest neighbor search is the engine, and the database around it is what turns a demo into a product.