Vector Databases Compared: pgvector, Qdrant, Weaviate, Milvus & More
Choosing a vector database for RAG: pgvector, Qdrant, Weaviate, Milvus, Pinecone and Chroma compared on hosting, filtering, hybrid search, scale and cost model.

Choosing a vector database comes down to five questions: where it runs (self-hosted or managed), how well it combines vector search with metadata filters, whether it does hybrid keyword-plus-vector search, how far it scales, and how you pay for it. For most business RAG systems we start with pgvector in an existing PostgreSQL database. We move to Qdrant, Weaviate or Milvus when scale or search features require it, and to Pinecone when a team wants fully managed infrastructure. Here is how the options compare as of September 2026.
What a vector database does in a RAG system
In retrieval-augmented generation, documents are split into chunks, turned into embedding vectors and stored. At question time, the query is embedded and the database returns the most similar chunks, usually through an approximate nearest neighbour (ANN) index such as HNSW. Then an LLM writes the answer.
The vector index is only part of the job. In real systems we also need:
- Metadata filtering: only search documents this user may see, for this customer, in this language.
- Hybrid search: combine semantic similarity with keyword matching (BM25 or sparse vectors). Pure vector search often misses exact product codes, names and error messages.
- Updates and deletes: documents change, and deleted data must actually disappear (important under GDPR).
- Operational basics: backups, access control, monitoring.
If you are new to RAG, read our RAG best practices first. Chunking and retrieval strategy usually matter more than the database brand.
The main options at a glance
| Database | Licence | Hosting | Hybrid search | Filtering | Best fit |
|---|---|---|---|---|---|
| pgvector | Open source (PostgreSQL extension) | Self-host or any managed Postgres that ships it | Vector + Postgres full-text search, combined in SQL | Full SQL WHERE, joins, row-level security |
Teams already on Postgres; small to mid-size corpora |
| Qdrant | Apache 2.0 (Rust) | Self-host (Docker) or Qdrant Cloud | Dense + sparse vectors | Rich JSON payload filters | Performance-focused self-hosting, strong filtering |
| Weaviate | BSD 3-Clause (Go) | Self-host or Weaviate Cloud | BM25 + vector built in | Structured filters, multi-tenancy | Built-in vectoriser modules, multi-tenant SaaS |
| Milvus | Apache 2.0 (LF AI & Data) | Lite / Standalone / Distributed, or Zilliz Cloud | Sparse vectors + BM25 full-text | Metadata filtering, multi-tenancy | Very large scale, Kubernetes-native, GPU indexing |
| Pinecone | Proprietary | Fully managed (serverless); BYOC option | Dense + sparse / full-text in one index | Metadata filters | Teams that want zero database operations |
| Chroma | Apache 2.0 | In-memory, local, client-server, or Chroma Cloud | Full-text + vector | Metadata and document filters | Prototypes, local tools, simple apps |
Other options you will come across include LanceDB (embedded, file-based), Elasticsearch/OpenSearch (if you already run them for search), Redis, and vector search built into MongoDB Atlas and the major cloud databases. If your team already operates one of these well, its vector features may be enough.
pgvector: the sensible default
pgvector adds a vector column type and ANN indexes (HNSW and IVFFlat) to PostgreSQL. It also supports halfvec (half precision), sparsevec and binary vectors, and several distance metrics. Recent versions add iterative index scans, which fix the old problem of filtered queries returning too few results.
Why we reach for it first:
- One database. Documents, metadata, permissions and vectors live together. Filtering by tenant or access rights is a normal SQL
WHEREclause or row-level security policy. - Transactions and deletes work the way your team already expects.
- Hybrid search is possible by combining pgvector with Postgres full-text search, for example with reciprocal rank fusion in SQL.
- Easy to host. It runs wherever Postgres runs, including your own servers. That matters for private RAG.
Limits to know: indexed vector columns support up to 2,000 dimensions (4,000 with halfvec), and large HNSW indexes need a lot of memory and take time to build. At tens of millions of vectors with heavy query load, a dedicated engine usually becomes easier to run.
Qdrant: fast, filter-friendly, self-hostable
Qdrant is a dedicated vector engine written in Rust. Strong points:
- Payload filtering designed to work with the vector index rather than as a post-filter.
- Dense, sparse and multi-vector support for hybrid and late-interaction retrieval.
- Quantization to cut memory use.
- Distributed mode with sharding and replication.
- A single Docker container to start, and a managed cloud with a free tier.
It is a good choice when you want a dedicated vector store you can self-host, with heavy filtering (multi-tenant apps, permission-aware search).
Weaviate: batteries included
Weaviate stores objects and vectors together and ships hybrid search (BM25 + vector), built-in vectoriser modules that call embedding providers for you, reranking and multi-tenancy. It offers REST, gRPC and GraphQL APIs. It fits teams that want the database to handle embedding and hybrid ranking with little glue code, or SaaS products that need clean per-tenant isolation.
Milvus: built for very large scale
Milvus is a distributed, Kubernetes-native system with a wide choice of index types (HNSW, IVF, DiskANN and others), GPU acceleration, hot/cold storage and native BM25 full-text search through sparse vectors. Milvus Lite runs inside a Python process for development, and Zilliz Cloud is the managed version. Choose it when you expect hundreds of millions to billions of vectors and have the platform skills to operate it, or pay for the managed service.
Pinecone: fully managed
Pinecone is a proprietary, managed-only vector database with a serverless architecture. You do not run servers. Its current platform supports metadata filtering, dense and sparse vectors and full-text search in one index, plus integrated embedding and reranking. There is also a bring-your-own-cloud option for running inside your own cloud account. It suits teams without database operations capacity who are comfortable with a SaaS vendor holding their vectors (check the available regions) and a usage-based bill.
Chroma: easiest to start
Chroma is the fastest way to get vectors into a prototype: pip install chromadb and you have an in-memory or local persistent store. It now also offers a client-server mode and a hosted cloud with full-text, regex and metadata filtering. It is great for notebooks, internal tools and small apps. For large multi-tenant production systems, compare it carefully with the options above.
Understanding the cost model
We will not quote prices here, because they change often. Check each vendor’s current pricing page. What matters is how you pay:
| Model | Who | You pay for | Watch out for |
|---|---|---|---|
| Self-hosted open source | pgvector, Qdrant, Weaviate, Milvus, Chroma | Servers (RAM matters most for HNSW), storage, your team’s ops time | Memory sizing, backups, upgrades |
| Managed cluster | Qdrant Cloud, Weaviate Cloud, Zilliz Cloud, managed Postgres | Provisioned capacity per hour/month | Paying for idle capacity |
| Serverless / usage-based | Pinecone, serverless tiers of others | Storage per GB, read and write units, sometimes inference tokens | Bills that grow with query volume; plan minimums |
Rules of thumb for estimating:
- Memory is the main cost driver for HNSW indexes. It scales with vector count × dimensions × bytes per dimension, plus index overhead. Smaller embedding dimensions, half precision and quantization reduce it a lot.
- Embedding generation is often a larger cost than storage. Re-embedding your whole corpus when you change models is a real expense, so plan for it.
- Operations time is a real cost of self-hosting. It is small for pgvector on an existing database and larger for a distributed Milvus cluster.
For the LLM side of the bill, see our LLM cost optimization guide.
How to choose: a decision guide
- Already on PostgreSQL, under a few million chunks? Use pgvector. Revisit only if you hit clear limits.
- Need data on your own infrastructure with strong filtering and more scale? Qdrant or Weaviate, self-hosted.
- Hundreds of millions of vectors or more? Milvus (or Zilliz Cloud), or a managed service sized for it.
- No ops capacity, speed to market matters most? Pinecone or a managed cloud tier of an open-source engine.
- Prototype or local tool? Chroma, or pgvector in a Docker container.
- Strict EU data residency or on-premise requirements? Prefer a self-hostable open-source option, and check the region options of any managed service.
Whatever you choose, keep the vector store behind a small internal interface in your code, and keep the source documents and embedding pipeline reproducible. Switching databases later is then a migration project, not a rewrite.
Key takeaways
- Retrieval quality depends more on chunking, hybrid search and filtering than on the database brand.
- pgvector is the pragmatic default for most business RAG systems.
- Qdrant and Weaviate are strong self-hostable engines. Milvus targets very large scale. Pinecone removes operations work. Chroma is ideal for getting started.
- Compare cost models (self-hosted, provisioned or usage-based) rather than headline prices, and include embedding and operations costs.
- Keep your data pipeline portable so you can change databases later.
If you are planning a RAG system and want help choosing and setting up the right retrieval stack, especially one that keeps data on your own infrastructure, see our private RAG service or contact us.