Pinecone vs Weaviate vs Qdrant: Vector Databases for AI Agents
Pinecone, Weaviate and Qdrant compared for AI agent RAG: features, cost, hybrid search and when to use each database.

TL;DR
- Pinecone: Very fast queries, zero ops, expensive at scale. Rating: 4.5/5
- Weaviate: Best hybrid search, flexible, moderate speed, mid-tier cost. Rating: 4.6/5
- Qdrant: Cheapest (self-hosted is free, managed is low-cost), fast, smaller ecosystem. Rating: 4.3/5
- Quick pick: Pinecone for ease, Weaviate for hybrid search, Qdrant for budget.
- Pinecone charges for what Qdrant does free if you self-host. Is it worth it? Here's how to decide.
# Pinecone vs Weaviate vs Qdrant: Vector Database Showdown
Your AI agent needs a vector database for RAG. Do you use Pinecone (everyone uses it), Weaviate (heard good things), or Qdrant (open-source, cheaper)?
This guide compares all three for a typical RAG agent (for example, around 1M chunks embedded with OpenAI text-embedding-3-small at 1,536 dimensions): performance characteristics, cost, and when to use each.
What to Compare
When you test these yourself, measure the same things on your own data:
- Query latency (p50, p95, p99)
- Recall@10 (accuracy: does the result contain relevant docs in the top 10?)
- Cost per million vectors
- Hybrid search capability
- Developer experience
Latency and recall vary a lot with hardware, index settings and query mix, so treat the comparisons below as general guidance rather than fixed numbers.
Pinecone
Verdict: Fastest queries, zero operations burden, most expensive.
Performance
Typically the lowest query latency of the three for managed deployments, with strong recall and high throughput per pod.
Why so fast? Purpose-built for vector search. Optimized indexing (proprietary algorithm), global edge network.
Cost
Pricing: A free tier for small projects, then paid plans that are the most expensive of the three for a production-grade index. Check Pinecone's pricing page for current figures, as plans change often.
Scaling: Cost per vector falls as you scale up, but the total bill still climbs faster than self-hosted alternatives.
Setup Experience
Installation: Zero. Sign up, get API key, start inserting vectors.
Indexing:
import pinecone
pinecone.init(api_key="...")
index = pinecone.Index("my-index")
# Upload 1M vectors
for i in range(0, 1_000_000,100):
batch = vectors[i:i+100]
index.upsert(vectors=batch)Developer experience: 10/10. Simplest API, great docs, works immediately.
Hybrid Search
Support: Partial. Supports sparse-dense hybrid via "sparse_values" parameter.
index.query(
vector=[0.1, 0.2, ...], # Dense embedding
sparse_vector={"indices": [10, 50], "values": [0.9, 0.7]}, # Sparse (keyword)
top_k=10
)Limitation: Manual BM25 calculation required. Not built-in like Weaviate.
Rating: 7/10 for hybrid search
Pros
- Fastest queries
- Zero ops (fully managed, auto-scaling)
- Global edge network (low latency worldwide)
- Best docs and DX
Cons
- Most expensive of the three
- Vendor lock-in (proprietary, can't self-host)
- Hybrid search clunky (manual sparse vector generation)
Rating: 4.5/5
Use Pinecone if: Budget not constrained, want fastest queries, prefer zero ops.
---
Weaviate
Verdict: Best hybrid search, flexible schema, good performance, mid-tier cost.
Performance
Slower than Pinecone and Qdrant on raw query latency, but often the best recall when you use hybrid search.
Why good recall? Hybrid search (vector + BM25) built-in. Finds docs missed by pure vector search.
Cost
Pricing (managed Weaviate Cloud): a free sandbox tier, then paid tiers that sit between Pinecone and Qdrant. Check Weaviate's pricing page for current figures.
Self-hosted: Free (open-source), but requires Kubernetes/Docker management.
vs Pinecone: Usually cheaper
vs Qdrant: Usually more expensive than Qdrant's managed cloud
Setup Experience
Managed (Weaviate Cloud):
import weaviate
client = weaviate.Client(
url="https://my-cluster.weaviate.network",
auth_client_secret=weaviate.AuthApiKey(api_key="...")
)
# Define schema
schema = {
"class": "Document",
"vectorizer": "none", # We provide embeddings
"properties": [
{"name": "content", "dataType": ["text"]},
{"name": "source", "dataType": ["string"]}
]
}
client.schema.create_class(schema)
# Upload vectors (batch import)
with client.batch as batch:
for doc in documents:
batch.add_data_object(
data_object={"content": doc.text, "source": doc.source},
class_name="Document",
vector=doc.embedding
)Developer experience: 8/10. More config than Pinecone, but flexible.
Hybrid Search
Support: Native. Best-in-class.
result = client.query.get(
"Document", ["content", "source"]
).with_hybrid(
query="What is RAG?",
alpha=0.7 # 0.7 = 70% vector, 30% BM25
).with_limit(10).do()Why superior? BM25 (keyword search) built-in. No manual sparse vector calculation.
Hybrid search catches edge cases (exact keyword matches, acronyms) vector search misses.
Rating: 10/10 for hybrid search
Advanced Features
1. Multi-tenancy: Built-in tenant isolation (separate namespaces per user)
2. Filtering: Filter by metadata before vector search
.with_where({
"path": ["source"],
"operator": "Equal",
"valueString": "wikipedia"
}).with_near_vector({
"vector": embedding
})3. Generative search: Combine vector search + LLM generation (RAG in one query)
.with_generate(
single_prompt="Summarize: {content}"
)Pros
- Best hybrid search (native BM25 + vector)
- Strong recall with hybrid search
- Flexible (multi-tenancy, filtering, generative search)
- Open-source (can self-host)
Cons
- Slower than Pinecone on raw latency
- More complex setup than Pinecone
- Mid-tier cost
Rating: 4.6/5
Use Weaviate if: Need hybrid search, want flexibility, recall matters more than latency.
---
Qdrant
Verdict: Cheapest (self-hosted or managed), fast, Rust-based, smaller ecosystem.
Performance
Why fast? Written in Rust (low-level performance), optimized HNSW index.
Usually faster than Weaviate and close to Pinecone, with good recall and throughput on modest hardware.
Cost
Managed (Qdrant Cloud): a free tier with limited throughput, then low-cost paid clusters. Check Qdrant's pricing page for current figures.
Self-hosted: Free (open-source). You pay only for the VM and storage, for example a 4 vCPU, 16GB RAM instance plus SSD storage.
vs Pinecone: Several times cheaper
vs Weaviate: Usually cheaper
Even self-hosted, including compute, it tends to be the cheapest of the three.
Setup Experience
Self-hosted (Docker):
docker run -p 6333:6333 qdrant/qdrantPython client:
from qdrant_client import QdrantClient
from qdrant_client.models import Distance, VectorParams
client = QdrantClient(host="localhost", port=6333)
# Create collection
client.create_collection(
collection_name="documents",
vectors_config=VectorParams(size=1536, distance=Distance.COSINE)
)
# Upload vectors
client.upsert(
collection_name="documents",
points=[
{"id": i, "vector": embedding, "payload": {"content": text}}
for i, (embedding, text) in enumerate(zip(vectors, texts))
]
)Developer experience: 8/10. Clean API, good docs, but smaller community than Pinecone/Weaviate.
Hybrid Search
Support: Yes (via sparse vectors)
from qdrant_client.models import SparseVector
client.search(
collection_name="documents",
query_vector=dense_embedding,
sparse_vector=SparseVector(indices=[10, 50], values=[0.9, 0.7]),
limit=10
)Implementation: Similar to Pinecone (manual sparse vector generation).
Not as smooth as Weaviate (no built-in BM25), but works.
Rating: 7/10 for hybrid search
Pros
- Cheapest (managed or self-hosted)
- Fast (close to Pinecone)
- Rust-based (low resource usage, stable)
- Open-source (self-host option)
Cons
- Smaller ecosystem (smaller community than Pinecone or Weaviate)
- Fewer integrations (works with major frameworks, but less coverage)
- Hybrid search not native (like Pinecone, requires manual BM25)
Rating: 4.3/5
Use Qdrant if: Budget-conscious, comfortable self-hosting, want good performance at low cost.
---
Performance Benchmark Summary
| Database | Latency | Recall | Cost | Best For |
|---|---|---|---|---|
| Pinecone | Lowest | High | Highest | Zero ops, speed-critical |
| Weaviate | Moderate | Highest (with hybrid) | Medium | Hybrid search, flexibility |
| Qdrant | Low | High | Lowest | Budget, self-hosting |
Decision Framework
Start
↓
Tight budget? → YES → Qdrant (managed) or self-host
↓ NO
↓
Need hybrid search? → YES → Weaviate (native BM25)
↓ NO
↓
Speed critical? → YES → Pinecone
↓ NO
↓
Prefer self-hosting? → YES → Qdrant or Weaviate (open-source)
↓ NO
↓
Want zero ops? → YES → Pinecone (fully managed, auto-scale)
↓
Default: Weaviate (best balance)Example Use Case: Customer Support RAG
Imagine a support knowledge base of a few hundred thousand docs and tens of thousands of queries a month. All three would handle it comfortably. The latency differences are too small for users to notice in a chat interface, and recall is similar once the chunks are well prepared. At that point cost usually decides it, which tends to favour Qdrant, with Weaviate the pick if keyword-heavy queries (product codes, error messages) matter.
Migration Path
Moving between databases:
# Export from Pinecone
vectors = []
for ids_batch in pinecone_index.list():
vectors.extend(pinecone_index.fetch(ids_batch).vectors)
# Import to Qdrant
qdrant_client.upsert(
collection_name="documents",
points=[{"id": v.id, "vector": v.values, "payload": v.metadata} for v in vectors]
)
# Migration time depends on collection size and batch settingsDowntime: 0 (run both in parallel, switch DNS/config when ready)
Frequently Asked Questions
Which has best scaling?
All three scale horizontally:
- Pinecone: Automatic (add pods)
- Weaviate: Add nodes to cluster
- Qdrant: Add nodes, supports sharding
At 10M+ vectors, the gap narrows somewhat, but Qdrant generally keeps its cost advantage, especially self-hosted.
Can I switch databases later?
Yes. All use standard vector format, and migrating around 1M vectors is typically a short job rather than a project.
Risk: Minimal. Switching cost is low.
What about pgvector (Postgres extension)?
Worth considering:
- Latency: Noticeably slower than the specialised databases at scale
- Recall: Usually a little lower than specialised DBs
- Cost: Cheapest if you already run Postgres
Use pgvector if: Already running Postgres, <100K vectors, low query volume.
Not recommended for: >1M vectors, high query rates, production RAG.
---
Bottom line: Pinecone for speed + zero ops, Weaviate for hybrid search + flexibility, Qdrant for budget + self-hosting. All three work well. Choose based on priorities.
Next: Read our Complete RAG Guide for full implementation with any vector database.
More from the blog
How to Set Up Claude Code on a VPS: A Complete Guide
Claude Code VPS setup, step by step: provisioning, authentication, tmux vs systemd, security, and an honest look at when a VPS beats running locally.
Claude Code Agent Teams: How to Run Them on a Schedule
Claude Code Agent Teams runs up to 10 parallel Claude instances against one task list. What it is, how it works, and how to schedule runs.
Stop doing the work around the work
OpenHelm connects to your tools, reads the context, and does the steps, so you sign off on the result instead of producing it. See how it covers an entire role’s weekly workload, check the pricing, or run it yourself with the free local app.