Skip to content
Reviews

Pinecone vs Weaviate vs Qdrant: Vector Databases for AI Agents

Pinecone, Weaviate and Qdrant compared for AI agent RAG: features, cost, hybrid search and when to use each database.

M
Max Beech· Founder
··11 min read
Pinecone vs Weaviate vs Qdrant: Vector Databases for AI Agents

TL;DR

  • Pinecone: Very fast queries, zero ops, expensive at scale. Rating: 4.5/5
  • Weaviate: Best hybrid search, flexible, moderate speed, mid-tier cost. Rating: 4.6/5
  • Qdrant: Cheapest (self-hosted is free, managed is low-cost), fast, smaller ecosystem. Rating: 4.3/5
  • Quick pick: Pinecone for ease, Weaviate for hybrid search, Qdrant for budget.
  • Pinecone charges for what Qdrant does free if you self-host. Is it worth it? Here's how to decide.

# Pinecone vs Weaviate vs Qdrant: Vector Database Showdown

Your AI agent needs a vector database for RAG. Do you use Pinecone (everyone uses it), Weaviate (heard good things), or Qdrant (open-source, cheaper)?

This guide compares all three for a typical RAG agent (for example, around 1M chunks embedded with OpenAI text-embedding-3-small at 1,536 dimensions): performance characteristics, cost, and when to use each.

What to Compare

When you test these yourself, measure the same things on your own data:

  • Query latency (p50, p95, p99)
  • Recall@10 (accuracy: does the result contain relevant docs in the top 10?)
  • Cost per million vectors
  • Hybrid search capability
  • Developer experience

Latency and recall vary a lot with hardware, index settings and query mix, so treat the comparisons below as general guidance rather than fixed numbers.

Pinecone

Verdict: Fastest queries, zero operations burden, most expensive.

Performance

Typically the lowest query latency of the three for managed deployments, with strong recall and high throughput per pod.

Why so fast? Purpose-built for vector search. Optimized indexing (proprietary algorithm), global edge network.

Cost

Pricing: A free tier for small projects, then paid plans that are the most expensive of the three for a production-grade index. Check Pinecone's pricing page for current figures, as plans change often.

Scaling: Cost per vector falls as you scale up, but the total bill still climbs faster than self-hosted alternatives.

Setup Experience

Installation: Zero. Sign up, get API key, start inserting vectors.

Indexing:

import pinecone

pinecone.init(api_key="...")
index = pinecone.Index("my-index")

# Upload 1M vectors
for i in range(0, 1_000_000,100):
    batch = vectors[i:i+100]
    index.upsert(vectors=batch)

Developer experience: 10/10. Simplest API, great docs, works immediately.

Hybrid Search

Support: Partial. Supports sparse-dense hybrid via "sparse_values" parameter.

index.query(
    vector=[0.1, 0.2, ...],  # Dense embedding
    sparse_vector={"indices": [10, 50], "values": [0.9, 0.7]},  # Sparse (keyword)
    top_k=10
)

Limitation: Manual BM25 calculation required. Not built-in like Weaviate.

Rating: 7/10 for hybrid search

Pros

  • Fastest queries
  • Zero ops (fully managed, auto-scaling)
  • Global edge network (low latency worldwide)
  • Best docs and DX

Cons

  • Most expensive of the three
  • Vendor lock-in (proprietary, can't self-host)
  • Hybrid search clunky (manual sparse vector generation)

Rating: 4.5/5

Use Pinecone if: Budget not constrained, want fastest queries, prefer zero ops.

---

Weaviate

Verdict: Best hybrid search, flexible schema, good performance, mid-tier cost.

Performance

Slower than Pinecone and Qdrant on raw query latency, but often the best recall when you use hybrid search.

Why good recall? Hybrid search (vector + BM25) built-in. Finds docs missed by pure vector search.

Cost

Pricing (managed Weaviate Cloud): a free sandbox tier, then paid tiers that sit between Pinecone and Qdrant. Check Weaviate's pricing page for current figures.

Self-hosted: Free (open-source), but requires Kubernetes/Docker management.

vs Pinecone: Usually cheaper

vs Qdrant: Usually more expensive than Qdrant's managed cloud

Setup Experience

Managed (Weaviate Cloud):

import weaviate

client = weaviate.Client(
    url="https://my-cluster.weaviate.network",
    auth_client_secret=weaviate.AuthApiKey(api_key="...")
)

# Define schema
schema = {
    "class": "Document",
    "vectorizer": "none",  # We provide embeddings
    "properties": [
        {"name": "content", "dataType": ["text"]},
        {"name": "source", "dataType": ["string"]}
    ]
}

client.schema.create_class(schema)

# Upload vectors (batch import)
with client.batch as batch:
    for doc in documents:
        batch.add_data_object(
            data_object={"content": doc.text, "source": doc.source},
            class_name="Document",
            vector=doc.embedding
        )

Developer experience: 8/10. More config than Pinecone, but flexible.

Hybrid Search

Support: Native. Best-in-class.

result = client.query.get(
    "Document", ["content", "source"]
).with_hybrid(
    query="What is RAG?",
    alpha=0.7  # 0.7 = 70% vector, 30% BM25
).with_limit(10).do()

Why superior? BM25 (keyword search) built-in. No manual sparse vector calculation.

Hybrid search catches edge cases (exact keyword matches, acronyms) vector search misses.

Rating: 10/10 for hybrid search

Advanced Features

1. Multi-tenancy: Built-in tenant isolation (separate namespaces per user)

2. Filtering: Filter by metadata before vector search

.with_where({
    "path": ["source"],
    "operator": "Equal",
    "valueString": "wikipedia"
}).with_near_vector({
    "vector": embedding
})

3. Generative search: Combine vector search + LLM generation (RAG in one query)

.with_generate(
    single_prompt="Summarize: {content}"
)

Pros

  • Best hybrid search (native BM25 + vector)
  • Strong recall with hybrid search
  • Flexible (multi-tenancy, filtering, generative search)
  • Open-source (can self-host)

Cons

  • Slower than Pinecone on raw latency
  • More complex setup than Pinecone
  • Mid-tier cost

Rating: 4.6/5

Use Weaviate if: Need hybrid search, want flexibility, recall matters more than latency.

---

Qdrant

Verdict: Cheapest (self-hosted or managed), fast, Rust-based, smaller ecosystem.

Performance

Why fast? Written in Rust (low-level performance), optimized HNSW index.

Usually faster than Weaviate and close to Pinecone, with good recall and throughput on modest hardware.

Cost

Managed (Qdrant Cloud): a free tier with limited throughput, then low-cost paid clusters. Check Qdrant's pricing page for current figures.

Self-hosted: Free (open-source). You pay only for the VM and storage, for example a 4 vCPU, 16GB RAM instance plus SSD storage.

vs Pinecone: Several times cheaper

vs Weaviate: Usually cheaper

Even self-hosted, including compute, it tends to be the cheapest of the three.

Setup Experience

Self-hosted (Docker):

docker run -p 6333:6333 qdrant/qdrant

Python client:

from qdrant_client import QdrantClient
from qdrant_client.models import Distance, VectorParams

client = QdrantClient(host="localhost", port=6333)

# Create collection
client.create_collection(
    collection_name="documents",
    vectors_config=VectorParams(size=1536, distance=Distance.COSINE)
)

# Upload vectors
client.upsert(
    collection_name="documents",
    points=[
        {"id": i, "vector": embedding, "payload": {"content": text}}
        for i, (embedding, text) in enumerate(zip(vectors, texts))
    ]
)

Developer experience: 8/10. Clean API, good docs, but smaller community than Pinecone/Weaviate.

Hybrid Search

Support: Yes (via sparse vectors)

from qdrant_client.models import SparseVector

client.search(
    collection_name="documents",
    query_vector=dense_embedding,
    sparse_vector=SparseVector(indices=[10, 50], values=[0.9, 0.7]),
    limit=10
)

Implementation: Similar to Pinecone (manual sparse vector generation).

Not as smooth as Weaviate (no built-in BM25), but works.

Rating: 7/10 for hybrid search

Pros

  • Cheapest (managed or self-hosted)
  • Fast (close to Pinecone)
  • Rust-based (low resource usage, stable)
  • Open-source (self-host option)

Cons

  • Smaller ecosystem (smaller community than Pinecone or Weaviate)
  • Fewer integrations (works with major frameworks, but less coverage)
  • Hybrid search not native (like Pinecone, requires manual BM25)

Rating: 4.3/5

Use Qdrant if: Budget-conscious, comfortable self-hosting, want good performance at low cost.

---

Performance Benchmark Summary

DatabaseLatencyRecallCostBest For
PineconeLowestHighHighestZero ops, speed-critical
WeaviateModerateHighest (with hybrid)MediumHybrid search, flexibility
QdrantLowHighLowestBudget, self-hosting

Decision Framework

Start
  ↓
Tight budget? → YES → Qdrant (managed) or self-host
  ↓ NO
  ↓
Need hybrid search? → YES → Weaviate (native BM25)
  ↓ NO
  ↓
Speed critical? → YES → Pinecone
  ↓ NO
  ↓
Prefer self-hosting? → YES → Qdrant or Weaviate (open-source)
  ↓ NO
  ↓
Want zero ops? → YES → Pinecone (fully managed, auto-scale)
  ↓
Default: Weaviate (best balance)

Example Use Case: Customer Support RAG

Imagine a support knowledge base of a few hundred thousand docs and tens of thousands of queries a month. All three would handle it comfortably. The latency differences are too small for users to notice in a chat interface, and recall is similar once the chunks are well prepared. At that point cost usually decides it, which tends to favour Qdrant, with Weaviate the pick if keyword-heavy queries (product codes, error messages) matter.

Migration Path

Moving between databases:

# Export from Pinecone
vectors = []
for ids_batch in pinecone_index.list():
    vectors.extend(pinecone_index.fetch(ids_batch).vectors)

# Import to Qdrant
qdrant_client.upsert(
    collection_name="documents",
    points=[{"id": v.id, "vector": v.values, "payload": v.metadata} for v in vectors]
)

# Migration time depends on collection size and batch settings

Downtime: 0 (run both in parallel, switch DNS/config when ready)

Frequently Asked Questions

Which has best scaling?

All three scale horizontally:

  • Pinecone: Automatic (add pods)
  • Weaviate: Add nodes to cluster
  • Qdrant: Add nodes, supports sharding

At 10M+ vectors, the gap narrows somewhat, but Qdrant generally keeps its cost advantage, especially self-hosted.

Can I switch databases later?

Yes. All use standard vector format, and migrating around 1M vectors is typically a short job rather than a project.

Risk: Minimal. Switching cost is low.

What about pgvector (Postgres extension)?

Worth considering:

  • Latency: Noticeably slower than the specialised databases at scale
  • Recall: Usually a little lower than specialised DBs
  • Cost: Cheapest if you already run Postgres

Use pgvector if: Already running Postgres, <100K vectors, low query volume.

Not recommended for: >1M vectors, high query rates, production RAG.

---

Bottom line: Pinecone for speed + zero ops, Weaviate for hybrid search + flexibility, Qdrant for budget + self-hosting. All three work well. Choose based on priorities.

Next: Read our Complete RAG Guide for full implementation with any vector database.

More from the blog

Stop doing the work around the work

OpenHelm connects to your tools, reads the context, and does the steps, so you sign off on the result instead of producing it. See how it covers an entire role’s weekly workload, check the pricing, or run it yourself with the free local app.