Skip to content
Academy

RAG Pipeline Optimization for Agent Accuracy

How chunking strategy, embedding models and retrieval methods affect RAG accuracy for AI agents, and how to benchmark them on your own queries.

M
Max Beech· Founder
··15 min read

TL;DR

  • Three levers decide RAG accuracy for agents: chunking strategy, embedding model, and retrieval method. Measure each against your own ground-truth queries.
  • A strong starting point: 500-token chunks with around 20% overlap + text-embedding-3-large + hybrid search (BM25 + vector).
  • Most impactful optimization: Hybrid retrieval. Least impactful per pound spent: Expensive embedding models, which cost several times more for a modest gain.

Jump to Evaluation method · Jump to Chunking strategies · Jump to Embedding models · Jump to Retrieval methods · Jump to Recommendations

# RAG Pipeline Optimization for Agent Accuracy: A Practical Guide

Every AI agent builder faces the same question: "How do I make my agent stop hallucinating and actually use the knowledge I gave it?"

The answer is almost always RAG (Retrieval-Augmented Generation): retrieve relevant context from your knowledge base, inject it into the LLM prompt, get better answers. Simple concept. Devilish implementation.

How do you chunk documents? Fixed-size? Semantic? Sentence-based?

Which embedding model? OpenAI's latest? Open-source alternatives?

How do you retrieve? Pure vector similarity? Keyword search? Both?

Most teams pick defaults, ship it, and hope for the best. A better approach is to measure accuracy, latency, and cost for each change against a fixed set of real queries. This guide covers what to test and what tends to win.

How to evaluate your pipeline

Build a representative query set

Sample real queries from production and make sure the mix reflects how your agents are actually used, for example:

  • Research queries: "What are best practices for X?"
  • Factual lookups: "What's our policy on Y?"
  • Troubleshooting: "How do I fix error Z?"
  • Comparison: "Difference between A and B?"

Note the size and make-up of the knowledge base too (product docs, internal wikis, support articles, meeting transcripts, languages), because the best configuration depends on it.

Ground truth labeling

For each query, establish ground truth by:

  1. Human experts manually answering the query using the full knowledge base
  2. Identifying which document chunks contain the answer
  3. Rating agent responses on 0-100 scale for correctness

Accuracy metric: Percentage of queries where agent response scored ≥85 (substantially correct).

What to vary

Vary three dimensions, one at a time:

1. Chunking strategy

  • Fixed 250 tokens, no overlap
  • Fixed 500 tokens, no overlap
  • Fixed 500 tokens, 20% overlap
  • Fixed 1000 tokens, no overlap
  • Semantic chunking (split on topic shifts)
  • Sentence-based (preserve sentence boundaries)

2. Embedding model

  • text-embedding-ada-002 (OpenAI, 1536d)
  • text-embedding-3-small (OpenAI, 1536d)
  • text-embedding-3-large (OpenAI, 3072d)
  • all-MiniLM-L6-v2 (open-source, 384d)
  • bge-large-en-v1.5 (open-source, 1024d)

3. Retrieval method

  • Pure vector similarity (cosine)
  • Pure keyword search (BM25)
  • Hybrid (vector + keyword, weighted combination)

Run every configuration on the same query sample for a fair comparison.

Baseline configuration

Default (what most teams start with):

  • Chunking: Fixed 1000 tokens, no overlap
  • Embedding: text-embedding-ada-002
  • Retrieval: Pure vector similarity
  • Top-k: 5 chunks

Chunking strategy results

Chunking strategy had the second-largest impact on accuracy after retrieval method.

Chunking strategyAccuracyLatencyNotes
Fixed 250 tokens, no overlapLowerLowToo granular, loses context
Fixed 500 tokens, no overlapGoodLowGood balance
Fixed 500 tokens, 20% overlapBestLowBest overall
Fixed 1000 tokens, no overlapBaselineMediumBaseline
Semantic chunkingGoodHigherSlower, good accuracy
Sentence-basedModerateLowPreserves coherence

Typical winner: 500-token chunks with 20% overlap

Why 500 tokens with overlap works

Problem with no overlap: Important concepts spanning chunk boundaries get split, reducing retrieval accuracy.

Example:

Chunk 1: "...our pricing model offers three tiers. Enterprise tier includes..."
Chunk 2: "...advanced analytics, dedicated support, and custom integrations."

Query: "What's included in Enterprise tier?"

Without overlap, Chunk 1 mentions "Enterprise" but doesn't list features. Chunk 2 lists features but doesn't mention "Enterprise." Neither chunk alone fully answers the query.

With 20% overlap:

Chunk 1: "...our pricing model offers three tiers. Enterprise tier includes advanced analytics, dedicated support..."
Chunk 2: "...Enterprise tier includes advanced analytics, dedicated support, and custom integrations. Pricing starts at..."

Now both chunks contain the full answer.

Overlap percentage impact

Overlap raises storage and retrieval cost roughly in proportion to the percentage you add. Accuracy tends to improve up to around 20% overlap and then flatten.

Diminishing returns after about 20%. That makes 20% a sensible default.

Semantic chunking considerations

Semantic chunking (splitting on topic shifts using NLP) does well, but rarely beats fixed chunks with overlap. Trade-offs:

Pros:

  • Preserves topic coherence
  • Handles variable-length documents well
  • Better for narrative content (meeting transcripts, articles)

Cons:

  • Slower (NLP analysis overhead)
  • Variable chunk sizes complicate batching
  • Requires tuning per content type

Recommendation: Use semantic chunking for unstructured narrative content (transcripts, blogs). Use fixed 500-token with overlap for structured docs (APIs, wikis, FAQs).

Embedding model comparison

Embedding model choice matters less than retrieval method or chunking, but still significant.

Embedding modelDimsRelative accuracyCost/1M tokens
ada-002 (baseline)1536Baseline$0.10
text-emb-3-small1536Better$0.02
text-emb-3-large3072Best$0.13
MiniLM-L6-v2 (OSS)384Lower~$0 (self-host)
bge-large-en-v1.5 (OSS)1024Similar to baseline~$0 (self-host)

Winner: text-embedding-3-large for accuracy, text-embedding-3-small for cost-effectiveness.

Model selection guidance

Use text-embedding-3-large if:

  • Accuracy is critical (compliance, medical, legal domains)
  • Cost isn't a constraint
  • You can use higher dimensions (3072)

Use text-embedding-3-small if:

  • High query volume (>1M/month)
  • Cost-sensitive
  • A small accuracy tradeoff vs. 3-large is acceptable

Use open-source (bge-large) if:

  • Can self-host (removes per-token embedding costs)
  • A modest accuracy tradeoff is acceptable
  • Data privacy requires on-prem

Dimensionality impact

text-embedding-3-large can return shortened embeddings (for example 768 or 1536 dimensions). Fewer dimensions cut storage and query latency, at some cost in accuracy.

Recommendation: Use full 3072 dimensions unless storage costs are prohibitive.

Retrieval method performance

Retrieval method had the largest impact on accuracy.

Retrieval methodAccuracyPrecisionRecallLatency
Pure vector similarityBaselineModerateModerateLow
Pure BM25 (keyword)LowerLowerGoodVery low
Hybrid (vector + BM25)HighestHighHighModerate

Hybrid search typically improves accuracy substantially over pure vector.

Why hybrid search wins

Vector search and keyword search fail in different ways:

Vector search weaknesses:

  • Struggles with exact matches (product codes, error messages)
  • Poor at rare terms not well-represented in embeddings
  • Misses queries with specific keyword requirements

Example query: "What's error code E4701?"

Vector search might return documents about "error handling" generally. Keyword search finds the exact code.

Keyword search (BM25) weaknesses:

  • No semantic understanding
  • Fails on paraphrases and synonyms
  • Sensitive to vocabulary mismatch

Example query: "How do I reset my password?"

Keyword search misses documents using "credential recovery" or "account access restoration" instead of exact phrase "reset password."

Hybrid combines strengths:

def hybrid_search(query: str, vector_weight: float = 0.7):
    """Combine vector and keyword search."""

    # Vector search
    query_embedding = embed_query(query)
    vector_results = vector_db.search(query_embedding, top_k=20)

    # Keyword search (BM25)
    keyword_results = bm25_index.search(query, top_k=20)

    # Combine scores (normalize first)
    combined_scores = {}
    for doc_id, score in vector_results:
        combined_scores[doc_id] = score * vector_weight

    for doc_id, score in keyword_results:
        combined_scores[doc_id] = combined_scores.get(doc_id, 0) + score * (1 - vector_weight)

    # Rank by combined score
    ranked = sorted(combined_scores.items(), key=lambda x: x[1], reverse=True)
    return ranked[:5]  # Top 5

Optimal weighting

Accuracy usually peaks with vector search carrying most of the weight and keyword search a meaningful minority. Pure vector and pure keyword both do worse than any sensible blend.

Recommendation: Use 70% vector, 30% keyword as default. Tune per use case.

Query type breakdown

Different query types favor different retrieval methods:

Query typeBest method
Factual lookupsHybrid
ResearchVector (90%) + Keyword (10%)
TroubleshootingKeyword (60%) + Vector (40%)
ComparisonVector

Insight: Troubleshooting queries benefit from higher keyword weighting because they often include specific error codes or log messages.

Combined optimization results

Combining the best option from each dimension:

Optimized pipeline:

  • Chunking: 500 tokens, 20% overlap
  • Embedding: text-embedding-3-large (3072d)
  • Retrieval: Hybrid (70% vector, 30% BM25)
  • Top-k: 5 chunks

Compared with the baseline, expect:

  • Clearly better accuracy, precision and recall
  • Somewhat higher latency (usually still acceptable for most use cases)
  • Higher cost per query, driven mainly by the larger embedding model

ROI: For most applications, the accuracy improvement justifies the extra cost. Check the per-query numbers on your own traffic.

Latency vs. accuracy trade-offs

Different use cases prioritize speed vs. accuracy differently.

Use caseAcceptable latencyTarget accuracyRecommended config
Chatbot (customer-facing)<300msModerateVector only, text-emb-3-small, 500 tokens no overlap
Internal knowledge search<500msHighHybrid, text-emb-3-large, 500 tokens 20% overlap
Compliance/Legal<1000msVery highHybrid + reranker, text-emb-3-large, semantic chunking
Batch processingNo constraintVery highFull optimization + GPT-4 verification

Adding a reranker

For use cases that need the highest accuracy, add a reranker stage:

1. Hybrid search retrieves top 20 candidates (cheap, fast)
2. Reranker (e.g., Cohere rerank, cross-encoder) reorders top 20 (expensive, accurate)
3. Select top 5 from reranked list

Impact:

  • Accuracy: a further improvement
  • Latency: noticeably higher
  • Cost: roughly doubles per query or more, depending on the reranker

Recommendation: Use reranker for high-stakes queries (legal, compliance, medical). Skip for general knowledge retrieval.

Cost optimization strategies

RAG costs add up at scale. Optimization strategies:

1. Tiered retrieval

Use cheap search first, escalate to expensive methods only if needed:

Query arrives
└─> Try BM25 keyword search (fast, cheap)
    └─> If confidence <0.8:
        └─> Try vector search
            └─> If confidence <0.8:
                └─> Try hybrid + reranker

Result: Queries that keyword search can answer never touch the more expensive stages.

2. Cache popular queries

Store results for frequently-asked questions:

from functools import lru_cache

@lru_cache(maxsize=1000)
def retrieve_with_cache(query: str):
    """Cache results for repeated queries."""
    # Normalize query (lowercase, remove punctuation)
    normalized = normalize(query)

    # Check cache
    if cached_result := cache.get(normalized):
        return cached_result

    # Perform retrieval
    result = hybrid_search(query)

    # Cache result
    cache.set(normalized, result, ttl=3600)  # 1 hour TTL
    return result

Result: Every cache hit skips retrieval entirely.

3. Use smaller embeddings for low-stakes queries

Route chatbot queries to text-emb-3-small, route compliance queries to text-emb-3-large:

def get_embedding_model(query_type: str):
    """Select embedding model based on query importance."""
    if query_type in ["compliance", "legal", "financial"]:
        return "text-embedding-3-large"
    else:
        return "text-embedding-3-small"  # 6.5× cheaper per token at list price

Result: Lower costs with minimal accuracy impact on low-stakes queries.

4. Batch embeddings

Embed in batches of 100-1000 instead of one-by-one:

# Bad: One at a time
for doc in documents:
    embedding = client.embeddings.create(input=doc, model="text-embedding-3-large")

# Good: Batched
batch_size = 100
for i in range(0, len(documents), batch_size):
    batch = documents[i:i+batch_size]
    embeddings = client.embeddings.create(input=batch, model="text-embedding-3-large")

Result: Far fewer API calls and less per-request overhead.

Failure mode analysis

Even an optimized pipeline gets some queries wrong. The failures usually fall into these buckets, roughly from most to least common:

Failure modeExample
Answer not in knowledge baseQuery: "What's our policy on X?" → No doc covers X
Requires multi-hop reasoningQuery needs info from 3+ disconnected chunks
Ambiguous query"How do I set it up?" → What's "it"?
Outdated informationRetrieved chunk is from old version of docs
Retrieval failure (bad chunks)Relevant chunks exist but weren't retrieved

Addressing failure modes

Answer not in KB:

  • Detect using confidence scoring: if top retrieval score <0.6, respond "I don't have information on that"
  • Avoid hallucination by refusing to answer instead of guessing

Multi-hop reasoning:

  • Use agentic RAG: retrieve, synthesize, retrieve again if needed
  • Or: expand context window to include more chunks (5 → 10)

Ambiguous queries:

  • Add clarification step: "Did you mean X or Y?"
  • Use conversation history to resolve pronouns ("it," "that," "this")

Outdated information:

  • Add metadata: last_updated timestamp on chunks
  • Prefer recent chunks when dates are close
  • Implement versioned knowledge base

Retrieval failure:

  • Add query expansion: rewrite query in multiple ways, retrieve for each
  • Use HyDE (Hypothetical Document Embeddings): generate a hypothetical answer, embed it, search for similar docs

Recommendations by use case

Customer support chatbot

Priority: Low latency, reasonable accuracy, low cost

Config:

  • Chunking: 500 tokens, 10% overlap
  • Embedding: text-embedding-3-small
  • Retrieval: Vector only (skip hybrid for speed)
  • Top-k: 3
  • Cache: Yes (1-hour TTL)

Expected: Reasonable accuracy, low latency, lowest cost per query

Internal knowledge assistant

Priority: High accuracy, moderate latency acceptable

Config:

  • Chunking: 500 tokens, 20% overlap
  • Embedding: text-embedding-3-large
  • Retrieval: Hybrid (70% vector, 30% keyword)
  • Top-k: 5
  • Reranker: Optional

Expected: High accuracy, moderate latency, moderate cost per query

Compliance/Legal document search

Priority: Maximum accuracy, latency not critical

Config:

  • Chunking: Semantic (preserve document structure)
  • Embedding: text-embedding-3-large (3072d)
  • Retrieval: Hybrid + Cohere reranker
  • Top-k: 10 → rerank to 5
  • Verification: GPT-4 checks answer against source

Expected: Highest accuracy, latency under a second, highest cost per query

Real-time code documentation

Priority: Very low latency, good accuracy

Config:

  • Chunking: Function-level (preserve code blocks)
  • Embedding: bge-large (self-hosted)
  • Retrieval: BM25 keyword (function names, class names)
  • Top-k: 3
  • Cache: Aggressive (24-hour TTL)

Expected: Good accuracy, very low latency, ~$0/query (self-hosted)

Implementation checklist

Week 1: Baseline measurement

  • [ ] Collect 100-500 representative queries
  • [ ] Establish ground truth answers
  • [ ] Measure baseline accuracy with current RAG setup
  • [ ] Measure baseline latency and cost

Week 2: Chunking optimization

  • [ ] Test 500 tokens with 0%, 10%, 20% overlap
  • [ ] Measure accuracy impact
  • [ ] Select optimal overlap percentage

Week 3: Retrieval upgrade

  • [ ] Implement BM25 keyword search
  • [ ] Build hybrid search combining vector + BM25
  • [ ] Test weight ratios (70/30, 60/40, 80/20)
  • [ ] Measure accuracy improvement

Week 4: Embedding optimization

  • [ ] Test text-embedding-3-large
  • [ ] Measure accuracy vs. cost trade-off
  • [ ] Decide on embedding model

Week 5: Production rollout

  • [ ] Deploy optimized config to 10% of traffic
  • [ ] Monitor accuracy, latency, cost for 1 week
  • [ ] If successful, roll out to 100%

Ongoing:

  • [ ] Monthly review of failure cases
  • [ ] Retune hybrid weights based on query distribution
  • [ ] Update knowledge base regularly

Tools and libraries

Vector databases:

  • Pinecone (managed, easy): Good for getting started
  • Weaviate (hybrid search built-in): Best for hybrid retrieval
  • Qdrant (open-source, fast): Good for self-hosting
  • PostgreSQL + pgvector (familiar stack): Good if already using Postgres

BM25 implementations:

  • Elasticsearch: Industry standard, mature
  • Typesense: Faster, simpler API
  • rank-bm25 (Python library): Lightweight, for prototyping

Rerankers:

  • Cohere Rerank API: Easiest, $1/1000 searches
  • Cross-encoders (ms-marco-MiniLM): Self-hostable
  • Voyage Rerank: Alternative to Cohere

Evaluation frameworks:

  • RAGAS: RAG evaluation metrics (faithfulness, relevance)
  • LangSmith: End-to-end RAG pipeline testing
  • PromptLayer: A/B testing for RAG configs

Key takeaways

  • Hybrid retrieval is the highest-leverage optimization, combining vector semantic search with keyword exactness.
  • 500-token chunks with 20% overlap outperform both smaller chunks (lose context) and larger chunks (noise).
  • Embedding model matters but not as much as retrieval method -text-embedding-3-large adds only a modest gain over 3-small for several times the cost.
  • Different use cases need different configs -chatbots prioritize speed, compliance prioritizes accuracy, batch processing optimizes for both.
  • Measurement is prerequisite to optimization -establish ground truth, measure baseline, test systematically.

---

RAG pipeline optimization isn't one-size-fits-all. The "best" configuration depends on your accuracy requirements, latency constraints, and cost budget. Start with hybrid retrieval (biggest bang for buck), dial in chunking strategy, then optimize embedding model if accuracy still falls short. Measure continuously and retune as your knowledge base and query distribution evolve.

Frequently asked questions

Q: Should I optimize RAG before or after prompt engineering?

A: Do basic prompt engineering first (clear instructions, few-shot examples) to establish a baseline. Then optimize RAG. Advanced prompt engineering can compensate for poor RAG but wastes tokens and increases costs.

Q: How often should I retune RAG parameters?

A: Review monthly for first 6 months, then quarterly. Retune immediately if you notice accuracy degradation or if your knowledge base content changes significantly (e.g., docs rewrite, new product launch).

Q: Can I use different RAG configs for different document types?

A: Yes! Route queries to specialized indices: structured docs use fixed chunking + keyword search, narrative content uses semantic chunking + vector search.

Q: What's the minimum dataset size to run meaningful RAG experiments?

A: 50-100 queries with ground truth answers. Below that, results aren't statistically significant. Above 500, diminishing returns on experiment value.

Further reading:

External references:

---

Frequently Asked Questions

Q: What skills do I need to build AI agent systems?

You don't need deep AI expertise to implement agent workflows. Basic understanding of APIs, workflow design, and prompt engineering is sufficient for most use cases. More complex systems benefit from software engineering experience, particularly around error handling and monitoring.

Q: What's the typical ROI timeline for AI agent implementations?

Many organisations see positive ROI within a few months of deployment. Initial productivity gains are often noticeable, with improvements compounding as teams optimise prompts and workflows based on production experience.

Q: How long does it take to implement an AI agent workflow?

Implementation timelines vary based on complexity, but most teams see initial results within 2-4 weeks for simple workflows. More sophisticated multi-agent systems typically require 6-12 weeks for full deployment with proper testing and governance.

More from the blog

Stop doing the work around the work

OpenHelm connects to your tools, reads the context, and does the steps, so you sign off on the result instead of producing it. See how it covers an entire role’s weekly workload, check the pricing, or run it yourself with the free local app.