Skip to content
Academy

Cost Optimization Strategies for LLM-Based Agents

Practical ways to cut agent running costs without losing quality: model selection, prompt compression, caching and batching.

M
Max Beech· Founder
··11 min read
Cost Optimization Strategies for LLM-Based Agents

TL;DR

  • LLM costs often dominate agent economics -for typical production systems, API calls are the largest share of operational spend.
  • The strategies below can cut cost per task substantially without hurting task success rate.
  • Biggest wins usually come from smart model routing, prompt compression, and response caching. Combined effect compounds.

Jump to Cost breakdown · Jump to Model selection · Jump to Prompt optimization · Jump to Caching · Jump to Batching

# Cost Optimization Strategies for LLM-Based Agents

Imagine an agent that is working well: tasks succeed, users are happy, and the monthly API bill grows with every task. Extrapolate that to your growth target and the bill can quickly outrun your budget.

You have two choices: accept that agent economics don't work at scale, or find a way to cut costs radically without breaking quality.

This guide covers how to do the second.

Understanding agent cost structure

Before optimizing, understand where money goes.

Typical agent cost breakdown

Cost componentTypical share of total
LLM API callsLargest
Tool API calls (enrichment, search)Significant
Infrastructure (hosting, DB)Small
Monitoring & logsSmall

LLM costs dominate. That's where optimization yields biggest returns.

LLM cost drivers

LLM cost = (Input tokens × Input price) + (Output tokens × Output price)

Example: GPT-4 Turbo call

  • Input: 2,500 tokens @ £0.01/1K tokens = £0.025
  • Output: 800 tokens @ £0.03/1K tokens = £0.024
  • Total: £0.049 per call

If a task averages 8 LLM calls → £0.39/task just for LLM

Three levers to pull:

  1. Reduce tokens (input + output)
  2. Use cheaper models (GPT-3.5 vs GPT-4)
  3. Reduce calls (fewer LLM invocations per task)

Strategy 1: Smart model selection

Not every task needs GPT-4 Opus. Route intelligently based on complexity.

Model tier strategy

TierModelsCost/1M tokensUse for
PremiumGPT-4, Claude Opus£30-60Complex reasoning, code generation, analysis
StandardGPT-4 Turbo, Claude Sonnet£10-15General tasks, multi-step workflows
EconomyGPT-3.5, Claude Haiku£0.50-3Classification, summarization, simple Q&A
BudgetMixtral, Llama 3 (self-hosted)~£0High-volume, latency-tolerant tasks

Implementation:

def select_model(task_complexity: str, task_type: str) -> str:
    """Route to appropriate model based on task."""

    # High-complexity tasks → premium models
    if task_complexity == "high" or task_type in ["code_generation", "research_synthesis"]:
        return "gpt-4-turbo"

    # Medium complexity → standard models
    elif task_complexity == "medium" or task_type in ["analysis", "planning"]:
        return "gpt-3.5-turbo"

    # Simple tasks → economy models
    elif task_type in ["classification", "extraction", "summarization"]:
        return "gpt-3.5-turbo"  # or claude-haiku

    # Default to standard
    return "gpt-4-turbo"

# Enhanced with confidence-based routing
def select_model_adaptive(task: dict, previous_attempts: int = 0) -> str:
    """Start cheap, escalate if needed."""

    # First attempt: try economy model
    if previous_attempts == 0:
        return "gpt-3.5-turbo"

    # If failed or low confidence, escalate to standard
    elif previous_attempts == 1:
        return "gpt-4-turbo"

    # Last resort: premium model
    else:
        return "gpt-4"

What routing typically changes

With tiered routing, most simple tasks move to economy models, a smaller share runs on standard models, and only the genuinely hard cases reach the premium tier. The blended cost per task falls.

Quality often holds or even improves: cheaper models handle simple tasks quickly without overthinking, and the premium model is reserved for cases that need it. Measure success rate before and after to confirm this on your own workload.

When to self-host

For very high volume, self-hosting open models (Llama 3, Mixtral) can make economic sense:

Break-even analysis:

ScenarioAPI-based (GPT-3.5)Self-hosted (Llama 3 70B)
Setup costNoneSignificant (GPU servers)
Monthly cost at moderate volumeLowerHigher (fixed infrastructure)
Monthly cost at high volumeGrows in line with usageGrows slowly

Break-even: Only at high volume; run the numbers with your own hardware quotes and API pricing.

Trade-offs:

  • ✅ Unlimited usage above break-even
  • ✅ Data stays in your infrastructure
  • ❌ Engineering overhead (deployment, monitoring)
  • ❌ Lower quality than GPT-4 (acceptable for many use cases)

Strategy 2: Prompt compression

Shorter prompts = lower input token costs.

Techniques

1. Remove redundancy

Before (182 tokens):

You are a helpful AI assistant designed to help users with customer support queries. Please analyse the following customer support ticket carefully and provide a detailed, helpful response that addresses all of the customer's concerns. Make sure your response is professional, empathetic, and actionable.

Customer query: [...]

After (89 tokens, -51%):

Analyse this support ticket and provide a professional, actionable response.

Query: [...]

Savings: £0.001 per call × 8 calls/task × 40K tasks/month = £320/month

2. Use structured formats

Before (verbose):

Please extract the following information from the document: the customer's name, their email address, their company name, their job title, and the date they signed up.

After (JSON schema):

Extract to JSON:
{"name": "", "email": "", "company": "", "title": "", "signup_date": ""}

Token reduction: Noticeably fewer tokens

3. Eliminate few-shot examples when possible

Few-shot examples (showing the model examples before the task) improve quality but cost tokens.

Test whether they're necessary:

def test_fewshot_necessity(task_sample: list, prompt_with_examples: str, prompt_without_examples: str):
    """A/B test few-shot vs zero-shot."""
    results_with = []
    results_without = []

    for task in task_sample:
        # With examples
        response_with = llm.complete(prompt_with_examples + task)
        results_with.append(evaluate_quality(response_with, task))

        # Without examples
        response_without = llm.complete(prompt_without_examples + task)
        results_without.append(evaluate_quality(response_without, task))

    print(f"With few-shot: {np.mean(results_with):.2%} quality")
    print(f"Without few-shot: {np.mean(results_without):.2%} quality")
    print(f"Token savings: {calculate_token_diff(prompt_with_examples, prompt_without_examples)}")

# Weigh the quality difference against the token savings

A common outcome: remove few-shot examples for simple tasks (classification, extraction), keep them for complex tasks (code generation, analysis).

Prompt caching

Some LLM providers (Anthropic Claude, OpenAI with prompt caching beta) allow caching prompt prefixes.

How it works:

# First call: full cost
response = client.messages.create(
    model="claude-3-sonnet",
    system="You are a customer support agent. Here's our knowledge base: [5,000 tokens of docs]",
    messages=[{"role": "user", "content": "How do I reset my password?"}]
)
# Cost: 5,100 input tokens

# Subsequent calls within 5 minutes: cached system prompt
response = client.messages.create(
    model="claude-3-sonnet",
    system="You are a customer support agent. Here's our knowledge base: [5,000 tokens of docs]",  # CACHED
    messages=[{"role": "user", "content": "How do I change my email?"}]
)
# Cost: 100 input tokens (only the new message)

Savings: 90%+ on input tokens for repeated prompts

Limitations:

  • Cache expires after 5 minutes (Anthropic) or 1 hour (OpenAI)
  • Only works if system prompt is identical across calls
  • Cache misses still cost full tokens

Use cases:

  • Chatbots (same knowledge base for all queries)
  • Document processing (same instructions, different docs)
  • Multi-turn conversations

Strategy 3: Intelligent caching

Cache LLM responses to avoid redundant calls.

Response caching for repeated queries

import hashlib
from functools import lru_cache

class LLMCache:
    """Cache LLM responses."""

    def __init__(self, ttl: int = 3600):
        self.cache = {}
        self.ttl = ttl

    def get(self, prompt: str, model: str) -> str|None:
        """Get cached response if exists."""
        cache_key = self._hash(prompt, model)
        entry = self.cache.get(cache_key)

        if entry and time.time() - entry["timestamp"] < self.ttl:
            return entry["response"]

        return None

    def set(self, prompt: str, model: str, response: str):
        """Cache response."""
        cache_key = self._hash(prompt, model)
        self.cache[cache_key] = {
            "response": response,
            "timestamp": time.time()
        }

    def _hash(self, prompt: str, model: str) -> str:
        """Generate cache key."""
        return hashlib.sha256(f"{model}:{prompt}".encode()).hexdigest()

# Usage
cache = LLMCache(ttl=3600)  # 1-hour TTL

def cached_llm_call(prompt: str, model: str):
    """Call LLM with caching."""
    # Check cache
    cached_response = cache.get(prompt, model)
    if cached_response:
        return cached_response

    # Cache miss, call LLM
    response = llm.complete(prompt, model=model)

    # Store in cache
    cache.set(prompt, model, response)

    return response

Results: Savings scale directly with your cache hit rate, which varies a lot by use case.

Expected cache hit rate by use case:

Use caseHit rateWhy
FAQ chatbotHighRepeated questions
Document summarizationLowUnique documents
Code reviewMediumCommon patterns
Customer supportMedium-HighSimilar queries

Semantic caching

Standard caching requires exact prompt match. Semantic caching matches similar prompts:

from sentence_transformers import SentenceTransformer
import numpy as np

class SemanticCache:
    """Cache based on semantic similarity."""

    def __init__(self, similarity_threshold: float = 0.95):
        self.embedder = SentenceTransformer('all-MiniLM-L6-v2')
        self.cache = []  # List of (embedding, response) tuples
        self.similarity_threshold = similarity_threshold

    def get(self, prompt: str) -> str|None:
        """Find semantically similar cached response."""
        if not self.cache:
            return None

        # Embed query
        query_embedding = self.embedder.encode(prompt)

        # Find most similar cached prompt
        for cached_embedding, cached_response in self.cache:
            similarity = np.dot(query_embedding, cached_embedding)

            if similarity > self.similarity_threshold:
                return cached_response

        return None

    def set(self, prompt: str, response: str):
        """Cache response with prompt embedding."""
        embedding = self.embedder.encode(prompt)
        self.cache.append((embedding, response))

        # Limit cache size
        if len(self.cache) > 1000:
            self.cache.pop(0)  # Remove oldest

# Example
cache = SemanticCache()

# First query
response_1 = llm.complete("How do I reset my password?")
cache.set("How do I reset my password?", response_1)

# Similar query (different wording) → cache hit!
response_2 = cache.get("What's the process for resetting my password?")
# Returns cached response_1 (95%+ similarity)

Trade-off: Embedding cost (£0.00002/query) vs. LLM call savings (£0.05/query) → 2,500× ROI

Strategy 4: Batching and parallelization

Process multiple items in one LLM call instead of many sequential calls.

Batch processing

Before (sequential, £0.40):

for email in emails:
    classification = llm.classify_email(email)
    # 10 emails × £0.04/call = £0.40

After (batched, £0.08):

batch_prompt = f"""
Classify these 10 emails as spam/not spam:

{format_emails(emails)}

Return JSON array: [{{"email_id": 1, "classification": "spam"}}, ...]
"""
classifications = llm.complete(batch_prompt)
# 1 call × £0.08 = £0.08 (-80% cost)

Limitations:

  • Batch size limited by context window (can't fit 1,000 emails)
  • Quality may degrade for very large batches (model loses focus)
  • Single failure affects entire batch

Optimal batch size: Test 5, 10, 25, 50 items and pick the largest size before quality starts to slip.

Parallel tool calls

Many agents make sequential tool calls. Enable parallelization:

Before (sequential, 3.2s latency):

result_1 = fetch_data_from_api_1()  # 800ms
result_2 = fetch_data_from_api_2()  # 1,200ms
result_3 = fetch_data_from_api_3()  # 1,200ms
# Total: 3,200ms

After (parallel, 1.2s latency):

import asyncio

results = await asyncio.gather(
    fetch_data_from_api_1(),
    fetch_data_from_api_2(),
    fetch_data_from_api_3()
)
# Total: 1,200ms (longest call)

Cost impact: Indirect -faster execution = better user experience = higher agent adoption = more value from agent investment.

Strategy 5: Output length control

LLMs often over-generate. Constrain output to save tokens.

Techniques

1. Explicit length limits

prompt = f"""
Summarise this article in EXACTLY 3 sentences. No more, no less.

Article: {article_text}
"""

2. Token limits (max_tokens parameter)

response = client.completions.create(
    model="gpt-4-turbo",
    prompt=prompt,
    max_tokens=100  # Hard cap at 100 output tokens
)

3. Structured outputs (JSON)

Before (free-form, 400 tokens average):

"The customer seems frustrated about the delayed shipment. They ordered on Jan 15th and expected delivery by Jan 20th but haven't received it yet..."

After (JSON, 80 tokens):

{
  "sentiment": "frustrated",
  "issue": "delayed_shipment",
  "order_date": "2024-01-15",
  "expected_delivery": "2024-01-20",
  "status": "not_received"
}

Savings: 80% fewer output tokens

Strategy 6: Streaming for user experience

Streaming doesn't reduce costs but improves perceived performance:

def stream_response(prompt: str):
    """Stream LLM response token-by-token."""
    for chunk in client.completions.create(
        model="gpt-4-turbo",
        prompt=prompt,
        stream=True
    ):
        yield chunk.choices[0].text

# Display to user immediately
for token in stream_response(user_query):
    print(token, end="", flush=True)

Benefit: User sees response start in 200ms instead of waiting 3s for full completion.

Cost: Identical to non-streaming

Strategy 7: Fine-tuning for efficiency

Fine-tuned models need shorter prompts to achieve same quality.

Example: Customer support classification

Base model (GPT-3.5,450-token prompt with examples):

You are a customer support classifier. Examples:
[10 examples, 400 tokens]

Classify this ticket: [50 tokens]

Cost: £0.00045/call

Fine-tuned model (GPT-3.5 fine-tuned on 500 examples):

Classify: [50 tokens]

Cost: £0.00005/call (90% cheaper)

Fine-tuning costs:

  • Training: £50 one-time (500 examples)
  • Inference: 10% cheaper per call
  • Break-even: 50,000 calls

When to fine-tune:

  • High-volume tasks (>10K/month)
  • Repeated patterns (classification, extraction, formatting)
  • Quality ceiling reached with prompting

Strategy 8: Monitoring and alerting

Track costs in real-time to catch spikes:

class CostMonitor:
    """Track LLM costs per task."""

    def __init__(self):
        self.task_costs = []

    def record_task_cost(self, task_id: str, cost: float):
        """Log task cost."""
        self.task_costs.append({
            "task_id": task_id,
            "cost": cost,
            "timestamp": datetime.utcnow()
        })

        # Alert if anomaly
        recent_avg = np.mean([t["cost"] for t in self.task_costs[-100:]])

        if cost > recent_avg * 3:  # 3× average cost
            self.alert_anomaly(task_id, cost, recent_avg)

    def alert_anomaly(self, task_id: str, cost: float, avg: float):
        """Alert on cost spike."""
        send_slack_alert(f"⚠️ Cost anomaly: Task {task_id} cost £{cost:.4f} (avg: £{avg:.4f})")

# Daily summary
def daily_cost_report():
    """Generate cost summary."""
    today_tasks = [t for t in monitor.task_costs if is_today(t["timestamp"])]

    report = {
        "total_cost": sum(t["cost"] for t in today_tasks),
        "task_count": len(today_tasks),
        "avg_cost_per_task": np.mean([t["cost"] for t in today_tasks]),
        "max_cost": max([t["cost"] for t in today_tasks]),
        "p95_cost": np.percentile([t["cost"] for t in today_tasks], 95)
    }

    return report

Combining the strategies

Baseline (no optimizations):

  • Model: GPT-4 Turbo for all tasks
  • Prompts: Verbose with few-shot examples
  • No caching
  • Sequential processing

Optimized (all strategies):

  • Smart model routing (GPT-3.5 → GPT-4 escalation)
  • Compressed prompts
  • Response caching
  • Batched processing where applicable
  • Structured outputs (JSON)

The LLM API line shrinks the most, because routing, compression, caching and structured outputs all act on it at once. Measure cost per task before and after each change so you know which strategies are paying off for your workload.

Implementation roadmap

Week 1: Baseline measurement

  • [ ] Instrument all LLM calls to log tokens and costs
  • [ ] Calculate current cost/task
  • [ ] Identify top 3 cost drivers

Week 2: Quick wins

  • [ ] Implement response caching
  • [ ] Compress prompts (remove redundancy)
  • [ ] Add max_tokens limits

Week 3: Model routing

  • [ ] Define task complexity tiers
  • [ ] Implement routing logic
  • [ ] A/B test quality vs baseline

Week 4: Advanced optimizations

  • [ ] Batch eligible tasks
  • [ ] Test fine-tuning for high-volume tasks
  • [ ] Set up cost monitoring and alerts

Month 2+:

  • [ ] Continuous optimization based on cost analytics
  • [ ] Explore self-hosting for very high volumes
  • [ ] Regular review of model pricing (providers update frequently)

Key takeaways

  • LLM costs dominate agent economics -optimizing inference costs is critical for scalability.
  • Smart model routing offers biggest single win -route simple tasks to cheap models, escalate complex tasks to expensive models.
  • Caching delivers immediate ROI -a cache lookup costs a tiny fraction of the LLM call it saves.
  • Optimizations compound -combining strategies multiplies the savings rather than simply adding them.
  • Quality doesn't have to suffer -with careful routing and testing, task success can hold steady while costs fall.

---

Agent economics improve dramatically with deliberate cost optimization. Start with model routing and caching for quick wins, then layer in prompt compression, batching, and fine-tuning as volume scales. The goal isn't minimum cost -it's maximum value per pound spent.

Frequently asked questions

Q: Will cheaper models hurt quality?

A: For many tasks, no. GPT-3.5 handles classification, extraction, and simple Q&A well for a fraction of GPT-4's cost. Test on your use case.

Q: How do I know if optimizations are working?

A: Track cost/task weekly. If cost drops but task success rate stays flat or improves, you're winning.

Q: Should I optimize before launching or after?

A: Get to product-market fit first. Optimize once you have consistent usage and understand cost drivers. Premature optimization wastes time.

Q: What's a good cost/task target?

A: Depends on value delivered. If agent saves £2 in human time per task, £0.50/task is excellent ROI. If it saves £0.50, you need <£0.10/task.

Further reading:

External references:

---

Frequently Asked Questions

Q: What's the typical ROI timeline for AI agent implementations?

Well-scoped implementations often pay back within a few months. Gains tend to compound as teams optimise prompts and workflows based on production experience.

Q: How do AI agents handle errors and edge cases?

Well-designed agent systems include fallback mechanisms, human-in-the-loop escalation, and retry logic. The key is defining clear boundaries for autonomous action versus requiring human approval for sensitive or unusual situations.

Q: How long does it take to implement an AI agent workflow?

Implementation timelines vary based on complexity, but most teams see initial results within 2-4 weeks for simple workflows. More sophisticated multi-agent systems typically require 6-12 weeks for full deployment with proper testing and governance.

More from the blog

Stop doing the work around the work

OpenHelm connects to your tools, reads the context, and does the steps, so you sign off on the result instead of producing it. See how it covers an entire role’s weekly workload, check the pricing, or run it yourself with the free local app.