Prompt Engineering for Production AI Agents: Techniques That Actually Work
Seven prompt engineering techniques that make agents more reliable in production: few-shot examples, structured output, chain-of-thought and more.

TL;DR
- Most prompt engineering advice is cargo cult nonsense. Here are 7 techniques that reliably earn their place.
- Few-shot examples (2-3): a clear accuracy gain over zero-shot on classification tasks
- Structured output format: JSON schema enforcement all but eliminates parsing errors
- Chain-of-thought: helps on reasoning tasks, but adds latency -use selectively
- Negative examples: Showing what NOT to do improves edge case handling
- Temperature tuning: 0.0-0.3 for consistent output, 0.7-1.0 for creative tasks
- Measure each change on your own production queries before adopting it
# Prompt Engineering for Production AI Agents
The internet is full of prompt engineering tips. "Add 'Let's think step by step!'" "Use role-playing!" "Say please!"
On production workloads (customer support, data extraction, content generation), most of these tricks make no measurable difference, and some make things worse.
Here are 7 techniques that actually move reliability metrics.
Technique 1: Few-Shot Examples (2-3 Optimal)
Claim: Showing examples improves performance.
Reality: True, but more isn't always better.
Test Setup
Task: Classify customer support tickets into categories (Bug, Feature Request, Question, Complaint)
Zero-shot (no examples):
Classify this ticket: {ticket_text}
Categories: Bug, Feature Request, Question, ComplaintFew-shot (3 examples):
Classify customer support tickets.
Examples:
Ticket: "App crashes when I upload images"
Category: Bug
Ticket: "Can you add dark mode?"
Category: Feature Request
Ticket: "How do I reset my password?"
Category: Question
Now classify:
Ticket: {ticket_text}
Category:What to expect
| Approach | Typical effect |
|---|---|
| Zero-shot | Baseline |
| 1 example | Noticeable improvement |
| 2-3 examples | Usually the sweet spot |
| 5+ examples | No better, sometimes worse, and more tokens |
Optimal: 2-3 well-chosen examples. More examples add noise and cost without reliably improving accuracy.
Why diminishing returns? LLMs pattern-match. 2-3 examples establish pattern. 10 examples create ambiguity (which pattern to follow?).
Implementation
def build_few_shot_prompt(task_description, examples, query):
"""
examples = [
{"input": "...", "output": "..."},
{"input": "...", "output": "..."}
]
"""
prompt = f"{task_description}\n\nExamples:\n"
for ex in examples[:3]: # Limit to 3
prompt += f"Input: {ex['input']}\nOutput: {ex['output']}\n\n"
prompt += f"Now:\nInput: {query}\nOutput:"
return promptPro tip: Choose diverse examples covering edge cases, not just happy path.
Technique 2: Structured Output Enforcement
Problem: LLMs return text. You need JSON. Without enforcement, parsing fails often enough to break pipelines.
Solution: Enforce output format in prompt + use structured output APIs.
Before (Unreliable)
prompt = """
Extract company name, revenue, and industry from this text:
{text}
Return as JSON.
"""
# Model returns:
"The company is Acme Corp. Their revenue is $50M. Industry: SaaS"
# Or: {"company": "Acme Corp", revenue: "$50M", "industry": "SaaS"} # Invalid JSON
# Or: Here's the extracted data: {"company": "Acme Corp", ...} # Extra textParse success rate: unreliable
After (Reliable)
prompt = """
Extract information and return ONLY valid JSON matching this schema:
{
"company_name": string,
"revenue_usd": number (no currency symbols),
"industry": string
}
Text: {text}
JSON:
"""
# Use OpenAI's response_format parameter
response = client.chat.completions.create(
model="gpt-4-turbo",
messages=[{"role": "user", "content": prompt}],
response_format={"type": "json_object"} # Enforces JSON
)Parse success rate: near-perfect
Results Table
| Method | Valid JSON | Correct Data | Production Ready |
|---|---|---|---|
| No guidance | Unreliable | Low | ❌ |
| Prompt: "Return JSON" | Better | Medium | ❌ |
| + Schema example | Good | Good | ⚠️ |
+ response_format | Very high | High | ✅ |
Technique 3: Chain-of-Thought (Use Selectively)
Claim: Adding "Let's think step by step" improves reasoning.
Reality: True for complex reasoning. Overkill for simple tasks.
When Chain-of-Thought Helps
Complex reasoning task (math word problem):
Without CoT:
Q: A shop sells pens at 3 for £2. How much do 12 pens cost?
A: £6 ❌ (incorrect: the model multiplied 12 by 0.5)With CoT:
Q: A shop sells pens at 3 for £2. How much do 12 pens cost?
Let's think step by step:
1. 12 pens is 12 / 3 = 4 groups of 3
2. Each group costs £2
3. 4 × £2 = £8
A: £8 ✅ (correct)Where it pays off
| Task Type | Benefit of CoT | Latency Impact |
|---|---|---|
| Math problems | Large | Significant |
| Logic puzzles | Large | Significant |
| Multi-step reasoning | Large | Significant |
| Simple classification | Negligible ❌ | Significant |
| Fact lookup | None ❌ | Significant |
Use CoT when: Multi-step reasoning, math, logic
Skip CoT when: Classification, lookup, simple Q&A
Cost-benefit: CoT adds noticeable latency and several times the output tokens. Only use it when the accuracy gain justifies the cost.
Technique 4: Negative Examples
Showing what NOT to do improves edge case handling.
Example: Email Classification
Without negative examples:
Classify emails as Spam or Not Spam.
Email: "URGENT: Your account will be suspended"
Classification: Spam ❌ (False positive - legitimate security alert)With negative examples:
Classify emails as Spam or Not Spam.
Example (Spam):
"Congratulations! You won $1M! Click here!!!"
→ Spam
Example (NOT Spam - even if urgent):
"Security alert: Unusual login detected from new device"
→ Not Spam
Email: "URGENT: Your account will be suspended"
Classification: Not Spam ✅ (Correct)The biggest gain shows up on edge cases, with fewer false positives, rather than on overall accuracy.
When to use: Tasks with tricky edge cases, high cost of false positives/negatives.
Technique 5: Temperature Tuning
Temperature controls randomness. Most people use default (1.0). Wrong for many tasks.
Temperature Guide
| Temperature | Behavior | Use Case |
|---|---|---|
| 0.0 | Deterministic, same output every time | Classification, data extraction, structured tasks |
| 0.3 | Mostly consistent, slight variation | Customer support, Q&A |
| 0.7 | Balanced creativity/consistency | Content summarization |
| 1.0 | Creative, diverse outputs | Content generation, brainstorming |
| 1.5+ | Very random, unpredictable | Creative writing, poetry |
Higher temperatures make responses less consistent and increase the risk of invented details, which matters most for customer-facing agents.
Recommendation: Start with 0.3 for most production agents. Adjust based on task:
- Increase (0.7-1.0) for creative tasks
- Decrease (0.0-0.1) for deterministic outputs
Technique 6: Explicit Constraints
Don't assume the model knows your constraints. State them explicitly.
Before (Implicit)
Summarize this article.Result: 800-word summary (way too long)
After (Explicit)
Summarize this article in exactly 3 sentences. Each sentence must be under 25 words.Result: 3 sentences, 72 words total ✅
Constraint Types to Specify
1. Length
- "In exactly 3 bullet points"
- "Under 100 words"
- "One paragraph"
2. Format
- "Return as numbered list"
- "Use markdown headings"
- "JSON only, no explanation"
3. Tone
- "Professional business tone"
- "Casual, friendly language"
- "Technical, for engineers"
4. Content restrictions
- "Do not mention competitors"
- "Avoid jargon"
- "Include at least one statistic"
Explicit constraints have the biggest effect on length requirements, followed by format and tone.
Technique 7: Iterative Refinement Pattern
For complex tasks, break into steps with validation.
Single-Shot (Less Reliable)
User query → [Agent generates final answer] → Return to userIterative Refinement (More Reliable)
Step 1: [Agent drafts answer]
Step 2: [Agent reviews draft for errors]
Step 3: [Agent revises if needed]
Step 4: Return to userThe review step catches many errors the first draft would have passed straight to the user.
Implementation
def iterative_answer(query):
# Step 1: Draft
draft_prompt = f"Draft an answer to: {query}"
draft = call_llm(draft_prompt)
# Step 2: Review
review_prompt = f"""
Review this draft answer for accuracy and completeness:
Query: {query}
Draft: {draft}
Issues (if any):
"""
review = call_llm(review_prompt)
# Step 3: Revise if issues found
if "Issue:" in review or "Error:" in review:
revise_prompt = f"""
Original query: {query}
Draft: {draft}
Issues found: {review}
Provide revised answer:
"""
final = call_llm(revise_prompt)
else:
final = draft
return finalCost: 2-3× LLM calls
Benefit: Fewer errors reach users
ROI: Worth it for high-stakes use cases (medical, legal, financial)
Prompt Template Library
Classification Template
CLASSIFICATION_TEMPLATE = """
Classify the input into one of these categories: {categories}
Examples:
{few_shot_examples}
Input: {input_text}
Category (one word only):
"""Data Extraction Template
EXTRACTION_TEMPLATE = """
Extract the following fields from the text. Return ONLY valid JSON.
Required schema:
{json_schema}
Text:
{input_text}
JSON:
"""Reasoning Template
REASONING_TEMPLATE = """
Answer this question by thinking step by step.
Question: {question}
Let's solve this step by step:
1.
"""What Doesn't Work
| Technique | Claimed Benefit | Actual Result | Status |
|---|---|---|---|
| "Be creative!" | Better outputs | No measurable difference | ❌ Myth |
| "You are an expert..." | Higher quality | Small, inconsistent effect | ❌ Overhyped |
| "Say please" | Politeness helps | No difference | ❌ Myth |
| ALL CAPS | Emphasis | No difference | ❌ Doesn't work |
| Emoji in prompts 🎯 | Engagement | No difference | ❌ Gimmick |
Stick to techniques you can measure.
Frequently Asked Questions
How much does prompt engineering actually matter vs model selection?
On many tasks, a cheaper model with well-optimised prompts gets close to a more expensive model with basic prompts.
But: The larger model can cost many times more per token. When the accuracy gap is small, prompt engineering the cheaper model is the better ROI.
Recommendation: Optimize prompts first. Upgrade model only if prompt optimization plateaus below requirements.
Should I version-control prompts?
Yes. Treat prompts like code:
# prompts/v1/customer_support.py
SYSTEM_PROMPT_V1 = """
You are a customer support agent...
"""
# prompts/v2/customer_support.py
SYSTEM_PROMPT_V2 = """
You are a helpful support agent. Answer using the knowledge base provided.
Use examples from context where possible.
"""Run A/B tests:
variant = random.choice(['v1', 'v2'])
prompt = SYSTEM_PROMPT_V1 if variant == 'v1' else SYSTEM_PROMPT_V2
# Track which variant performs better
log_metric('prompt_version', variant, accuracy)How do I measure prompt quality?
Key metrics:
- Task success rate: Did agent complete the task correctly?
- Format compliance: Output matches expected format (JSON, specific length, etc.)
- Hallucination rate: Factually incorrect or invented information
- User satisfaction: If customer-facing, track ratings
Evaluation pipeline:
def evaluate_prompt(prompt_template, test_cases):
results = []
for case in test_cases:
response = call_llm(prompt_template.format(**case['input']))
results.append({
'correct': response == case['expected_output'],
'valid_format': validate_format(response),
'has_hallucination': detect_hallucination(response, case['context'])
})
return {
'accuracy': sum(r['correct'] for r in results) / len(results),
'format_compliance': sum(r['valid_format'] for r in results) / len(results),
'hallucination_rate': sum(r['has_hallucination'] for r in results) / len(results)
}---
Bottom line: Prompt engineering isn't magic, but these 7 techniques consistently earn their place. Start with few-shot examples and structured output (biggest wins). Add chain-of-thought selectively. Test everything.
Next: Read our Agent Testing Strategies guide to build evaluation pipelines for prompt optimization.
More from the blog
How to Set Up Claude Code on a VPS: A Complete Guide
Claude Code VPS setup, step by step: provisioning, authentication, tmux vs systemd, security, and an honest look at when a VPS beats running locally.
Claude Code Agent Teams: How to Run Them on a Schedule
Claude Code Agent Teams runs up to 10 parallel Claude instances against one task list. What it is, how it works, and how to schedule runs.
Stop doing the work around the work
OpenHelm connects to your tools, reads the context, and does the steps, so you sign off on the result instead of producing it. See how it covers an entire role’s weekly workload, check the pricing, or run it yourself with the free local app.