Library · Power-user techniques

Prompt Caching and the Batch API: economies of scale

Engineer60 minUpdated: October 2026
47 of 105 in the library

Time: ~30 min theory + 30 min practice


The gist

Two ways to save money at scale: Prompt Caching is like a gym membership (you pay once to get in, then go many times), and the Batch API is like a wholesale order from a factory (cheaper per unit, but you wait up to 24 hours for delivery). Each one on its own gives you a 50-90% discount. Combined, cached input tokens can cost as little as 5% of the base price (and cache reads on Opus 5.5 and Fable 5.1 are even cheaper). Current prices and versions: What's current.


Key concepts

  • Prompt Caching: caching the repeating parts of your prompt on Anthropic's servers
  • cache_control: the {"type": "ephemeral"} marker that says what to cache
  • TTL: how long the cache lives: 5 minutes (standard, "5m") or 1 hour ("1h", more expensive to write)
  • Cache pricing: a 5m write costs 1.25x the base price, a 1h write costs 2x, and a read costs 0.1x (90% savings; on Opus 5.5 a read is 0.05x, and on Fable 5.1 it's cheaper still)
  • Batch API: send up to 100,000 requests (or 256 MB) at once with a 50% discount
  • Batch statuses: in_progress → ended (most batches finish in under an hour, 24 hours at most)

Theory

Part 1: Prompt Caching

How caching works

Every time you send a request to Claude, you pay for all the tokens: the system prompt, the context, the examples, the user's message. If your system prompt is 3,000 tokens and you make 1,000 requests a day, that's 3 million tokens just for the repeating context.

Prompt Caching saves the prompt on Anthropic's servers. There are two TTLs (how long the cache lives):

TTL Write cost Read cost When to use it
5 minutes ("5m", default) 1.25x the base price (+25%) 0.1x the base price (−90%) Frequent requests, chatbots, real-time APIs
1 hour ("1h") 2x the base price (+100%) 0.1x the base price (−90%) Batch processing, long tasks with extended thinking (>5 min), infrequent requests
Type this into the chat
First request: you pay to write to the cache (1.25x for 5m or 2x for 1h)
Requests 2-N (within the TTL): you pay about 10% to read from the cache (5% on Opus 5.5, even less on Fable 5.1) + the full price for uncached tokens

🎨 Picture this: The 5-minute cache is a morning pass at a coffee shop: your second coffee costs 10% of the price, but only while you're still in the shop. The 1-hour cache is an all-day pass: more expensive up front, but it lasts much longer.

How a cached request is structured

Option 1: Automatic caching (top-level cache_control)

The simplest way: add cache_control at the request level, and the API figures out what to cache on its own:

python
import anthropic

client = anthropic.Anthropic()

response = client.messages.create(
    model="claude-sonnet-5-5",
    max_tokens=2048,
    cache_control={"type": "ephemeral"},  # ← automatic caching
    system="You are a specialized assistant for analyzing contracts...",
    messages=[{"role": "user", "content": "Analyze this contract: [text]"}]
)

Option 2: Explicit breakpoints (precise control)

For precise control, put cache_control on specific content blocks:

python
# System prompt: 2000+ tokens (long, repeated)
SYSTEM_PROMPT = """
You are a specialized assistant for analyzing commercial lease agreements.
Respond in English only.

ANALYSIS RULES:
1. Always check the lease term, start date and end date
2. Highlight the early termination terms
3. Note any penalties and late fees
4. Check whether there's a rent escalation clause
5. Note each party's responsibility for repairs
...another 1500 tokens of rules and examples...
"""

response = client.messages.create(
    model="claude-sonnet-5-5",
    max_tokens=2048,
    system=[
        {
            "type": "text",
            "text": SYSTEM_PROMPT,
            "cache_control": {"type": "ephemeral"}  # ← marker on a specific block
        }
    ],
    messages=[
        {
            "role": "user",
            "content": "Analyze this contract: [contract text]"
        }
    ]
)

# Check that the cache is working
usage = response.usage
print(f"Input tokens: {usage.input_tokens}")
print(f"Tokens written to cache: {usage.cache_creation_input_tokens}")
print(f"Tokens read from cache: {usage.cache_read_input_tokens}")

Option 3: 1-hour TTL (for batch jobs and long tasks with extended thinking)

python
response = client.messages.create(
    model="claude-sonnet-5-5",
    max_tokens=2048,
    cache_control={
        "type": "ephemeral",
        "ttl": "1h"  # ← 1 hour instead of 5 minutes (writing costs more, reading costs the same)
    },
    system="A long system prompt...",
    messages=[{"role": "user", "content": "Request..."}]
)

What's worth caching

Cache it Don't cache it
System prompt The user's request
Few-shot examples (5-10 of them) Users' personal data
Long instructions Dynamic data (time, IDs)
Knowledge base (RAG context) Short parts that change
Legal/technical rules Variable parts of a template

Minimum size for caching (depends on the model; as of October 2026):

Models Minimum tokens
Fable 5.1, Opus 5.5, Sonnet 5.5 512 tokens
Sonnet 5, Sonnet 4.6, Sonnet 4.5, Opus 4.8 1,024 tokens
Opus 4.7 2,048 tokens
Haiku 4.5, Opus 4.6, Opus 4.5 4,096 tokens

Below the minimum, the cache silently isn't created (no error). Check: if cache_creation_input_tokens and cache_read_input_tokens are both 0, caching didn't work.

The caching price model (as of October 2026)

Sonnet 5.5 (a typical choice, $2/MTok input):

Token type 5 min TTL 1 hour TTL
Regular input tokens $2 / 1M $2 / 1M
Cache write $2.50 / 1M (+25%) $4 / 1M (+100%)
Cache read $0.20 / 1M (−90%) $0.20 / 1M (−90%)
Output tokens $10 / 1M $10 / 1M

All current models (as of October 2026):

Model Base input 5m write 1h write Cache read
Fable 5.1 $10/MTok $12.50/MTok $20/MTok see the pricing page
Opus 5.5 $4/MTok $5/MTok $8/MTok $0.20/MTok
Sonnet 5.5 $2/MTok $2.50/MTok $4/MTok $0.20/MTok
Haiku 4.5 $1/MTok $1.25/MTok $2/MTok $0.10/MTok

The formula: 5m write = 1.25x base, 1h write = 2x base, read = 0.1x base (0.05x on Opus 5.5, and less than that on Fable 5.1). Prices change from version to version, but the principle stays the same: What's current.

On the first request you pay a little more to create the cache. Starting with the second request within the TTL, you save about 90% on cached tokens.

A savings example

Scenario (Sonnet 5.5 prices as of October 2026): 1,000 requests a day, a 3,000-token system prompt, a 500-token response. Requests come often enough that the cache gets refreshed every 5 minutes.

Type this into the chat
Without caching:
  Input: 1000 × 3000 = 3,000,000 tokens × $2/1M = $6.00/day
  Output: 1000 × 500 = 500,000 tokens × $10/1M = $5.00/day
  Total: $11.00/day

With caching (5-minute TTL, one session):
  Cache write (once): 3000 × $2.50/1M = $0.0075
  Cache reads (999 times): 999 × 3000 × $0.20/1M = $0.599
  Output (1000 times): $5.00/day
  Total: $5.61/day

Savings: ~49%

The example shows the calculation method. Plug in your own numbers from the What's current page: the savings depend on how much of your total cost comes from repeated input.

Multiple cache points (breakpoints)

🎨 Picture this: Bookmarks in a thick book: one at the introduction, one in the middle, one at the end. You don't read from the beginning every time; you jump straight to the spot you need. Breakpoints are bookmarks for the API: it reads the cache from the right point instead of recomputing everything from scratch.

You can cache several blocks in one request. The maximum is 4 breakpoints (explicit cache_control):

python
response = client.messages.create(
    model="claude-sonnet-5-5",
    max_tokens=1024,
    system=[
        {
            "type": "text",
            "text": BASE_RULES,           # Base rules (always)
            "cache_control": {"type": "ephemeral"}  # breakpoint 1
        },
        {
            "type": "text",
            "text": DOMAIN_KNOWLEDGE,     # Domain knowledge (changes sometimes)
            "cache_control": {"type": "ephemeral"}  # breakpoint 2
        }
    ],
    messages=[{
        "role": "user",
        "content": user_question
    }]
)

The order the cache is checked in: the API checks the cache in this order: tools → system → messages. Each breakpoint caches everything up to and including it (cumulatively).

What can and can't be cached

Can be cached Can't be cached
Tool definitions (the tools array) Empty text blocks
System messages Thinking blocks with explicit cache_control
Text messages (user and assistant) Sub-content (citations inside documents)
Images and documents in user messages
Tool use / tool result blocks

What invalidates the cache

Changing any of these "breaks" the cache, and the next request creates a new one:

  • Tool definitions
  • Thinking and effort parameters (on some models)
  • Switching tool_choice
  • Changing images in the prompt

Part 2: The Batch API

When you need the Batch API

🎨 Picture this: A delivery service: you can call a cab right now and pay three times as much, or hand over 1,000 packages for next-day delivery and get a 50% bulk discount. The Batch API is the bulk courier: not urgent, but cheap.

The Batch API is for tasks that don't need an answer right this second. You process a whole stack of requests, get the results a few hours later, and pay half as much.

Regular API Batch API
Answer in 1-5 seconds Most < 1 hour, max 24 hours
Full price 50% off everything
One request Up to 100,000 requests (or 256 MB)
Synchronous Asynchronous

Ideal use cases:

  • Classifying 5,000 customer reviews
  • Generating descriptions for 2,000 products
  • Analyzing 1,000 résumés
  • Translating 3,000 articles
  • SEO optimization for 500 pages
  • Large-scale evaluations (thousands of test cases)

What you can send in a batch: any Messages API request: vision, tool use, system messages, multi-turn, extended thinking, any beta features. The exception: fast mode isn't available in batches. Each request is processed independently, so you can mix different types in one batch.

Tip: for batches with a shared system prompt, use the 1-hour cache ("ttl": "1h"): a batch usually takes longer than 5 minutes, and a 5-minute cache will expire.

Sending a batch

python
import anthropic

client = anthropic.Anthropic()

# Preparing the requests. Haiku 4.5 may be retired from the API no earlier than 10/15/2026:
# before you run this, check the model deprecations page and use the current cheap model
requests = []
products = load_products_from_db()  # your 1000 products

for i, product in enumerate(products):
    requests.append({
        "custom_id": f"product-{product['id']}",  # your ID for matching
        "params": {
            "model": "claude-haiku-4-5-20251001",  # Haiku for batches: cheaper
            "max_tokens": 500,
            "messages": [{
                "role": "user",
                "content": f"""Write an SEO description for this product:
Name: {product['name']}
Category: {product['category']}
Specs: {product['specs']}

The description should be 100-150 words and include keywords."""
            }]
        }
    })

# Send the batch
batch = client.messages.batches.create(requests=requests)

print(f"Batch created: {batch.id}")
print(f"Status: {batch.processing_status}")  # in_progress
print(f"Requests in batch: {batch.request_counts.processing}")

Checking the status and getting the results

python
import time

batch_id = batch.id

# Wait for it to finish (polling)
while True:
    batch_status = client.messages.batches.retrieve(batch_id)
    
    if batch_status.processing_status == "ended":
        print("Batch finished!")
        print(f"Succeeded: {batch_status.request_counts.succeeded}")
        print(f"Errors: {batch_status.request_counts.errored}")
        break
    
    print(f"Processing: {batch_status.request_counts.processing} requests...")
    time.sleep(60)  # check once a minute

# Get the results
results = {}
for result in client.messages.batches.results(batch_id):
    if result.result.type == "succeeded":
        results[result.custom_id] = "".join(b.text for b in result.result.message.content if b.type == "text")
    else:
        results[result.custom_id] = None
        print(f"Error for {result.custom_id}: {result.result.error}")

# Save to the database
save_descriptions_to_db(results)

Batch statuses

Code
in_progress  → requests are being processed
ending       → wrapping up (some are still running)
ended        → all done, results available

For each request:
  succeeded   → OK, there's a result
  errored     → error (rate limit, invalid request)
  expired     → the request wasn't processed within 24 hours
  canceled    → the batch was canceled manually

🎨 Picture this: Polling a batch's status is like tracking a package on the carrier's website: you check once a minute to see "in transit" or "delivered," instead of waiting for the driver to call.

Important: batch results are available for 29 days after creation. After that you can still see the batch itself, but you can't download the results.

Canceling a batch

python
# Changed your mind? Cancel while there's still time
client.messages.batches.cancel(batch_id)

Canceled and unprocessed requests aren't billed.

Batch API pricing

Everything is 50% of the standard prices, both input and output:

Model (as of October 2026) Batch input Batch output
Fable 5.1 $5/MTok $25/MTok
Opus 5.5 $2/MTok $10/MTok
Sonnet 5.5 $1/MTok $5/MTok
Haiku 4.5 $0.50/MTok $2.50/MTok

Batch API + the output-300k-2026-03-24 beta header: up to 300,000 output tokens per request for Opus 5.5, Sonnet 5.5 and several earlier models (the usual limit for a synchronous request is 128k on Fable 5.1, Opus 5.5 and Sonnet 5.5, and 64k on Haiku 4.5). A single answer like that can take more than an hour to generate, so plan for the full 24-hour window.


The combo: Batch + Caching = maximum savings

python
# Shared system prompt, cached with a 1-hour TTL (the batch takes > 5 min!)
ANALYSIS_SYSTEM = """[4500+ tokens of analysis rules: for Haiku 4.5 the cache minimum is 4096 tokens]"""

requests = []
for doc in documents:  # 5000 documents
    requests.append({
        "custom_id": f"doc-{doc['id']}",
        "params": {
            "model": "claude-haiku-4-5-20251001",
            "max_tokens": 300,
            "system": [
                {
                    "type": "text",
                    "text": ANALYSIS_SYSTEM,
                    "cache_control": {
                        "type": "ephemeral",
                        "ttl": "1h"  # ← 1 hour! The batch takes longer than 5 minutes
                    }
                }
            ],
            "messages": [{"role": "user", "content": doc['text']}]
        }
    })

batch = client.messages.batches.create(requests=requests)

Why "1h" and not "5m"? A batch is processed asynchronously. If it takes 20 minutes, a 5-minute cache expires after the first few requests, and the remaining 4,500 documents pay full price. The 1-hour cache costs more to write (2x), but less overall.

An illustration using a hypothetical $100 (actual savings depend on how much of your input repeats):

Method Discount Final price
Regular API 0% $100
Batch only -50% $50
Caching only -45% (average) $55
Batch + Caching -90% to -95% $5-10

🎨 Picture this: Two kinds of savings at a store. Prompt Caching is a loyalty card: your second coffee is cheaper. The Batch API is a wholesale order: buy 1,000 units and get the wholesale price. Use both and you get the biggest discount.


Practice

Assignment: Batch processing with caching

  1. Prepare a list of 10 short texts (reviews, descriptions, anything):

    python
    texts = [
        "Great service, I recommend it to everyone!",
        "Delivery was 3 days late, not great.",
        # ... 8 more texts
    ]
  2. Create a batch for sentiment classification (positive/negative/neutral):

    python
    SENTIMENT_PROMPT = """
    Classify the sentiment of the text.
    Answer with a single word: positive, negative or neutral.
    Don't add any explanation.
    """  # ~50 tokens: below the minimum, so add more rules and examples
  3. Add cache_control to the system prompt (expand it with examples: Sonnet 5.5 needs at least 512 tokens, Haiku 4.5 at least 4,096)

  4. Send the batch and start a polling loop to check the status

  5. When it's done: print the custom_id + result for each text

  6. Check usage in the responses: do you see cache_read_input_tokens?

Goal: go through the full Batch API cycle and see the savings in real numbers.



Additional paid API features (as of October 2026)

Besides model tokens, the Claude API has some separately billed services:

Feature Price What it does
Web Search $10 / 1,000 searches + tokens Claude searches the internet while answering
Web Fetch Free (tokens only) Claude reads a URL you give it
Code Execution billed per container-hour, with a free monthly allowance (see the pricing page) Runs Python inside the response
Code Execution + Web Free When used together with Web Search/Fetch
Managed Agents billed per session-hour + tokens (see the pricing page) Anthropic-hosted agents (you pay only for running time)
US-only data residency a surcharge on all tokens (see the pricing page) A legal requirement to keep data in the US

Fast mode (research preview) for Opus 5.5: noticeably faster, but twice as expensive: $8/MTok input and $40/MTok output versus $4 and $20 in regular mode. It doesn't work with the Batch API. Use it when speed matters more than cost.

🎨 Picture this: Web Search is like a paid search engine for Claude. Code Execution is like a sandbox where Claude can actually run Python and see the result (instead of guessing). Managed Agents are a rented server for your agent, billed by the hour like a cab with the meter running.


Tools and resources


Key takeaways

Prompt Caching pays for itself on the second request. Two TTLs: 5 minutes (default, write +25%) and 1 hour (write +100%). A read is usually 0.1x (−90%), and even cheaper on Opus 5.5 and Fable 5.1. The cache minimum depends on the model: from 512 to 4,096 tokens. If the prompt is shorter, the cache silently isn't created. The Batch API: up to 100,000 requests at once, 50% off. Most batches finish in under an hour. For batches, use the 1-hour cache ("ttl": "1h"): a 5-minute cache will expire before the batch finishes. Combine both methods for large-scale jobs: savings of 90-95% off the base price.


Next lesson

→ Loop and Scheduled Tasks: when to use which

The mark stays in this browser only and is never sent anywhere. My progress