# Prompt caching explained: cut your LLM costs

> Prompt caching reuses a repeated prompt prefix at a fraction of the cost. How it works, where to put cache_control, what silently breaks it and how to verify hits.

Source: https://devaiper.com/blog/prompt-caching-explained
Published: 2026-10-08
Topics: Prompt caching, Claude API, Cost

**Short answer:** Prompt caching lets the API reuse a repeated prefix of your prompt, billing cached reads at a fraction of normal input price. Put stable content first, mark it with cache_control, and check cache_read_input_tokens to confirm it works.

If every request starts with the same 8,000-token system prompt or the same document, you are paying to process those tokens again and again. **Prompt caching** stops that: the API remembers the prefix, and later requests that begin the same way read it from cache at a discount.

## How it works

Caching is a **prefix match**. The request is rendered in a fixed order (tools, then system, then messages). Everything up to your cache breakpoint is the prefix. If a later request starts with exactly the same bytes, it is a hit.

> **Diagram:** How prompt caching works. The first request writes the prefix to the cache at a slightly higher price. Later requests with the identical prefix read it from the cache at a much lower price, and only the new question is processed at full price.
> Request 1 (prefix + question A) → Cache write (prefix stored) → Request 2 (same prefix + question B) → Cache read (prefix at a fraction of the price)
> Only the part after the breakpoint is billed at full input price on a hit.

## Turn it on

The simplest way is top-level automatic caching:

```ts
import Anthropic from "@anthropic-ai/sdk";

const client = new Anthropic();

const response = await client.messages.create({
  model: "claude-opus-5-5",
  max_tokens: 1024,
  cache_control: { type: "ephemeral" }, // caches the last cacheable block
  system: "You are an expert on the following handbook... (long text)",
  messages: [{ role: "user", content: "What is the refund policy?" }],
});
```

Or place the breakpoint yourself on a specific block, with an optional longer lifetime:

```ts
system: [
  {
    type: "text",
    text: LONG_HANDBOOK,
    cache_control: { type: "ephemeral", ttl: "1h" }, // default is 5 minutes
  },
],
```

## Verify it, do not assume

```ts
console.log(response.usage.cache_creation_input_tokens); // written to cache (costs a bit more)
console.log(response.usage.cache_read_input_tokens);     // served from cache (cheap)
console.log(response.usage.input_tokens);                // normal-price tokens
```

On the first call you should see a creation number. On the second, with an identical prefix, a read number. If reads stay at zero across repeated calls, something is invalidating the prefix.

> **Diagram:** Relative cost of a long prompt prefix. A normal request costs 100. A cache write costs a little more, around 125. A cache read costs about 10.
> Normal input: 100 (baseline); Cache write (first call): about 125; Cache read (later calls): about 10
> Approximate multipliers. Check current pricing; the break-even is a couple of reuses.

## What silently breaks the cache

| Culprit | Why it breaks | Fix |
|---|---|---|
| `new Date()` or a timestamp in the system prompt | Prefix changes every call | Move it after the breakpoint |
| A request ID or user name early in the prompt | Same | Put per-request data last |
| JSON with unsorted keys | Bytes differ | Serialize deterministically |
| A changing tool list or order | Tools are the first part of the prefix | Keep the tool set stable |
| Prefix shorter than the model's minimum | Too short to cache | Cache only large prefixes |
| TTL expired | Default is 5 minutes | Use the 1-hour TTL for slow traffic |

## When it pays off

- Chat apps with a long system prompt.
- "Ask questions about this document" tools.
- Agents that resend a growing conversation and a large tool list.
- Evals that run hundreds of cases against the same instructions.

It does not help when every request has a unique long prefix.

## Design for caching

Put **stable** content first (instructions, tool definitions, reference documents) and **volatile** content last (the user's question, timestamps). That is also good [context engineering](https://devaiper.com/blog/what-is-context-engineering). New to the API? Start with [your first call](https://devaiper.com/blog/claude-api-first-call-typescript).

## FAQ

### What is prompt caching?

A feature where the API stores a processed prefix of your prompt, such as a long system prompt or document, so later requests that start with the same prefix reuse it at a lower price and faster.

### How much does prompt caching save?

Cached reads are billed at roughly a tenth of normal input price, while writing to the cache costs somewhat more than normal input. It pays off when the same prefix is reused several times.

### Why is my cache hit rate zero?

Caching is an exact prefix match. A timestamp, request ID, reordered JSON keys or a changed tool list anywhere in the prefix invalidates everything after it. Also check the prefix meets the model's minimum cacheable length.

### How do I know prompt caching is working?

Look at response.usage. cache_creation_input_tokens shows tokens written to the cache and cache_read_input_tokens shows tokens served from it.

