Prompt caching explained: cut your LLM costs
If every request starts with the same 8,000-token system prompt or the same document, you are paying to process those tokens again and again. Prompt caching stops that: the API remembers the prefix, and later requests that begin the same way read it from cache at a discount.
How it works#
Caching is a prefix match. The request is rendered in a fixed order (tools, then system, then messages). Everything up to your cache breakpoint is the prefix. If a later request starts with exactly the same bytes, it is a hit.
Turn it on#
The simplest way is top-level automatic caching:
import Anthropic from "@anthropic-ai/sdk";
const client = new Anthropic();
const response = await client.messages.create({
model: "claude-opus-5-5",
max_tokens: 1024,
cache_control: { type: "ephemeral" }, // caches the last cacheable block
system: "You are an expert on the following handbook... (long text)",
messages: [{ role: "user", content: "What is the refund policy?" }],
});Or place the breakpoint yourself on a specific block, with an optional longer lifetime:
system: [
{
type: "text",
text: LONG_HANDBOOK,
cache_control: { type: "ephemeral", ttl: "1h" }, // default is 5 minutes
},
],Verify it, do not assume#
console.log(response.usage.cache_creation_input_tokens); // written to cache (costs a bit more)
console.log(response.usage.cache_read_input_tokens); // served from cache (cheap)
console.log(response.usage.input_tokens); // normal-price tokensOn the first call you should see a creation number. On the second, with an identical prefix, a read number. If reads stay at zero across repeated calls, something is invalidating the prefix.
What silently breaks the cache#
| Culprit | Why it breaks | Fix |
|---|---|---|
new Date() or a timestamp in the system prompt | Prefix changes every call | Move it after the breakpoint |
| A request ID or user name early in the prompt | Same | Put per-request data last |
| JSON with unsorted keys | Bytes differ | Serialize deterministically |
| A changing tool list or order | Tools are the first part of the prefix | Keep the tool set stable |
| Prefix shorter than the model's minimum | Too short to cache | Cache only large prefixes |
| TTL expired | Default is 5 minutes | Use the 1-hour TTL for slow traffic |
When it pays off#
- Chat apps with a long system prompt.
- "Ask questions about this document" tools.
- Agents that resend a growing conversation and a large tool list.
- Evals that run hundreds of cases against the same instructions.
It does not help when every request has a unique long prefix.
Design for caching#
Put stable content first (instructions, tool definitions, reference documents) and volatile content last (the user's question, timestamps). That is also good context engineering. New to the API? Start with your first call.
Frequently asked questions
What is prompt caching?
A feature where the API stores a processed prefix of your prompt, such as a long system prompt or document, so later requests that start with the same prefix reuse it at a lower price and faster.
How much does prompt caching save?
Cached reads are billed at roughly a tenth of normal input price, while writing to the cache costs somewhat more than normal input. It pays off when the same prefix is reused several times.
Why is my cache hit rate zero?
Caching is an exact prefix match. A timestamp, request ID, reordered JSON keys or a changed tool list anywhere in the prefix invalidates everything after it. Also check the prefix meets the model's minimum cacheable length.
How do I know prompt caching is working?
Look at response.usage. cache_creation_input_tokens shows tokens written to the cache and cache_read_input_tokens shows tokens served from it.
Prefer plain text? Read this page as Markdown.