Claude API

Prompt caching explained: cut your LLM costs

If every request starts with the same 8,000-token system prompt or the same document, you are paying to process those tokens again and again. Prompt caching stops that: the API remembers the prefix, and later requests that begin the same way read it from cache at a discount.

How it works#

Caching is a prefix match. The request is rendered in a fixed order (tools, then system, then messages). Everything up to your cache breakpoint is the prefix. If a later request starts with exactly the same bytes, it is a hit.

Request 1prefix +question ACache writeprefix storedRequest 2same prefix +question BCache readprefix at afraction of theprice
Only the part after the breakpoint is billed at full input price on a hit.

Turn it on#

The simplest way is top-level automatic caching:

ts
import Anthropic from "@anthropic-ai/sdk";

const client = new Anthropic();

const response = await client.messages.create({
  model: "claude-opus-5-5",
  max_tokens: 1024,
  cache_control: { type: "ephemeral" }, // caches the last cacheable block
  system: "You are an expert on the following handbook... (long text)",
  messages: [{ role: "user", content: "What is the refund policy?" }],
});

Or place the breakpoint yourself on a specific block, with an optional longer lifetime:

ts
system: [
  {
    type: "text",
    text: LONG_HANDBOOK,
    cache_control: { type: "ephemeral", ttl: "1h" }, // default is 5 minutes
  },
],

Verify it, do not assume#

ts
console.log(response.usage.cache_creation_input_tokens); // written to cache (costs a bit more)
console.log(response.usage.cache_read_input_tokens);     // served from cache (cheap)
console.log(response.usage.input_tokens);                // normal-price tokens

On the first call you should see a creation number. On the second, with an identical prefix, a read number. If reads stay at zero across repeated calls, something is invalidating the prefix.

Normal input100 (baseline)Cache write (first call)about 125Cache read (later calls)about 10
Approximate multipliers. Check current pricing; the break-even is a couple of reuses.

What silently breaks the cache#

CulpritWhy it breaksFix
new Date() or a timestamp in the system promptPrefix changes every callMove it after the breakpoint
A request ID or user name early in the promptSamePut per-request data last
JSON with unsorted keysBytes differSerialize deterministically
A changing tool list or orderTools are the first part of the prefixKeep the tool set stable
Prefix shorter than the model's minimumToo short to cacheCache only large prefixes
TTL expiredDefault is 5 minutesUse the 1-hour TTL for slow traffic

When it pays off#

  • Chat apps with a long system prompt.
  • "Ask questions about this document" tools.
  • Agents that resend a growing conversation and a large tool list.
  • Evals that run hundreds of cases against the same instructions.

It does not help when every request has a unique long prefix.

Design for caching#

Put stable content first (instructions, tool definitions, reference documents) and volatile content last (the user's question, timestamps). That is also good context engineering. New to the API? Start with your first call.

Frequently asked questions

What is prompt caching?

A feature where the API stores a processed prefix of your prompt, such as a long system prompt or document, so later requests that start with the same prefix reuse it at a lower price and faster.

How much does prompt caching save?

Cached reads are billed at roughly a tenth of normal input price, while writing to the cache costs somewhat more than normal input. It pays off when the same prefix is reused several times.

Why is my cache hit rate zero?

Caching is an exact prefix match. A timestamp, request ID, reordered JSON keys or a changed tool list anywhere in the prefix invalidates everything after it. Also check the prefix meets the model's minimum cacheable length.

How do I know prompt caching is working?

Look at response.usage. cache_creation_input_tokens shows tokens written to the cache and cache_read_input_tokens shows tokens served from it.