Tools

Prompt caching calculator

What caching the stable start of your prompt would save, per day, on a model with published cache rates. Priced from the same rate table Vigil's proxy bills with, under a model of cache behaviour that is written out below so you can disagree with it. An estimate, not a measurement.

Only models with a published cache-write rate are listed.

Estimated saving per day

$21.59

Without caching, per day
$42.00
With caching, per day (1 write, 1,999 reads)
$20.41
If every call missed the cache, per day
$48.00
One call, prefix as plain input
$0.0210
One call that writes the prefix
$0.0240
One call that reads the prefix
$0.0102

Calls arrive every 43 s per stream against a 5-minute TTL, so each stream writes once and reads thereafter.

Minimum cacheable prefix on this model: 1,024 tokens.

An estimate under the assumptions stated below this calculator, from the registry row retrieved 2026-09-11. Not a measured saving.

Price one call in full, with endpoint class and monthly volume: AI token calculator

The assumptions

  • A call is a cached prefix (system prompt, tools, fixed examples), a fresh part (the user turn and anything that changes) and an output. Without caching the prefix is billed as input on every call.
  • With caching, the first call of a stream writes the prefix at the write rate and each later call within the TTL reads it at the read rate. Anthropic refreshes the TTL on every read at no charge, so a stream whose calls arrive at least once per TTL writes once a day and reads thereafter.
  • Calls are assumed evenly spaced. If the spacing exceeds the TTL, every call is a write; the calculator shows that case and calls it what it is, a loss.
  • “Streams” is the number of distinct prefixes live at once: one per agent, or per distinct system prompt. Each writes once a day in the steady case.
  • Below the model's minimum cacheable prefix, nothing is cached. The minimums shown are the published ones for the Claude family; for other platforms the floor is not checked here.
  • Rates are the registry's current window for the model at the moment you load the page; dated changes resolve as the proxy would resolve a call made now.

Whether your prefix actually matches on every call is the other half of the question: a timestamp, a reordered tool array or a changed system prompt misses the cache however favourable the arithmetic. The posts below cover each.

Questions

Why can caching cost more than not caching?
A cache write is priced above the input rate (1.25 times on Anthropic’s 5-minute tier, 2 times on the 1-hour tier), and a read far below it. The saving comes from reads. If your calls arrive less often than the TTL, the entry expires between them, every call is a write, and you pay the premium with no reads to recoup it. The calculator shows that case as a negative saving.
What counts as the cached prefix?
The bytes at the start of the request that are identical on every call: the system prompt, tool definitions and any fixed examples. Anything that changes, a timestamp, the user turn, retrieved documents, has to come after it, or the prefix match fails. Below the model’s minimum cacheable length nothing is cached at all.
Is this what Vigil would save me?
No. It is an estimate from published rates under the assumptions printed beside the result. Vigil measures the real figure from your traffic: in Shadow it reports what caching would have saved on the calls you actually made, and switched on, it reports the saving the provider’s bill shows. No savings figure of Vigil’s is quoted here because none is backed.

Caching on Anthropic, Vertex, Mistral and Bedrock with an API key is where Vigil can place the markers for you; see Anthropic base URL and the other providers.