Mechanics

LLM cost mechanics

How LLM API calls are actually priced, cached and failed, worked through one mechanism at a time. Most of these posts are about prompt caching: why a cache write costs more than the read it was meant to save, why a timestamp or a reordered tool array misses the cache on every call, what the minimum cacheable prompt length is, and when the 1-hour cache tier is worth paying for.

The rest cover what breaks around the call: streamed responses that report zero tokens, a 125-second edge timeout on long calls, error bodies that arrive as HTML, and model IDs that do not match a rate table. There is also a guide to LLM monitoring as a whole. Each post ends with what to check in your own traffic.

Start here

Anthropic prompt caching: a write costs 12.5x a read

The two prices of prompt caching, which most of the posts here build on.

All 12 posts

To see cache writes, reads and cost per call for your own traffic, start with LLM proxy setup. Monitoring is free.