LLM cost mechanics
How LLM API calls are actually priced, cached and failed, worked through one mechanism at a time. Most of these posts are about prompt caching: why a cache write costs more than the read it was meant to save, why a timestamp or a reordered tool array misses the cache on every call, what the minimum cacheable prompt length is, and when the 1-hour cache tier is worth paying for.
The rest cover what breaks around the call: streamed responses that report zero tokens, a 125-second edge timeout on long calls, error bodies that arrive as HTML, and model IDs that do not match a rate table. There is also a guide to LLM monitoring as a whole. Each post ends with what to check in your own traffic.
Start here
Anthropic prompt caching: a write costs 12.5x a read
The two prices of prompt caching, which most of the posts here build on.
All 12 posts
Multi-provider LLM observability: where to start
Multi-provider LLM observability is one event per model call, across every provider: cost, latency, errors, injection and PII. What to instrument first.
LLM Monitoring: How to Track AI Performance and Cost
LLM monitoring tracks four layers in production: performance, quality, cost and safety. What to measure in each, and where traces, evals and a proxy fit.
LLM model aliases, and the rates that don't resolve
The model ID on a response is often not the one you sent. A pricing engine must resolve the right aliases and refuse the tempting ones, or fail closed.
OpenAI streaming usage is zero unless you ask for it
Chat Completions sends no usage on a streamed call unless you set stream_options include_usage, so each streamed call meters as zero tokens and $0.00.
Prompt caching not working: check for a timestamp
Prompt caching matches exact bytes, so a clock in your system prompt misses the cache on every call. Why it is the most common cache failure, and the fix.
LLM API error handling when the error body is HTML
A timed-out upstream returned Cloudflare's HTML error page, not a JSON error. Two rules for LLM API errors: store the code, not the message; trust neither.
The 125-second LLM API timeout and the streaming fix
A CDN edge closes idle connections at about 125 seconds, and a long non-streaming LLM call looks idle. Your SDK gets a 524 page, not the provider's reply.
Prompt cache TTL: a 1-hour write costs 2x base input
Anthropic bills 1-hour cache writes at 2x base input, against 1.25x for the 5-minute TTL. Confusing them understates that line's cost by 37.5%, as we did.
Prompt caching with tools: order is part of the key
Tool definitions sit ahead of the system prompt in the cached prefix, so one changed tool, or the same tools reordered, invalidates everything after them.
Prompt caching minimum tokens: below it, no cache
Prompt caching has a floor: below a model's minimum, the API silently ignores your cache breakpoint. The minimums, the counting trap, and the rescue.
Your prompt cache hit rate hides avoidable writes
A healthy prompt cache hit rate can hide writes of bytes that already had a live entry, each paying 1.25x where a 0.1x read would do. How to find yours.
Anthropic prompt caching: a write costs 12.5x a read
On Anthropic's default tier prompt caching has two prices: a write at 1.25x the input rate, a read at 0.1x. A write is what you pay when no read was there.
To see cache writes, reads and cost per call for your own traffic, start with LLM proxy setup. Monitoring is free.