Notes from inside the request path
Caching mechanics, cost arithmetic, provider changes that move a bill, and what breaks in a proxy that cannot be allowed to break. Every post is something you can use whether or not you ever point an SDK at us.
RSS · posts are dated to when the finding was made
- Mechanics8 min read
LLM Monitoring: How to Track AI Performance and Cost
LLM monitoring tracks four layers in production: performance, quality, cost and safety. What to measure in each, and where traces, evals and a proxy fit.
- The honesty line5 min read
Proxy vs SDK observability: in-band and out-of-band
Proxy vs SDK observability: an SDK reports what your code believes happened, a proxy records what crossed the wire, and each knows what the other can't.
- Provider notes5 min read
Prompt caching providers: inject, steer or observe
Which LLM providers let a proxy create prompt cache hits, which only let it raise the odds, and which it can only watch: three verdicts, by provider.
- Mechanics5 min read
LLM model aliases, and the rates that don't resolve
The model ID on a response is often not the one you sent. A pricing engine must resolve the right aliases and refuse the tempting ones, or fail closed.
- Mechanics5 min read
OpenAI streaming usage is zero unless you ask for it
Chat Completions sends no usage on a streamed call unless you set stream_options include_usage, so each streamed call meters as zero tokens and $0.00.
- Mechanics7 min read
Prompt caching not working: check for a timestamp
Prompt caching matches exact bytes, so a clock in your system prompt misses the cache on every call. Why it is the most common cache failure, and the fix.
- Provider notes5 min read
Claude on Bedrock vs Vertex: not the same product
Claude on Anthropic, Bedrock and Vertex differs in auth, model IDs, price per endpoint class and which prompt cache it hits. Merge them and figures drift.
- Provider notes5 min read
Bedrock API keys are bearer tokens. Caching cares.
SigV4 signs the request body, so no proxy can add a cache breakpoint without the secret key. A Bedrock API key signs no body, and that decides caching.
- Provider notes5 min read
LLM token cost is a function, not a number
A per-token price is a function of platform, model ID, endpoint class, billing mode, context tier and time. Every shortcut gives a confident wrong number.
- Building5 min read
Cloudflare Browser Integrity Check blocked our site
Cloudflare's Browser Integrity Check sent real browsers a 403 and logged nothing. An IP skip rule hid it for a week, while every synthetic check passed.
- Building5 min read
Usage metering: a quota the customer could reset
Our usage metering counted rows customers could delete, so the product's own "clear history" button reset their quota. The fix: a count that only goes up.
- Building5 min read
GDPR account deletion that never worked for anyone
Our GDPR account deletion referenced three tables that didn't exist, so every run failed with 42P01, after cancelling the customer's Stripe subscription.
- Building5 min read
Fail open vs fail closed: a missing env var, no auth
Our API-key check defaulted to off when its env var was unset, empty or misspelled, so deleting one variable disabled auth. The fix, and a rule about 401s.
- Building5 min read
Stripe webhook not working: a redirect nobody saw
Our apex domain 308-redirected to www, and Stripe does not follow redirects, so every webhook since launch died one hop before our well-tested code.
- Findings5 min read
Anthropic web search pricing: the cost nobody meters
Anthropic bills web search at $10 per 1,000 requests, and advisor, fallback and compaction tokens outside top-level usage. Token-only metering misses both.
- Findings6 min read
Nine AI dashboard metrics that were quietly wrong
We checked every figure on our AI monitoring dashboard against the code behind it. Nine definitions failed, one a savings total that was a rolling window.
- The honesty line6 min read
LLM cost accuracy: why we show a dash, not $0.00
A zero is a claim and a missing figure is an absence. For LLM cost accuracy, our dashboard shows a dash wherever a plausible number would have been a lie.
- Mechanics5 min read
LLM API error handling when the error body is HTML
A timed-out upstream returned Cloudflare's HTML error page, not a JSON error. Two rules for LLM API errors: store the code, not the message; trust neither.
- Mechanics5 min read
The 125-second LLM API timeout and the streaming fix
A CDN edge closes idle connections at about 125 seconds, and a long non-streaming LLM call looks idle. Your SDK gets a 524 page, not the provider's reply.
- Mechanics5 min read
Prompt cache TTL: a 1-hour write costs 2x base input
Anthropic bills 1-hour cache writes at 2x base input, against 1.25x for the 5-minute TTL. Confusing them understates that line's cost by 37.5%, as we did.
- Mechanics5 min read
Prompt caching with tools: order is part of the key
Tool definitions sit ahead of the system prompt in the cached prefix, so one changed tool, or the same tools reordered, invalidates everything after them.
- Mechanics5 min read
Prompt caching minimum tokens: below it, no cache
Prompt caching has a floor: below a model's minimum, the API silently ignores your cache breakpoint. The minimums, the counting trap, and the rescue.
- Mechanics5 min read
Your prompt cache hit rate hides avoidable writes
A healthy prompt cache hit rate can hide writes of bytes that already had a live entry, each paying 1.25x where a 0.1x read would do. How to find yours.
- Mechanics5 min read
Anthropic prompt caching: a write costs 12.5x a read
On Anthropic's default tier prompt caching has two prices: a write at 1.25x the input rate, a read at 0.1x. A write is what you pay when no read was there.