LLM provider differences
The same model is not the same product everywhere, and these notes are about the differences that change a bill or a cache. Claude on Anthropic’s own API, on AWS Bedrock and on Google Vertex differs in authentication, model IDs, price per endpoint class and where the prompt cache lives. Bedrock’s two ways to authenticate decide whether anything in the request path can add a cache breakpoint.
A per-token price depends on the platform, endpoint class and billing mode, not just the model name. And providers differ in what a proxy can do about prompt caching at all: create hits, raise the odds, or only watch. Each note says what to check before you compare prices or move traffic.
Start here
LLM token cost is a function, not a number
What a per-token price depends on; the other notes are cases of it.
All 4 posts
Prompt caching providers: inject, steer or observe
Which LLM providers let a proxy create prompt cache hits, which only let it raise the odds, and which it can only watch: three verdicts, by provider.
Claude on Bedrock vs Vertex: not the same product
Claude on Anthropic, Bedrock and Vertex differs in auth, model IDs, price per endpoint class and which prompt cache it hits. Merge them and figures drift.
Bedrock API keys are bearer tokens. Caching cares.
SigV4 signs the request body, so no proxy can add a cache breakpoint without the secret key. A Bedrock API key signs no body, and that decides caching.
LLM token cost is a function, not a number
A per-token price is a function of platform, model ID, endpoint class, billing mode, context tier and time. Every shortcut gives a confident wrong number.
To price one call on each platform side by side, use the token cost calculator.