Google Vertex through the Vigil proxy
Google Cloud’s endpoint for Claude, under a Vertex project and location, billed through Google Cloud. Point your SDK's base URL at Vigil, add one header, and every Google Vertex call is logged with its cost, tokens, latency, errors and agent, then forwarded to the Google endpoint your path's locations/ segment names unchanged, unless you switch prompt-cache optimisation on.
What Vigil records per Google Vertex call
One row per call: the model id as Google Vertex returned it, input and output tokens, cache reads and writes where the response reports them, latency, the HTTP status and the provider's error code when it fails, and the agent named in the X-Vigil-Agent header. Likely prompt injections and personal data in prompts are flagged on the Errors page. A cost spike is flagged when a call costs over three times the agent's seven-day median.
Cost is computed from the proxy's rate registry, which prices 11 models on this platform today. A model the registry has no row for is logged with its cost left empty, never estimated.
- claude-fable-5
- claude-fable-5-1
- claude-haiku-4-5@20251001
- claude-opus-4-5@20251101
- claude-opus-4-6
- claude-opus-4-7
- claude-opus-4-8
- claude-opus-5
- claude-sonnet-4-5@20250929
- claude-sonnet-4-6
- claude-sonnet-5
Price per token, by model: Claude Opus 5, Claude Sonnet 5, Claude Haiku 4.5, Claude Fable 5.1, Claude Opus 4.8, Claude Sonnet 4.6.
Setup
Every route has the shape https://api.vigil.tools/{user_id}/{provider}/{path}. For Google Vertex the provider segment is vertex, so the base URL is https://api.vigil.tools/{user_id}/vertex and a request path such as /v1/projects/{project}/locations/{location}/publishers/anthropic/models/{model}:rawPredict is forwarded unchanged.
Two headers: X-Vigil-Key, your Vigil key from the dashboard, and X-Vigil-Agent, the name the dashboard groups this traffic under. Your Google Vertex key stays where it was, in GOOGLE_APPLICATION_CREDENTIALS.
import { AnthropicVertex } from "@anthropic-ai/vertex-sdk";
const client = new AnthropicVertex({
projectId: process.env.GOOGLE_CLOUD_PROJECT, // unchanged
region: "us-east5", // unchanged
baseURL: "https://api.vigil.tools/{user_id}/vertex/v1",
defaultHeaders: {
// MUST be X-Vigil-Key: Authorization carries the OAuth token the SDK mints.
"X-Vigil-Key": "vk_your_vigil_key",
"X-Vigil-Agent": "my-agent",
},
});
const message = await client.messages.create({
model: "claude-haiku-4-5@20251001",
max_tokens: 1024,
messages: [{ role: "user", content: "Hello!" }],
});The header name and value are certain; this SDK’s plumbing for setting them is not — it has not been smoke-tested against a live call. If it does not work, the curl format is the ground truth for exactly what Vigil expects. The location in the URL picks the Google endpoint AND the price: `global` is list price, while `us`/`eu` (multi-region) and any single region such as `us-east5` carry +10%. Sonnet 5 and Opus 5 are not served on regional endpoints at all — use global, us or eu for those, or Vigil will refuse the call rather than pick an endpoint for you.
from anthropic import AnthropicVertex
client = AnthropicVertex(
project_id=os.environ["GOOGLE_CLOUD_PROJECT"], # unchanged
region="us-east5", # unchanged
base_url="https://api.vigil.tools/{user_id}/vertex/v1",
default_headers={
"X-Vigil-Key": "vk_your_vigil_key",
"X-Vigil-Agent": "my-agent",
},
)
message = client.messages.create(
model="claude-haiku-4-5@20251001",
max_tokens=1024,
messages=[{"role": "user", "content": "Hello!"}],
)The header name and value are certain; this SDK’s plumbing for setting them is not — it has not been smoke-tested against a live call. If it does not work, the curl format is the ground truth for exactly what Vigil expects. The location in the URL picks the Google endpoint AND the price: `global` is list price, while `us`/`eu` (multi-region) and any single region such as `us-east5` carry +10%. Sonnet 5 and Opus 5 are not served on regional endpoints at all — use global, us or eu for those, or Vigil will refuse the call rather than pick an endpoint for you.
curl https://api.vigil.tools/{user_id}/vertex/v1/projects/$GOOGLE_CLOUD_PROJECT/locations/us-east5/publishers/anthropic/models/claude-haiku-4-5@20251001:rawPredict \
-H "content-type: application/json" \
-H "Authorization: Bearer $(gcloud auth print-access-token)" \
-H "X-Vigil-Key: vk_your_vigil_key" \
-H "X-Vigil-Agent: my-agent" \
-d '{
"anthropic_version": "vertex-2023-10-16",
"max_tokens": 1024,
"messages": [{ "role": "user", "content": "Hello!" }]
}'The location in the URL picks the Google endpoint AND the price: `global` is list price, while `us`/`eu` (multi-region) and any single region such as `us-east5` carry +10%. Sonnet 5 and Opus 5 are not served on regional endpoints at all — use global, us or eu for those, or Vigil will refuse the call rather than pick an endpoint for you.
Replace {user_id} and vk_your_vigil_key with the values on your Connect page. The snippets above are generated by the same code as that page.
Prompt caching on Google Vertex
Vigil adds cache breakpoints to your requests — it reads the prompt structure and marks where the reusable prefix ends, so the provider caches exactly that much. This is the mechanism the Optimise page measures and reports savings for.
With optimisation On: Vigil adds cache markers to the stable start of the prompt. In Shadow, Vigil measures what caching would have saved and changes nothing. Off records the call and nothing more.
Questions
- Does Vigil see or store my Google Vertex API key?
- No, not the key itself. Your Google Vertex key passes through the proxy with the request and is never written to disk. A call record can keep a short one-way digest of the credential, so that one key’s cache is kept apart from another’s; it cannot be turned back into the key. Vigil identifies you by the X-Vigil-Key header and your user id in the URL.
- Does routing Google Vertex through Vigil add latency?
- Some. The proxy runs on Cloudflare’s edge, so the added hop is short, but reading the request body to place cache markers takes time, and there is a bounded lookup for your account state. The dashboard shows total latency per call. For long calls, stream: Cloudflare’s edge closes non-streaming connections after roughly 125 seconds.
- Which Google Vertex models does Vigil price?
- 11 models with a published rate in the registry: claude-fable-5, claude-fable-5-1, claude-haiku-4-5@20251001, claude-opus-4-5@20251101, claude-opus-4-6, claude-opus-4-7, claude-opus-4-8, claude-opus-5, claude-sonnet-4-5@20250929, claude-sonnet-4-6, claude-sonnet-5. A call to any other Google Vertex model is logged and monitored, with its cost left empty rather than guessed.
Other providers: Anthropic · AWS Bedrock · OpenAI · Google Gemini · Mistral · xAI Grok · DeepSeek · Together AI · Fireworks AI · Groq · Cerebras · Baseten · Moonshot · Z.ai · Cloudflare Workers AI · OpenRouter