OpenAI through the Vigil proxy
OpenAI’s API at api.openai.com: Chat Completions and Responses, with automatic prompt caching on the provider’s side. Point your SDK's base URL at Vigil, add one header, and every OpenAI call is logged with its cost, tokens, latency, errors and agent, then forwarded to https://api.openai.com unchanged.
What Vigil records per OpenAI call
One row per call: the model id as OpenAI returned it, input and output tokens, cache reads and writes where the response reports them, latency, the HTTP status and the provider's error code when it fails, and the agent named in the X-Vigil-Agent header. Likely prompt injections and personal data in prompts are flagged on the Errors page. A cost spike is flagged when a call costs over three times the agent's seven-day median.
Cost is computed from the proxy's rate registry, which prices 13 models on this platform today. A model the registry has no row for is logged with its cost left empty, never estimated.
- chat-latest
- gpt-4.1-mini
- gpt-4o
- gpt-4o-2024-05-13
- gpt-4o-mini
- gpt-5.3-codex
- gpt-5.6-cyber
- gpt-5.6-luna
- gpt-5.6-sol
- gpt-5.6-terra
- gpt-6-astra
- gpt-6-luna
- gpt-6-sol
Price per token, by model: GPT-6 Sol, GPT-6 Luna, GPT-5.6 Sol, GPT-5.6 Terra, GPT-5.3 Codex, GPT-4o, GPT-4o mini.
Setup
Every route has the shape https://api.vigil.tools/{user_id}/{provider}/{path}. For OpenAI the provider segment is openai, so the base URL is https://api.vigil.tools/{user_id}/openai and a request path such as /v1/chat/completions is forwarded unchanged.
Two headers: X-Vigil-Key, your Vigil key from the dashboard, and X-Vigil-Agent, the name the dashboard groups this traffic under. Your OpenAI key stays where it was, in OPENAI_API_KEY.
import OpenAI from "openai";
const client = new OpenAI({
apiKey: process.env.OPENAI_API_KEY, // unchanged
baseURL: "https://api.vigil.tools/{user_id}/openai/v1",
defaultHeaders: {
// MUST be X-Vigil-Key: Authorization already carries your OpenAI key.
"X-Vigil-Key": "vk_your_vigil_key",
"X-Vigil-Agent": "my-agent",
},
});
const completion = await client.chat.completions.create({
model: "gpt-6-luna",
messages: [{ role: "user", content: "Hello!" }],
});from openai import OpenAI
client = OpenAI(
base_url="https://api.vigil.tools/{user_id}/openai/v1",
default_headers={
"X-Vigil-Key": "vk_your_vigil_key",
"X-Vigil-Agent": "my-agent",
},
)
completion = client.chat.completions.create(
model="gpt-6-luna",
messages=[{"role": "user", "content": "Hello!"}],
)curl https://api.vigil.tools/{user_id}/openai/v1/chat/completions \
-H "content-type: application/json" \
-H "Authorization: Bearer $OPENAI_API_KEY" \
-H "X-Vigil-Key: vk_your_vigil_key" \
-H "X-Vigil-Agent: my-agent" \
-d '{
"model": "gpt-6-luna",
"messages": [{ "role": "user", "content": "Hello!" }]
}'Replace {user_id} and vk_your_vigil_key with the values on your Connect page. The snippets above are generated by the same code as that page.
Prompt caching on OpenAI
This provider caches automatically, on its own terms. Vigil records what it reports and prices it — there is nothing for Vigil to add to the request.
With optimisation On: the provider caches automatically or not at all; Vigil records cache use and does not change the request. In Shadow, Vigil measures what caching would have saved and changes nothing. Off records the call and nothing more.
Questions
- Does Vigil see or store my OpenAI API key?
- No, not the key itself. Your OpenAI key passes through the proxy with the request and is never written to disk. A call record can keep a short one-way digest of the credential, so that one key’s cache is kept apart from another’s; it cannot be turned back into the key. Vigil identifies you by the X-Vigil-Key header and your user id in the URL.
- Does routing OpenAI through Vigil add latency?
- Some. The proxy runs on Cloudflare’s edge, so the added hop is short, but reading the request body to place cache markers takes time, and there is a bounded lookup for your account state. The dashboard shows total latency per call. For long calls, stream: Cloudflare’s edge closes non-streaming connections after roughly 125 seconds.
- Which OpenAI models does Vigil price?
- 13 models with a published rate in the registry: chat-latest, gpt-4.1-mini, gpt-4o, gpt-4o-2024-05-13, gpt-4o-mini, gpt-5.3-codex, gpt-5.6-cyber, gpt-5.6-luna, gpt-5.6-sol, gpt-5.6-terra, gpt-6-astra, gpt-6-luna, gpt-6-sol. A call to any other OpenAI model is logged and monitored, with its cost left empty rather than guessed.
Other providers: Anthropic · AWS Bedrock · Google Vertex · Google Gemini · Mistral · xAI Grok · DeepSeek · Together AI · Fireworks AI · Groq · Cerebras · Baseten · Moonshot · Z.ai · Cloudflare Workers AI · OpenRouter