Mechanics6 min read

Multi-provider LLM observability: where to start

Multi-provider LLM observability is one event per model call, across every provider: cost, latency, errors, injection and PII. What to instrument first.

Multi-provider LLM observability means recording every model call your product makes, whichever provider serves it, as one event with one shape: what it cost, how long it took, whether it failed or was retried, and whether it carried an injection attempt or leaked personal data. It is what lets you answer questions about the whole stack from one place instead of five vendor dashboards.

You shipped. Requests flow through OpenAI, Anthropic, a self-hosted model, maybe a routing layer, and one or two agents that call all of it. Then the questions start: which call blew up the bill last Tuesday, why p95 latency doubled, and whether a pasted support transcript put a customer's email somewhere it should not be. If you cannot answer those in under a minute, you do not have observability. You have logs.

This post covers what observability means once several providers and agents are involved, why multi-provider stacks break naive monitoring, and the order to instrument things so the first dashboard you build is one you keep. For the layers of monitoring in general, see LLM monitoring.

Five signals, tied to the same event

Traditional APM watches a service. LLM observability watches a conversation between your product and models you do not control, and that changes what you need to record.

For every request you want five things tied together: the tokens and dollars it cost, the latency it added, whether it errored or was retried, whether the input carried an injection attempt, and whether sensitive data crossed your boundary. Observability is the ability to slice all five by provider, model, feature, user and agent, from the same event.

If your data lives in five vendor dashboards and a spreadsheet, you have instrumentation. You do not have observability.

Why multi-provider stacks break naive monitoring

Single-provider logging tools assume one naming scheme, one way of counting tokens and one failure mode. The moment you route across providers, those assumptions fail:

  • Cost is not comparable. Token counts, cache pricing and streaming behaviour differ per provider, and even the same model is priced differently per platform. A raw token sum says nothing until it is converted to dollars at the right rate: see LLM token cost.
  • Latency has no shared baseline. Time to first token and total duration mean different things for streaming and non-streaming calls, and across providers.
  • Failures hide across hops. A retry on provider B hides an outage on provider A. Without a shared request ID, you are debugging two systems that never met.
  • Agents multiply the surface. One user action fans out into many model calls. Per-call metrics miss the cost and latency of the whole chain.

Fix the shape of the data once, with one event and one schema for every provider, and the rest gets easier.

The five signals to instrument first

Do not try to capture everything. In a multi-provider stack, instrument in this order.

1. Cost per request, in dollars

Start with spend, because it is what stakeholders ask about first and the signal most likely to be quietly broken. Capture input tokens, output tokens, cached tokens and the model that actually served the call, then convert to dollars at that provider's current rate. Aggregate by feature and by customer, so you can see which capability is expensive and which account is unprofitable.

2. Latency, split at the boundaries

Record time to first token and total duration separately, with the provider and model that served the call. Splitting them tells you whether the slowness is your prompt, the provider or the network. Track p50 and p95: averages hide the tail your users actually feel.

3. Errors and retries, attributed to the failing hop

Log the status, the error class and the retry count against one request ID that survives every hop. When a request touches three providers and two agents, that ID is the only thing stitching the story together.

4. Prompt injection attempts

Flag inputs that look like manipulation: instructions to ignore earlier rules, to reveal the system prompt, or to change the agent's goal. Injection detection is not a filter to bolt on later. It is telemetry you want from the first day, so you can see who is probing and how often.

5. Personal data crossing your boundary

Detect emails, phone numbers, card numbers and similar patterns in prompts and completions. The point is not only to prevent leaks, but to know whether any happened, and where.

A beginner's instrumentation checklist

  1. Put a proxy between your app and every provider, so all calls share one event schema. Why the vantage point matters is in proxy vs SDK observability.
  2. Emit one request ID per user action and carry it through every hop and agent.
  3. Capture cost, latency and status from the first day.
  4. Add injection and personal-data flags next, once the basics are flowing.
  5. Group everything by provider, model, feature, user and agent.
  6. Set a baseline before you optimise anything.

The last step is the one teams skip. You cannot tune latency or spend against numbers you never recorded.

What a healthy baseline looks like

With that checklist done, a healthy multi-provider stack shows a few recognisable shapes:

  • Spend is attributed, with no large "unknown" bucket.
  • p95 latency is stable per provider, and streaming calls report time to first token.
  • Error and retry counts sit near zero outside known outages.
  • Injection flags appear as a small, watched count. A count that stays at exactly zero for months is worth checking rather than assuming.
  • Personal-data findings are rare, and each one is explained.

You are looking for a dashboard you trust enough to act on: one screen with cost, latency, errors and safety flags for every provider at once.

Where Vigil fits

You do not need a rewrite to get there. Vigil is an LLM proxy: change one base URL, add one header, and every call through it is logged per agent with its cost, latency and errors, across the providers you already use.

It flags two kinds of safety finding, and it is worth being exact about both. Prompt injection is checked in the latest user message of each call, either as an explicit override ("ignore previous instructions") or as several weaker indicators together. Personal data is checked in the model's response: emails, phone numbers, US social security numbers and card numbers that appear in the output without having been in the request. Prompts are not scanned for personal data, and nothing is blocked or redacted; these are flags. Both show on the Errors page on every plan, and the dedicated Threats and PII pages come with Plus and above.

Monitoring is free. Paid plans switch optimisation on, priced by the spend Vigil optimises; the plans and their limits are on the pricing page.

What to do

Pick the question you could not answer last week, whether it was a cost, a latency jump or a leak, and work out which of the five signals would have answered it. Instrument cost and latency first, add the safety flags next, and record a baseline before you change a single setting. To route your existing calls through a proxy that records all five, start with LLM proxy setup.