LiteLLM vs Vigil
What LiteLLM does and what Vigil does, side by side. Every LiteLLM fact below was read on LiteLLM's own pages on 2026-10-09 and links to the page it came from; a feature its pages do not mention is marked “not stated”, never “no”. Where LiteLLM is stronger, this page says so.
What LiteLLM is, and what Vigil is
LiteLLM, in its own words: “The AI Gateway for platform teams. One API for 100+ LLMs, with spend tracking, auth, and guardrails.”litellm.ai. Both: a Python SDK in your process, or a self-hosted proxy your OpenAI client points at (base_url http://0.0.0.0:4000).
Vigil is a proxy for AI API calls. Change one base URL, add one header, and every call is logged with its cost, tokens, latency, errors and agent, then forwarded unchanged. Where the provider supports explicit prompt caching, Vigil can place the cache markers itself when optimisation is switched on. It stores no prompt text and caches no responses.
- LiteLLM says: enterprise pricing is "never per token".litellm.ai
- LiteLLM says: "Free forever · self-hosted"; no hosted cloud tier appears on the pricing page.litellm.ai
Feature by feature
The LiteLLM column was read on LiteLLM's own pages on 2026-10-09; each cell links to the page. “Not stated” means the pages read do not mention it, which is not the same as “no”. The Vigil column is read from Vigil's code and docs.
| Feature | LiteLLM | Vigil |
|---|---|---|
| Integration | SDK (from litellm import completion) or proxy: run litellm, point the OpenAI SDK at it.github.com | Proxy. Change the base URL and add the X-Vigil-Key header; no SDK. |
| Open source | MIT, except the enterprise/ directory under the BerriAI Enterprise license.raw.githubusercontent.com | No. Vigil is not open source. |
| Hosting | Self-host only: "It runs in your cloud, so prompts and responses never leave your environment."docs.litellm.ai | Cloud only: a Cloudflare Worker in front of the provider, a dashboard on Vercel, data in Supabase (Frankfurt, EU). |
| Cost per call | Yes: an x-litellm-response-cost header, a spend-logs table, spend by key, user and team.docs.litellm.ai | Yes: cost, tokens, latency, status and agent on every call, from a rate registry where every rate carries its source, date and provenance; a dash where no rate is published. |
| Providers | The homepage says "100+ LLMs", "140+ LLM providers" and "1,800+ models" in different places.litellm.ai | 17 providers routed by segment. |
| Prompt caching | Passes your markers through and translates them for Bedrock and Vertex; can also inject cache_control at configured points ("automatically cache system messages").docs.litellm.ai | Creates cache hits: Anthropic, Google Vertex, Mistral get cache markers placed by Vigil when optimisation is on; AWS Bedrock when the request carries an API key rather than a SigV4 signature. On other providers Vigil records the provider’s own caching. |
| Response caching | Yes: Redis, semantic (Redis, Valkey, Qdrant), S3, GCS, local and disk back ends.docs.litellm.ai | No. Vigil does not cache responses. |
| Routing and fallbacks | Routing strategies (shuffle, usage, latency, least-busy, cost), fallbacks, retries with backoff, cooldowns.docs.litellm.ai | Model routing is in beta: Sends a conversation to a cheaper model from the same provider — one tier down — when its opening's size and tools are in a band the agent's routing level covers and the cheaper model can take the request; every call of the conversation gets the same answer. No provider fallbacks or load balancing. |
| Prompt management | Through integrations (.prompt files, Langfuse, Humanloop); no versioning of its own.docs.litellm.ai | No prompt management or versioning. |
| Evals | Not stated on the pages read. | No evaluations. A score can be sent for a call through the feedback endpoint. |
| Tracing | Exports to Langfuse, Arize Phoenix, LangSmith and OpenTelemetry; no trace UI of its own claimed.litellm.ai | One record per call. No spans or traces. |
| Alerts | Budget, slow and failed calls, model outage; Slack, Discord, Microsoft Teams. Region-outage alerting is an enterprise feature.docs.litellm.ai | Cost spikes, failures and findings are flagged in the dashboard. Nothing is sent by email, Slack or push. |
| Prompt injection and PII | Through guardrail integrations: "PII masking and prompt-injection guardrails (Presidio, Lakera, and more)"; several require an enterprise license.docs.litellm.ai | Flags likely prompt injections and personal data in prompts, on every plan. Nothing is blocked or redacted. |
| Budgets and rate limits | max_budget and budget_duration per key, user and team; requests and tokens per minute. Per-model key budgets need an enterprise license.docs.litellm.ai | No spend caps or rate limits on your provider traffic. Optimisation can be switched off per agent. |
| Seats | Organisations, teams and user roles; organisation and team admin roles are premium features.docs.litellm.ai | Seats per plan: 1 on Free, 3 on Plus, unlimited from Pro. Teammates are invited; there are no role tiers. |
| Compliance | "LiteLLM is SOC 2 Type II audited"; ISO 27001 on the enterprise page.docs.litellm.ai | No SOC 2 or ISO 27001 report. A DPA is published. |
| Data residency | Self-hosted in your own cloud or air-gapped; no vendor region.docs.litellm.ai | Account and telemetry stored in the EU (Supabase, Frankfurt). The proxy runs on Cloudflare’s edge. No prompt text or responses are stored. |
Where LiteLLM is stronger
- Open source under MIT, self-hosted in your own cloud, so prompts never leave it.
- The widest routing: strategies, fallbacks, retries, cooldowns, and a response cache with many back ends.
- Budgets and rate limits per key, user and team, and alerts to Slack, Discord and Teams.
- Prompt-cache injection at configured points, plus translation of markers for Bedrock and Vertex.
Where Vigil differs
- It creates prompt-cache hits rather than reporting them, on the providers whose caching lets a proxy place markers, and measures what caching would have saved on the others.
- Every rate it prices with carries its source, its retrieval date and a provenance label, and a cost it cannot stand behind is a dash, not a number.
- Monitoring is not gated by plan: every agent, every provider, on Free. Paid plans switch optimisation on.
- It does less: no response cache, no evals, no prompt management, no traces, no alerts outside the dashboard.
Pricing, side by side
LiteLLMlitellm.ai
- Open Source $0: self-hosted, no caps stated: 100+ providers, virtual keys, teams, spend tracking, budgets, fallbacks
- Enterprise Annual, by usage: "sized to your annual gateway request capacity"; SSO, SCIM, audit logs, 24/7 support with SLAs; no figure published
Vigil vigil.tools/pricing
- Free $0: every agent monitored; 10,000 logged calls a month, 7 days of history; every optimisation On during your 7-day trial, as Plus, then Shadow only
- Plus $29/mo: every optimisation On up to $100 saved a billing month, then Shadow until the next; 500,000 logged calls, 30 days
- Pro $99/mo: every optimisation On up to $500 saved a billing month, then Shadow until the next; 2,000,000 logged calls, 90 days
- Supreme $299/mo: every optimisation On up to $2,000 saved a billing month, then Shadow until the next; Unlimited logged calls, 365 days
- Enterprise Custom: by contract
Monthly, excl. VAT. Monitoring is free on every plan; what a paid plan buys is a higher limit on the money Vigil saves you per billing month before optimisations go to Shadow.
Other comparisons: Helicone vs Vigil · Portkey vs Vigil · Langfuse vs Vigil · OpenRouter vs Vigil