Every AI call, measured, and rerouted when a model fails
See every AI call your company makes.
Fig. 1 · what you see once your AI calls go through proxium
your apps→proxium→your providers
Cost alerts
proxium · #ai-spend · 14:02
Sales is at 80% of its daily budget: $40.12 of $50.
Every request
| App | Task | Route | Took | Cost | Result |
|---|---|---|---|---|---|
| support-bot | reply-ticket-4411 | Anthropic API | 1.4 s | $0.011 | ok |
| code-review | review-pr-872 | OpenAI API → Anthropic API | 6.2 s | $0.084 | escalated |
| sales-agent | draft-email-58 | OpenRouter API, 429, tried again | 3.4 s | $0.006 | retried |
| nightly-batch | tag-doc-19230 | cache | 12 ms | $0 | cache hit |
Cost by team, today
Failures, and why
- rate limited
- timed out
- server error
What to change
- Move nightly-batch to a cheaper modelSame provider, $0.30 instead of $2.50 per million tokens.about $170 a month
- Turn on the cache for Supportsupport-bot asks the same questions, and pays for each one.about $90 a month
- Let Jev pick the tierShort questions go to
trivialinstead ofstandard.about $60 a month
Example data. The teams, apps and numbers are made up.
Set it up
Your apps change two lines.- Add your provider keys.OpenAI, Anthropic, Groq, Mistral, DeepSeek and the rest. They're sealed on the server, and your apps never hold them.
- Point your apps at proxium.Change the base URL and the key. Ask for a tier
you named, or send a
provider/modeland it passes straight through. - Watch the calls come in.Move a tier to another model in the console and every app that asks for it follows, with no deploy.
client = OpenAI(
base_url="https://proxium.tech/v1",
api_key=os.environ["PROXIUM_KEY"],
)
r = client.chat.completions.create(
model="standard", # a tier you named, not a model id
messages=[...],
)