llm·yard

Every AI call, measured, and rerouted when a model fails

See every AI call your company makes.

Fig. 1 · what you see once your AI calls go through proxium

your apps→proxium→your providers

Cost alerts

proxium · #ai-spend · 14:02

Sales is at 80% of its daily budget: $40.12 of $50.

Every request

AppTaskRouteTookCostResult
support-botreply-ticket-4411Anthropic API1.4 s$0.011ok
code-reviewreview-pr-872OpenAI API → Anthropic API6.2 s$0.084escalated
sales-agentdraft-email-58OpenRouter API, 429, tried again3.4 s$0.006retried
nightly-batchtag-doc-19230cache12 ms$0cache hit

Cost by team, today

  • Engineeringcode-review $42.10
  • Salessales-agent $40.1280% of its budget
  • Supportsupport-bot $18.40
  • Datanightly-batch $6.30

Failures, and why

  • OpenAI API 212
  • OpenRouter API 31
  • Anthropic API 6
  • rate limited
  • timed out
  • server error

What to change

  • Move nightly-batch to a cheaper modelSame provider, $0.30 instead of $2.50 per million tokens.about $170 a month
  • Turn on the cache for Supportsupport-bot asks the same questions, and pays for each one.about $90 a month
  • Let Jev pick the tierShort questions go to trivial instead of standard.about $60 a month

Example data. The teams, apps and numbers are made up.

Set it up

Your apps change two lines.
  1. Add your provider keys.OpenAI, Anthropic, Groq, Mistral, DeepSeek and the rest. They're sealed on the server, and your apps never hold them.
  2. Point your apps at proxium.Change the base URL and the key. Ask for a tier you named, or send a provider/model and it passes straight through.
  3. Watch the calls come in.Move a tier to another model in the console and every app that asks for it follows, with no deploy.
client = OpenAI(
    base_url="https://proxium.tech/v1",
    api_key=os.environ["PROXIUM_KEY"],
)

r = client.chat.completions.create(
    model="standard",     # a tier you named, not a model id
    messages=[...],
)

Point one app at it and look at a day of its calls.