Skip to main content

Get started with proxium

Who calls

Developercurl, OpenAI SDK
Coding agentClaude Code, Codex, Cursor · MCP
Your appOpenAI or Anthropic SDK
proxiumone base URLone virtual key

LLM APIs · your keys

  • OpenAI
  • Anthropic
  • Groq
  • OpenRouter
  • Mistral
  • DeepSeek
  • Together AI
  • Google Gemini
  • your own server

Memory

Project memorycoding agents read and write it over MCP

a dashed box connects over MCP

proxium is an OpenAI-compatible gateway. Your apps send their model calls to proxium with one virtual key. proxium calls your provider with your own provider key. If a model fails, proxium tries the next model of the chain. It records the cost of each call.

Change two lines​

Point the SDK that your app uses at https://proxium.tech/v1, and send your virtual key as the API key.

app.py
import os
from openai import OpenAI

client = OpenAI(
base_url="https://proxium.tech/v1",
api_key=os.environ["PROXIUM_KEY"],
)

answer = client.chat.completions.create(
model=os.environ["PROXIUM_MODEL"],
messages=[{"role": "user", "content": "Say hello in five words."}],
)
print(answer.choices[0].message.content)

PROXIUM_KEY is a virtual key of your project. PROXIUM_MODEL is a model id from GET /v1/models, or a tier name of your routing policy. The Quickstart shows how to get both.

Choose your path​

What proxium does with a chat call​

  1. It finds the project of the virtual key.
  2. It resolves the model field to a chain of models. See Routing.
  3. It looks for the answer in the response cache. A cache hit makes no provider call and costs nothing by default. See The response cache.
  4. It counts the call against the ceilings that you set for the calling application. See Budgets and limits.
  5. It calls the models of the chain in order, with your provider key, until one answers.
  6. It records the cost: the token counts from the answer, times the price of the model. See Track spend.

The full order, with the reason for each step, is in The request flow.