Use the OpenAI SDK
This guide tells you how to send OpenAI API calls through proxium. You need a virtual key from the Keys screen of the console. The examples read it from the environment variable PROXIUM_KEY.
Set the base URL and the key
Set the base URL to https://proxium.tech/v1. Set the API key to your virtual key. proxium reads the key from the Authorization: Bearer header, which the OpenAI SDK sends.
- Python
- Node
- HTTP
import os
from openai import OpenAI
client = OpenAI(
base_url="https://proxium.tech/v1",
api_key=os.environ["PROXIUM_KEY"],
)
import OpenAI from "openai";
const client = new OpenAI({
baseURL: "https://proxium.tech/v1",
apiKey: process.env.PROXIUM_KEY,
});
curl https://proxium.tech/v1/models \
-H "Authorization: Bearer $PROXIUM_KEY"
Name a model
The model field takes one of two forms.
| Form | Example | What proxium does |
|---|---|---|
A concrete provider/model id | t.<project>.openai/gpt-4o-mini | Calls that model first |
| A tier name | standard | Resolves the tier to an ordered chain of models, and calls the first one |
To get the concrete ids, send GET /v1/models. The answer lists each provider/model id that your project can use, in data[].id. The Routing screen of the console shows the tiers and their chains.
proxium resolves a tier name in this order. It uses the first chain that it finds:
- The chain of the tier for the application that the
x-proxium-sourceheader names. - The chain of the tier for the virtual key.
- The chain of the tier for the project.
- The default chain of the tier.
If a model fails, proxium tries the next model of the chain. Then it tries other models of the providers that your project can use. This applies to a concrete id too. proxium tries at most 4 models for one call. The Requests screen of the console shows which model answered. For the details, refer to Routing.
Send a chat call
response = client.chat.completions.create(
model="standard",
messages=[{"role": "user", "content": "Summarise this ticket in one line."}],
)
print(response.choices[0].message.content)
The answer is in the OpenAI chat completion shape. A cache hit returns the stored answer, and proxium does not call a provider.
Stream the answer
Set stream to true. proxium sends the answer as server-sent events, in the OpenAI chunk format.
- Python
- Node
- HTTP
stream = client.chat.completions.create(
model="standard",
messages=[{"role": "user", "content": "Write a haiku about queues."}],
stream=True,
)
for chunk in stream:
if chunk.choices and chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="", flush=True)
print()
const stream = await client.chat.completions.create({
model: "standard",
messages: [{ role: "user", content: "Write a haiku about queues." }],
stream: true,
});
for await (const chunk of stream) {
process.stdout.write(chunk.choices[0]?.delta?.content ?? "");
}
process.stdout.write("\n");
curl -N https://proxium.tech/v1/chat/completions \
-H "Authorization: Bearer $PROXIUM_KEY" \
-H "Content-Type: application/json" \
-d '{"model": "standard", "stream": true, "messages": [{"role": "user", "content": "Write a haiku about queues."}]}'
proxium adds stream_options.include_usage to the request that it sends to the provider. It uses the usage to record the cost. Thus the stream can end with a chunk that has usage and an empty choices list. The examples above check choices before they read it for this reason.
A cache hit on a streamed call also comes as server-sent events. POST /v1/responses streams too. Embeddings, moderations and rerank do not stream.
Make embeddings
Send the id of an embedding model. To get one, add an embedding vendor, such as the OpenAI (embeddings) preset. Then read the id from GET /v1/models.
import os
result = client.embeddings.create(
model=os.environ["PROXIUM_EMBED_MODEL"],
input=["The first text.", "The second text."],
)
print(len(result.data), len(result.data[0].embedding))
Set PROXIUM_EMBED_MODEL to an id such as t.<project>.openai-embed/text-embedding-3-small.
Name your application
Send the x-proxium-source header with a name for your application. proxium uses the name to select the routing chain and the budget of that application. It also records the cost against it. If you do not send the header, proxium uses the project slug.
Set the header one time on the client:
- Python
- Node
- HTTP
client = OpenAI(
base_url="https://proxium.tech/v1",
api_key=os.environ["PROXIUM_KEY"],
default_headers={"x-proxium-source": "support-bot"},
)
const client = new OpenAI({
baseURL: "https://proxium.tech/v1",
apiKey: process.env.PROXIUM_KEY,
defaultHeaders: { "x-proxium-source": "support-bot" },
});
curl https://proxium.tech/v1/chat/completions \
-H "Authorization: Bearer $PROXIUM_KEY" \
-H "x-proxium-source: support-bot" \
-H "Content-Type: application/json" \
-d '{"model": "standard", "messages": [{"role": "user", "content": "Hello"}]}'
The Requests screen of the console can show the calls of one application. For the other headers, refer to Request headers.
Call the endpoints that have no /v1 path
These endpoints are at the origin, https://proxium.tech, and not under /v1. The OpenAI SDK has no method for them. Call them with plain HTTP.
| Endpoint | Use |
|---|---|
GET /status/peek | Read the budget use of the key before a call |
GET /usage/me | Read the spend of the project for this month |
POST /video/generations | Start a video |
POST /video/operations | Read the state of a video |
POST /video/download | Download a finished video |
curl https://proxium.tech/usage/me \
-H "Authorization: Bearer $PROXIUM_KEY"
Handle errors
proxium sends each error in the OpenAI error shape. The SDK raises it as an exception. The code field holds the reason, for example no_route or tenant_capped. For the list of codes, refer to Errors.