Use the Anthropic SDK
Proxium speaks the Anthropic Messages API at POST /v1/messages. A Messages call can go to any vendor of your project, not only Anthropic. Proxium translates the call and the answer.
| Setting | Value |
|---|---|
| Base URL | https://proxium.tech, without /v1. The SDK adds /v1/messages |
| API key | A virtual key of your project, from Keys in the console |
| Model | A model id or a tier name of your project. Route requests to models explains both |
Connect the client
To connect, set the base URL and give your virtual key as the API key.
- Python
- Node
- curl
import os
from anthropic import Anthropic
client = Anthropic(
base_url="https://proxium.tech",
api_key=os.environ["PROXIUM_KEY"],
)
import Anthropic from "@anthropic-ai/sdk";
const client = new Anthropic({
baseURL: "https://proxium.tech",
apiKey: process.env.PROXIUM_KEY,
});
curl https://proxium.tech/v1/messages \
-H "x-api-key: $PROXIUM_KEY" \
-H "Content-Type: application/json" \
-d "{\"model\": \"$PROXIUM_MODEL\", \"max_tokens\": 256, \"messages\": [{\"role\": \"user\", \"content\": \"Hello\"}]}"
Send a message
A request is a list of messages and a limit on the length of the reply. Give model and max_tokens: without one of them, Proxium answers 400.
- Python
- Node
message = client.messages.create(
model=os.environ["PROXIUM_MODEL"],
max_tokens=512,
system="You answer in one sentence.",
messages=[{"role": "user", "content": "What is a virtual key?"}],
)
print(message.content[0].text)
const message = await client.messages.create({
model: process.env.PROXIUM_MODEL,
max_tokens: 512,
system: "You answer in one sentence.",
messages: [{ role: "user", content: "What is a virtual key?" }],
});
console.log(message.content[0].text);
The answer has the Anthropic message shape:
| Field | Holds |
|---|---|
content[0].text | The reply |
model | The model that answered |
usage | The token counts. Proxium prices the request from them |
Stream the answer
With streaming, the reply arrives in small pieces while the model writes it, so your app can show it as it grows. To stream, use the stream helper of the SDK, or set stream to true. Proxium sends the pieces as Anthropic server-sent events.
- Python
- Node
- curl
with client.messages.stream(
model=os.environ["PROXIUM_MODEL"],
max_tokens=512,
messages=[{"role": "user", "content": "Write a haiku about queues."}],
) as stream:
for text in stream.text_stream:
print(text, end="", flush=True)
print()
const stream = client.messages.stream({
model: process.env.PROXIUM_MODEL,
max_tokens: 512,
messages: [{ role: "user", content: "Write a haiku about queues." }],
});
stream.on("text", (text) => process.stdout.write(text));
await stream.finalMessage();
process.stdout.write("\n");
curl -N https://proxium.tech/v1/messages \
-H "x-api-key: $PROXIUM_KEY" \
-H "Content-Type: application/json" \
-d "{\"model\": \"$PROXIUM_MODEL\", \"max_tokens\": 512, \"stream\": true, \"messages\": [{\"role\": \"user\", \"content\": \"Write a haiku about queues.\"}]}"
Check which fields reach the model
| Proxium passes on | Proxium drops |
|---|---|
system, messages, tools, tool_choice, temperature, top_p, top_k, stop_sequences, stream, metadata.user_id | cache_control, container, mcp_servers, service_tier |
Text, image, tool_use and tool_result blocks | thinking, redacted_thinking and document blocks |
Proxium drops cache_control, so Anthropic prompt caching does not apply to a call through Proxium.
Handle errors
An error comes in the Anthropic error shape, and the SDK raises the class for its status. A body that is too large is the one exception: it gets a 413 with a plain-text body.
{
"type": "error",
"error": {
"type": "authentication_error",
"message": "invalid or revoked key"
}
}
| Status | error.type |
|---|---|
| 400, 415, 422 | invalid_request_error |
| 401 | authentication_error |
| 402, 403 | permission_error |
| 404 | not_found_error |
| 413 | request_too_large |
| 429 | rate_limit_error |
| 529 | overloaded_error |
| Any other status | api_error |
A 429 has a retry-after header with the seconds to wait. The message names the limit, for example source budget exceeded for 'support-bot' (calls). Errors lists each reason.
Next steps
- Route requests to models: model ids, tiers and the fallback order.
- Track spend: name each app with
x-proxium-source, and see what it costs.