Set budgets and limits
Many apps can share one key. If one of them loops, or gets a traffic spike, it can spend money for all of them. To stop that, give the app a limit. Above its limit, Proxium refuses the calls of that app. The other apps keep running.
You can also give a key its own limits. Above them, Proxium refuses every call of that key.
What Proxium checks before each call
- 1Did the key reach one of its limits?The limits that you set for the key on Keys. A new key has noneyes →Refused: 429 tenant_capped
- 2Did support-bot reach its own limit?The limit that you set for this app on Overviewyes →Refused: 429 source_capped
A refused call never reaches a vendor, so it costs nothing.
Set a limit for one app
An app is the name that it sends in the x-proxium-source header. You can give each app two limits:
| Limit | Means | Example |
|---|---|---|
| Calls per hour | The most calls in any 60 minutes | 500 |
| Dollars per day | The most spend in any 24 hours | 5 |
You set both on the console screen Overview, in the panel Per-app burn. There is no API for them.
To give support-bot at most 500 calls an hour and $5 a day:
- Make sure the app sends
x-proxium-source: support-bot. Track spend shows how. - Open Overview, and select the row support-bot in Per-app burn.
- Type
500in Calls per hour, and5in Dollars per day. - Select Set ceiling.
The Ceiling column of the row shows the new limits. 0 in a field means no limit. To remove both limits, select the row, then Remove ceiling.
An app shows in Per-app burn only after it made a call in the selected time window. If you do not see it, make one call, or select 30d.
The limits apply at once on the server that saved them, and within 30 seconds on the others.
Set the limits of a key
A key can have four limits. A new key has none.
| Limit | Means | Example |
|---|---|---|
| Calls per hour | The most calls in any 60 minutes | 1000 |
| Requests per minute | The most calls in any 60 seconds | 60 |
| Tokens per minute | The most tokens in any 60 seconds | 100000 |
| Dollars per day | The most spend in any 24 hours | 20 |
To give a key at most 60 calls a minute and $20 a day:
- Open Keys.
- On the row of the key, select Limits.
- Type
60in Requests per minute, and20in Dollars per day. Leave the other fields empty. - Select Save limits.
The Limits column of the row shows the new limits. An empty field means no limit. The key of an endpoint is on Keys too.
The limits apply to the next call of the key.
When an app reaches its limit
Proxium answers the call with HTTP 429. The retry-after header says to wait 60 seconds. The body says which app, and which limit:
{
"error": {
"message": "source budget exceeded for 'support-bot' (calls)",
"type": "source_capped",
"code": "source_capped"
}
}
(calls) means the calls of the last hour. (cost) means the spend of the last day.
Your app should wait, then try again. With the OpenAI SDK:
- Python
- Node
import time
import openai
try:
response = client.chat.completions.create(model="standard", messages=messages)
except openai.RateLimitError as e:
wait = int(e.response.headers.get("retry-after", "60"))
time.sleep(wait)
response = client.chat.completions.create(model="standard", messages=messages)
async function ask() {
return client.chat.completions.create({ model: "standard", messages });
}
let response;
try {
response = await ask();
} catch (e) {
if (e.status !== 429) throw e;
const wait = Number(e.headers?.["retry-after"] ?? 60);
await new Promise((r) => setTimeout(r, wait * 1000));
response = await ask();
}
On /v1/messages, the body has the Anthropic shape, with the type rate_limit_error. The message is the same.
How the limits are counted
- Calls per hour counts each call that Proxium let through in the last 60 minutes. A refused call does not count.
- Dollars per day adds the cost of the calls in the last 24 hours.
- Proxium checks the spend before a call, not during it. So the last call that starts under the limit can end a little above it.
- A free cache hit counts toward neither limit. See The response cache.
- All Proxium servers count together. Two servers cannot both let through the last allowed call.
Check the key before a call
GET https://proxium.tech/status/peek checks your key, and says if your next call will be refused. It does not count as a call, and it costs nothing. Use it before a batch of work.
curl https://proxium.tech/status/peek \
-H "Authorization: Bearer $PROXIUM_KEY"
A wrong or revoked key gets 401 invalid_key. A valid key gets this answer:
{
"tenant_capped": false,
"budget": { "calls_pct": 0, "cost_pct": 0 },
"credits": { "over": false },
"providers": [{ "provider": "t.acme.groq", "state": "closed" }],
"any_provider_available": true
}
| Field | Means |
|---|---|
tenant_capped | true: the key reached a limit, and the next call gets 429 tenant_capped |
budget.calls_pct, budget.cost_pct | How much of the limits of the key is used, in percent. 0 for a key with no limit |
credits.over | true: the project reached a monthly cap. No project has one today |
providers | Each vendor of the project: closed is working, open is skipped. See Routing |
any_provider_available | true: at least one vendor is not skipped, or the list is empty |
/status/peek does not see the limits of an app. A call can get 429 source_capped while tenant_capped is false.
A check before a batch:
import os
import requests
state = requests.get(
"https://proxium.tech/status/peek",
headers={"Authorization": f"Bearer {os.environ['PROXIUM_KEY']}"},
timeout=5,
).json()
if state["tenant_capped"] or not state["any_provider_available"]:
raise RuntimeError("Proxium will refuse this call now. Hold the work.")
Related pages
- Monitor your project: watch the spend and the failures.
- Track spend: see what each app spent.
- Errors: every 429.