Limits
These are the values on proxium.tech today. A limit of 0 means no limit.
The limits you meet first
| Limit | Value |
|---|---|
| Request body | 32 MiB |
x-proxium-timeout-ms | At most 120000 ms |
| Models in one plan | 4 |
| Retry-After that Proxium waits | At most 30 s |
| App limits | None until you set them |
Size
| Limit | Value | At the limit |
|---|---|---|
| Request body | 32 MiB | 413, plain-text body |
| Vendor answer | 64 MiB | The call fails. Proxium tries the next model |
| Stored copy of an answer | 32,768 characters | Proxium cuts the stored copy. Your app gets the full answer |
| Model id of a media call | 200 characters | 400 bad_model |
A media model id can hold only A-Z, a-z, 0-9, ., _, - and :.
Time
| Limit | Value | At the limit |
|---|---|---|
x-proxium-timeout-ms | At most 120000 ms | A larger value is lowered to the maximum |
| All calls of one request, with the header | The header value | 503 upstream_failed |
| All calls of one request, without the header | No limit | |
| One call, not streamed | 120 s | Timeout. Proxium tries the next model |
| Connect to a vendor | 10 s | The call fails. Proxium tries the next model |
| Silence in a stream | 300 s | Proxium ends the stream |
- Proxium ignores an
x-proxium-timeout-msof0, a negative value, or text. - When the header time ends, the message of the
503starts withclient deadline exceeded. - For a stream, the header time ends at the first piece. After that, only the silence limit applies. A stream has no limit on its total time.
- The silence limit also runs while the vendor prepares its first piece.
Retries and failover
| Limit | Value | At the limit |
|---|---|---|
| Models in one plan | 4 by default, 1 to 4 | 503 upstream_failed |
| Extra calls to one model | 1 by default, 0 to 3 | Proxium tries the next model |
| First wait before an extra call | 200 ms by default, 0 to 5000 ms. It doubles for each next call | — |
| Retry-After that Proxium waits | 30000 ms at most by default, 0 to 30000 ms | Proxium waits, then calls again |
| Failures that open a breaker | 5 in 10 minutes | Proxium skips the vendor |
| Pause after normal failures | 30 s | One trial call |
| Pause after a rejected key or used quota | 30 minutes | One trial call |
| Trial calls that close a breaker | 2 good answers | The vendor is used again |
| Failed calls in a row that retire a model | 3 by default, 1 to 100 | The project's calls skip the model |
| Time a retired model stays retired | 60 minutes by default, 1 to 10080 | The model is back |
| Time a vendor stays retired after "no credits" or "key rejected" on your own key | 60 minutes by default, 1 to 10080 | The vendor is back, and its first call is the test call |
| Failed calls in a row across its models that retire a vendor | 5 by default, 1 to 100 | The project's calls skip every model of the vendor |
| Time a vendor stays retired after failed calls | 60 minutes by default, 1 to 10080 | The vendor is back |
- An extra call to a model happens only after a 5xx, 408 or "warming" answer. After a 429, by default, the next call goes to a model at another vendor of the plan. With no other vendor, the same model gets one more call.
- A 429 never opens a breaker, and never counts toward a retirement.
- Your project sets each "by default" value of this table in its failover policy, on Routing › When a call fails. See Set your failover policy.
- Proxium does not skip a fallback model for its price. A fallback can cost more than your first model.
Spend and rate
| Limit | Value | At the limit |
|---|---|---|
| App: Calls per hour | You set it. Default 0 | 429 source_capped |
| App: Dollars per day | You set it. Default 0 | 429 source_capped |
| Key: calls per hour | You set it on Keys. Default 0 | 429 tenant_capped |
| Key: requests and tokens per minute | You set them on Keys. Default 0 | 429 tenant_capped |
| Key: dollars per day | You set it on Keys. Default 0 | 429 tenant_capped |
| Project: monthly cap | None during the open beta |
- Each
429carriesretry-after: 60. - Calls per hour: the last 60 minutes, rolling.
- Dollars per day: the last 24 hours, rolling, in hourly buckets.
- Requests and tokens per minute: the last 60 seconds, rolling.
- You set the app limits on Overview › Per-app burn, and the key limits on Keys. See Set budgets and limits.
Names and lists
| Limit | Value | At the limit |
|---|---|---|
| Key name, on Keys | 60 characters | The name is dropped. The key has no name |
| Endpoint name | 60 characters | 400 bad_name |
| Vendor name | 40 characters | 400 bad_name |
| Models of one vendor | 200 | 400 too_many_models |
A vendor name holds only a-z, 0-9 and -, with no - at the start or the end.
How soon a change applies
| Change | Applies |
|---|---|
| Routing, app limits, the cache price | At once on one server, within 30 s on all |
| The spend that the monthly cap reads | Every 30 s |
| The model list of each vendor | Every 21600 s |