Skip to main content

The request flow

A chat call to proxium passes through a fixed series of stages. This page describes those stages in order, and the reason for the order. It applies to /v1/chat/completions and /v1/messages. The /v1/messages route translates the Anthropic shape onto the same flow.

The flow at a glance​

Authenticate​

proxium reads the bearer key and finds the project of that key. The project comes from the key only. No header can name a different project.

A missing key gets 401 missing_key. An unknown, revoked or expired key gets 401 invalid_key. proxium checks the key before it parses the body. A caller with no valid key therefore gets 401, and not an error about the body.

Add memories​

If the project has memory on and the call sends x-proxium-memory: recall, proxium adds the project's memories to the messages. This happens before the cache lookup. A call that carries one end user's memories never uses the near-match layer of the cache. See Use project memory.

Pick the tier and resolve the chain​

proxium resolves the model field to an ordered chain of models. It tries a sender rule, then a key rule, then the project chain, then the default chain. See Route requests to models.

When auto-classify is on, a request for auto gets a concrete tier from a classifier call before this stage. Auto-classify is not active on proxium.tech today.

proxium resolves the chain before the cache lookup for one reason. The cache key holds the resolved model. Two requests that resolve to the same model can share one cache entry. A request that resolves to a different model never gets an answer from another model.

Look up the cache​

proxium looks up the request in the response cache. The cache has an exact layer and a near-match layer. See Caching.

On a hit, proxium returns the stored answer at once. A streamed request gets the stored answer as a stream of events.

The cache lookup comes before the budget reservation. Because of this order, a hit uses no call slot, calls no provider and costs $0.00. A hit also bypasses the circuit breakers, because no provider takes part.

A project can charge a fraction of the cost that a hit saved. It sets this under What a cached answer costs on the Settings screen. Then a hit goes through the budget reservation, uses a call slot and records the charged amount.

If the key already used its full window, proxium skips the near-match layer. The near-match layer makes a paid embedding call, and a call that proxium will refuse must not buy one. proxium does not record that embedding call against your project. The exact layer still answers. With the default charge of 0, a capped key still gets an exact hit.

Reserve the budget​

proxium checks the limits in this order:

  1. The monthly cap of the project. No project on proxium.tech has one during the open beta.
  2. The ceilings of the key: calls per hour, requests per minute and tokens per minute.
  3. The ceilings of the sender, when the project set one for the x-proxium-source value.

For a key or sender ceiling, the check and the reservation are one atomic step, shared by all proxium servers. Two servers cannot both pass a call that only one slot allows. The reservation survives a restart of a server.

A refusal stops the flow here, before any provider call. The caller gets 429, with the code tenant_capped or source_capped. See Set budgets and limits.

Check each entry and call it​

proxium builds a plan of at most four entries from the chain. It drops every model that the key may not use.

For each entry, proxium makes three checks before the call:

  • The deadline. If x-proxium-timeout-ms ran out, proxium stops and answers 502.
  • The credential. If the project registered its own key for the provider, proxium uses it. Otherwise it uses the credential that proxium holds.
  • The circuit breaker. If the breaker of that credential is open, proxium skips the entry.

proxium adds the provider credential on the server, at each attempt. Your app never holds it.

The failover price guard is not active on proxium.tech today. A failover entry can cost more than hop 1.

x-proxium-timeout-ms bounds the whole plan, and not one attempt. A caller can ask proxium to give up sooner. A caller cannot ask for more than the maximum. See Limits.

Record the cost​

proxium reads the token counts from the answer of the model that served the call. It prices those counts with the price of that model. If failover changed the model, the record names the model that answered.

proxium adds the cost to the windows of the project and the key. It also adds the cost to the sender window when the sender has a ceiling. Then it stores the answer in the cache, and returns the provider's answer to your app without change. See Track spend.

A streamed call follows the same flow. proxium records its cost when the stream ends.

Other endpoints​

The embeddings, responses, moderations and rerank routes take a shorter path. They authenticate, reserve the budget, resolve the chain, fail over and record the cost in the same way. They have no response cache, no auto-classify and no soft cap.

On these routes, proxium reserves the budget before it resolves the chain.