Skip to main content

Route requests to models

The model field of each request decides which model serves it. This guide shows how proxium reads that field, and how you change the result in the console.

Send a tier name or a model id​

The model field holds one of two things:

  • A tier name, for example standard. proxium resolves the tier to an ordered chain of models. The Routing screen lists the tiers you can send.
  • A concrete provider/model id, for example openai/gpt-4o-mini. proxium does not resolve a tier. It tries that model first. If that model fails, proxium fails over to other models that the project can use. Refer to Read the failover order.

To see the provider/model ids your project can use, call GET https://proxium.tech/v1/models. See List models.

curl https://proxium.tech/v1/chat/completions \
-H "Authorization: Bearer $PROXIUM_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"standard","messages":[{"role":"user","content":"hello"}]}'

proxium finds the request type from the path you call, not from the body. A chat call uses the Text chains. An embeddings call uses the Embeddings chains. A rule for one request type never applies to another.

How proxium resolves a tier​

proxium tries the narrowest scope first. The first scope that has a chain for the tier wins.

  1. Sender rule. A chain for the x-proxium-source value of the request.
  2. Key rule. A chain for the key of the request. An endpoint's own routing is a key rule.
  3. Project chain. The chain that your project set for the tier.
  4. Default chain. The chain that every project uses until it sets its own.
  5. Passthrough. No scope has a chain, so proxium uses the model value as a model id.

If the request has no x-proxium-source header, proxium skips step 1. The Scoped rules panel on the Routing screen shows this order as precedence: source → key → account → default. In that text, account is your project.

Use the meta tiers​

proxium reserves two tier names: auto, classifier. They do not name a fixed weight.

  • auto asks proxium to choose the tier. When auto-classify is on, a classifier call picks one of the configured chat tiers per request. If that call fails, proxium uses standard. When auto-classify is off, auto resolves through its own chain, like any other tier.
  • classifier is the chain that the classifier call runs on. It serves no request directly.

Auto-classify is one setting for all projects. Today it is off on proxium.tech, so auto resolves through its own chain. On the Routing screen, each meta tier row shows in use or not used, with the reason.

If auto-classify is on, a key rule or an endpoint rule written for auto never matches. The reason: proxium replaces auto with a concrete tier before it resolves the route. Write rules for the concrete tiers.

Know what the key tier does​

Each key has a tier. If you name no tier when you create a key, the key gets trial. The Keys screen calls this field Tier, and the Endpoints screen calls it Budget tier.

The key tier does not choose a model. The model field of each request chooses the route. See Budgets and limits.

Set your project's own chain for a tier​

Every project uses the default chain for a tier until it sets its own. You do this on the This project view of the Routing screen.

  1. Open the Routing screen in the console.
  2. Find the table for the request type, for example Text routing.
  3. On the row of the tier, click Change.
  4. Under Pick the model to try first, choose the first model.
  5. Under Pick a fallback, add each failover model in the order you want.
  6. Click Save for this project.

The row now shows custom, and the Resolves to column shows your chain. To return to the default chain, click Use default on the row. A saved empty chain also returns the tier to the default.

proxium applies the change on the server that saved it at once. Every other server applies it within 30 seconds.

Route one application differently​

Use a sender rule when one application must use a different chain from the rest of the project. The application identifies itself with the x-proxium-source header.

  1. Open the Routing screen.
  2. In the Scoped rules panel, click Add a rule.
  3. Set Applies to to A sender.
  4. In Sender name, type the value that the application sends in x-proxium-source.
  5. Select the Tier, and build the chain.
  6. Click Save rule.

proxium applies a rule as it applies a project chain: at once on one server, and within 30 seconds on all.

Then send the header from that application:

curl https://proxium.tech/v1/chat/completions \
-H "Authorization: Bearer $PROXIUM_KEY" \
-H "Content-Type: application/json" \
-H "x-proxium-source: inbox-worker" \
-d '{"model":"standard","messages":[{"role":"user","content":"hello"}]}'

A sender rule belongs to your project only. Another project that uses the same sender name does not read or change your rule.

To route one key differently, set Applies to to A key and type the key prefix. The key must belong to your project. On the Endpoints screen, the routing of an endpoint is a key rule of this type.

Read the failover order​

proxium builds one attempt plan for each request. The plan holds at most four entries.

  1. proxium adds the models of the resolved chain, in chain order. Hop 1 is first.
  2. If the plan has fewer than four entries, proxium adds other models that your project can use, in catalog order.

proxium skips a model of another project, and a model outside the allowed_models list of the key.

For each entry, proxium tries the model. If the answer is a 5xx, a 408, a 429 or a "warming" answer, proxium tries the same model one more time. It waits for the Retry-After value of the provider, up to 30 seconds. Then it moves to the next entry.

If every entry fails, the caller gets 502 with the code upstream_failed. The message gives the reason of the last attempt. If the plan is empty, the caller gets 400 with the code no_route.

A streamed answer cannot fail over after its first byte reaches the caller.

To limit the time of the whole plan, send x-proxium-timeout-ms. proxium caps the value at 120000 ms. If you send no header, each buffered attempt has a limit of 120 seconds, and the plan as a whole has none. See Limits.

Read the circuit breaker​

proxium keeps one circuit breaker for each provider credential. If your project uses its own key for a provider, that key has its own breaker.

StateWhat proxium does
closedSends calls to the provider.
openSkips the provider and moves to the next entry of the plan.
half_openSends trial calls. Two successes close the breaker.

The breaker opens after five failures inside ten minutes. A rejected credential or an exhausted quota opens the breaker of a credential that one project owns at the first failure. An open breaker allows a trial call after 30 seconds, or after 30 minutes for a rejected credential or an exhausted quota.

A 429, a "warming" answer, an unsupported model and a non-retryable request error do not move the breaker.

To read the breaker states before a call, use GET https://proxium.tech/status/peek. See Budgets and limits.

Know the price guard state​

proxium has a failover price guard. It drops a failover model that costs more than a set multiple of hop 1. The guard is not active on proxium.tech today.

warning

Failover can reach a model that costs more than hop 1. The catalog step can also add models that are not in your chain. Check the price of each model in the Model catalog on the Routing screen.