Debug a failed call
When a call fails, first find out who refused it: Proxium, before any vendor call, or the vendors that Proxium tried.
1. Read the status and the code
Status and code | Who refused | Vendor called | Next step |
|---|---|---|---|
401 missing_key, 401 invalid_key | Proxium | No | Check the key. See Errors |
400 no_route | Proxium | No | Read the message. Your project has no model for this request |
429 tenant_capped, 429 source_capped | Proxium | No | A ceiling. Wait retry-after seconds. See Set budgets and limits |
502 upstream_failed | The vendors | Yes, each model of the plan | Do step 2 |
The message of 502 upstream_failed gives the reason of the last attempt, for example:
{
"error": {
"message": "client deadline exceeded",
"type": "upstream_failed",
"code": "upstream_failed"
}
}
2. Find the call on the Requests screen
- In the console, open Requests.
- Select the Time window:
today,7dor30d. - Optional: to see one app only, select it in the app bar. All apps shows every app.
The Attempt log has one row for each attempt. A call that failed over has more than one row:
| When | Model | Asked for | Try | Outcome | Took |
|---|---|---|---|---|---|
| 14:02:11 | t.acme.groq/llama-3.3-70b-versatile | standard | 1 | error | 1,820 ms |
| 14:02:13 | t.acme.mistral/mistral-small-latest | standard | 2 | success | 640 ms |
The first model failed, and Proxium tried the second model of the chain. Your app got the answer of the second model.
| Outcome | Meaning |
|---|---|
success | The model answered |
error | The model answered with an error, or the connection failed |
timeout | The model did not answer in time |
circuit_open | Proxium skipped the vendor, because its circuit breaker was open. See Routing |
An image, speech or video call has one row, with Try 1: these routes do not fail over. A video that fails while it renders gets a second row, with the reason video_failed, when your app polls it. The drill-down of a media call shows the request as JSON, and the error text of a failed call. It never shows the image, the audio or the video.
3. Read what went out and what came back
Select a row of the Attempt log. The row opens, with each attempt of the call.
Each attempt starts with a strip: the outcome, the try, the model, the tier that you asked for, the finish reason and the tokens.
| You see | It means |
|---|---|
| finish stop | The model finished its answer |
| finish length | The model reached its output limit |
| finish tool_calls | The model asked your app to run a tool |
| "No answer text: the model stopped at its output limit…" | The model used all its output tokens to reason, and wrote no answer. Raise max_tokens |
For a failed attempt, Came back comes first: it is what the vendor answered. Sent upstream shows the request as messages. The messages that earlier calls sent are one folded row. The messages new in this call are open.
To see the exact JSON, select Raw. Then Copy request and Copy answer copy it.
The row shows a prompt only if the project stored it. A new project stores the prompts of failed attempts only. Step 4 shows how to change that.
4. Keep the prompts that you need
The setting What to store, in Settings › Stored prompts and answers, decides what the drill-down can show:
| Value | The drill-down shows |
|---|---|
off | No prompt and no answer |
errors (default) | The prompt and the answer of each failed attempt |
all | The prompt and the answer of each attempt |
Any member can change it. Data handling says how long Proxium keeps them.
5. Find a pattern
To see if one vendor or one model fails often, read the panels under the Attempt log:
| Panel | It shows |
|---|---|
| What failed | The failed attempts, by outcome, vendor and reason, with a count |
| Who is unreliable | Each vendor, with its share of the failures and the kinds of failure |
| Model health | Each model: attempts, 429 answers, 410 Gone answers, other errors and the average time of a success |
If one model fails often, put a different model first in the chain. See Route requests to models.
Related pages
- Errors: every code, with its cause and fix.
- Route requests to models: retries and failover.