Lower your costs
The Improve screen reads the traffic of your project, and lists what to change. Each finding shows the numbers that it comes from.
Read the findings
Open Improve in the console. The readouts on top count the findings:
| Readout | Shows |
|---|---|
| Findings | All findings |
| Worth doing first | The findings of high severity |
| Worth a look | The other findings |
| Worth fixing | The estimated saving of each month, in dollars, when Proxium can compute one |
What to change lists the findings, high severity first. There are three kinds:
| Kind | Proxium flags | Severity |
|---|---|---|
| Cost | One of your largest spend lines this month, at least $10, when the same vendor has a chat model that costs less than 70% of it | High from $25 of estimated saving each month |
| Cache | A project with at least 100 calls and $5 of spend this month, where the cache saved less than 2% | High from $20 of estimated saving |
| Reliability | A vendor with at least 20 attempts in the last 7 days, where 15% or more failed | High from 50% failed |
If the list says Nothing to flag, your traffic has none of these patterns.
Move traffic to a cheaper model
For a cost finding:
- Read the cheaper model that the finding names, and its price.
- Open Routing › Model catalog, and compare Cost, per Mtok of the two models.
- Test the cheaper model on a few real prompts of the app.
- In Text routing, select Change on the tier of the app, and put the cheaper model first.
- Select Save for this project.
Your apps need no change, because they send the tier name. To move one app only, add a rule for that app. See Route requests to models.
Let the cache answer repeated calls
For a cache finding: the cache answers only a request that is the same, field for field, as an earlier one. These changes give more hits:
| Change | Reason |
|---|---|
| Keep the system prompt the same on each call | A changed character is a new request |
Send the same temperature and tools each time | Both fields are part of the cache key |
| Put a date or an id in a message only when the answer depends on it | Each new value is a new request |
Overview › Saved by cache shows what the cache saved. The response cache explains the key.
Fix a vendor that fails often
For a reliability finding:
- Open Requests, and read What failed and Who is unreliable for that vendor.
- If the reason is a rejected key or a used quota, fix the key at the vendor, or add it again on Providers.
- Else, move the vendor down in your chains, so that a better vendor answers first.
Each failed attempt adds time before the answer. Debug a failed call shows how to read the attempts.
Related pages
- Track spend: where the money goes.
- Set budgets and limits: cap the spend of one app.