Track spend
Proxium records the cost of each call that a model answered. You read the cost in the console, or from the API.
Where to read the spend
| You want to know | Where |
|---|---|
| The spend of today, 7 days or 30 days | Console: Overview › Spent |
| The spend of each application | Console: Overview › Per-app burn |
| The spend of each virtual key | Console: Overview › Per-key burn |
| The spend of each model | Console: Overview › Model split. API: by_model of /usage/me |
| The spend of each task | API: by_task of /usage/me |
| What the cache saved | Console: Overview › Saved by cache, and Requests › Cache |
| The spend of this month, from your code | API: GET /usage/me |
Each console screen has a Time window: today, 7d or 30d.
How Proxium computes the cost
Proxium reads the token counts from the usage block of the answer. It multiplies each count by the price of the model that answered.
Example: a call to a model at $0.15 for each million input tokens and $0.60 for each million output tokens.
| Tokens | Count | Price for each million | Cost |
|---|---|---|---|
| Input | 1,200 | $0.15 | $0.00018 |
| Output | 300 | $0.60 | $0.00018 |
| The call | $0.00036 |
Two more token types have their own price:
| Tokens | Price |
|---|---|
| Prompt-cache read | The cache-read price of the model. If it has none, the input price |
| Prompt-cache write | The cache-write price of the model. If it has none, the input price |
A vendor counts its cache-read tokens inside the input total. Proxium takes them out of the input count first, so it prices each token one time.
To see the price of a model, open Routing › Model catalog, column Cost, per Mtok.
Cases that change the cost
| Case | Recorded cost |
|---|---|
| The model has no price, or a price of $0 | $0.00. A new model that Proxium finds at a vendor starts at $0 |
A stream with no usage block | Proxium asks the vendor for one with stream_options.include_usage. If none arrives, it counts 4 characters as 1 token |
| A second model answered after a failure | The cost of the model that answered |
| A cache hit | $0.00 by default. See What a hit costs |
Name each application
To see the spend of each application, send the x-proxium-source header with the name of the application. Set it one time, on the client:
- Python
- Node
- curl
client = OpenAI(
base_url="https://proxium.tech/v1",
api_key=os.environ["PROXIUM_KEY"],
default_headers={"x-proxium-source": "support-bot"},
)
const client = new OpenAI({
baseURL: "https://proxium.tech/v1",
apiKey: process.env.PROXIUM_KEY,
defaultHeaders: { "x-proxium-source": "support-bot" },
});
curl https://proxium.tech/v1/chat/completions \
-H "Authorization: Bearer $PROXIUM_KEY" \
-H "Content-Type: application/json" \
-H "x-proxium-source: support-bot" \
-d "{\"model\": \"$PROXIUM_MODEL\", \"messages\": [{\"role\": \"user\", \"content\": \"hello\"}]}"
A call with no header counts under the name of the project.
The name also selects the ceiling of the application, and its routing rule. See Set budgets and limits and Route requests to models.
To group calls by a task of your own, also send x-proxium-task-id, for example ticket-4812. /usage/me then gives the spend of each task in by_task.
Give each virtual key the name of its application when you create it. Per-key burn shows a key without a name as unnamed.
Read the spend from the API
GET https://proxium.tech/usage/me gives the spend of the project in the current calendar month. The path has no /v1.
curl https://proxium.tech/usage/me \
-H "Authorization: Bearer $PROXIUM_KEY"
{
"tenant_slug": "acme",
"period": "month",
"credits": { "unlimited": true, "used_millicents": 184250, "used_usd": 1.8425 },
"requests": 412,
"tokens": 903114,
"by_model": [
{ "provider": "t.acme.groq", "model": "llama-3.3-70b-versatile", "cost_usd_millicents": 184250, "tokens": 903114, "requests": 412 }
],
"by_task": [
{ "task_id": "ticket-4812", "cost_usd_millicents": 2210, "tokens": 10840, "requests": 3 }
]
}
| Field | Holds |
|---|---|
credits.used_usd | The spend of this month, in dollars |
credits.used_millicents | The same spend, in millicents. 100,000 millicents are $1 |
credits.unlimited | true: the project has no monthly cap. During the open beta it is always true |
requests, tokens | The calls and the tokens of this month |
by_model | The spend of each model, highest first |
by_task | The spend of each x-proxium-task-id, highest first. At most 100 rows |
/usage/me has no split by application. For that, use Per-app burn in the console. Read the usage of the project gives every field.
Related pages
- Set budgets and limits: cap the spend of one application.
- The response cache: what a cache hit costs.
- The request flow: when Proxium records the cost.