What happens to a request between your app and the provider
Fig. 1 · one request, end to end
- Your apps hold no provider keyThey send one base URL and one virtual key.
- Guardrail screen6 classes, off until you give one an action. A block costs nothing.
- Response cacheExact, then semantic. A hit ends here, at $0.00, using no call slot.
- Budget reserveDurable, survives a restart, downroutes at 90% instead of cutting off.
- RouterTier to model, with a failover chain, a breaker per endpoint and a 3.0× price guard.
- Cost meterPriced at the provider's own rate, per model.
- The provideropenai, anthropic, openrouter, or your own key. Real credentials live only here.
- Back to your appA byte-faithful OpenAI response, with the real cost recorded per call.
Eight steps, in order
Every endpoint runs them.| # | Step | What happens | If it fails |
|---|---|---|---|
| 1 | Authenticate | The bearer virtual key resolves the project. Identity is never taken from a client header. | 401 |
| 2 | Screen | 6 classes, redact or block, set per project. Off until you give a class an action. | 403 guardrail_blocked |
| 3 | Look up the cache | Exact first, then semantic. A hit returns here. | Treated as a miss. Both layers fail open |
| 4 | Reserve the budget | Durable, and it survives a restart. | Downroute at 90%, refuse at 100% |
| 5 | Route | Sender override, then the project default, then straight through. | 400 no models available for route |
| 6 | Prove the endpoint | The breaker must be closed and the price inside the 3.0× guard. | The next entry in the chain |
| 7 | Inject the credential | Server-side only. The real key never reaches a client. | 502 |
| 8 | Record the cost | Read from the provider’s answer and priced per model. | — |
Redaction runs first so a card number never reaches a cache embedding, a stored request body or the provider. One rewrite covers all three. A blocked request stops there too, before the budget step, so it uses no call slot.
Routing
The first rule that matches wins.| Order | Rule | Source | Example |
|---|---|---|---|
| 1 | Sender override | axon_route_overrides | x-axon-source: batch asking for standard goes somewhere else |
| 2 | Project default | axon_tier_defaults | Everything else asking for standard |
| 3 | Passthrough | the request itself | openrouter/qwen-3-235b is used as sent |
You name the tiers. Most teams set up
trivial, standard and heavy, and you can call them
anything you like, because a tier is a row in your routing table. A key can point at any tier
the gateway serves. Two names are reserved:
auto and classifier pick a model for each request. A new
key starts on trial until you move it.
If a project has no vendor key of its own and no grant on the shared pool, it can't
route anywhere. The key still authenticates and the tier still resolves, then every candidate is
skipped, and the call fails with 400 no models available for route.
A routing change or a revoked key reaches every replica within 30 seconds.
The response cache
Off until you set CACHE_ENABLED=true.| Layer | Store | Key or match | Cost of a hit |
|---|---|---|---|
| Exact | Dragonfly | sha256(tenant + model + endpoint + body) | $0.00 |
| Semantic | pgvector | cosine ≥ 0.95, inside 3600s, same tenant and model | $0.00 |
The cache is checked before the budget, so a hit makes no provider call and uses no call slot. A semantic hit is copied into the exact layer, and the next repeat is found by its hash. On a miss, the embedding from the lookup is reused for the write, so each miss is embedded once. If either layer is down, the request goes to the provider.
CACHE_EMBED_DIM must equal the vector(N) column width. The shipped
column is vector(1024). If they disagree, every embedding is dropped:
the semantic layer stops working, the exact layer keeps hitting, and the deployment looks
healthy.