llm·yard

What happens to a request between your app and the provider

Fig. 1 · one request, end to end

app · inbox app · batch app · agents no provider key on any of them 1 key AX−1 guardrail screen 6 classes, off until you give one an action 403 · blocked nothing spent, no call slot response cache · exact + semantic looked up before the reserve, so a hit is free hit ends here $0.00, no call slot used budget reserve durable · survives a restart · downroutes at 90% router tier → model · failover chain circuit breaker per endpoint · 3.0× price guard cost meter priced at the provider’s own rate openai platform pool anthropic wire, translated openrouter platform pool your own key BYOK, sealed at rest real credentials live only here byte-faithful OpenAI response · real cost recorded per call
  1. Your apps hold no provider keyThey send one base URL and one virtual key.
  2. Guardrail screen6 classes, off until you give one an action. A block costs nothing.
  3. Response cacheExact, then semantic. A hit ends here, at $0.00, using no call slot.
  4. Budget reserveDurable, survives a restart, downroutes at 90% instead of cutting off.
  5. RouterTier to model, with a failover chain, a breaker per endpoint and a 3.0× price guard.
  6. Cost meterPriced at the provider's own rate, per model.
  7. The provideropenai, anthropic, openrouter, or your own key. Real credentials live only here.
  8. Back to your appA byte-faithful OpenAI response, with the real cost recorded per call.

Eight steps, in order

Every endpoint runs them.
#StepWhat happensIf it fails
1AuthenticateThe bearer virtual key resolves the project. Identity is never taken from a client header.401
2Screen6 classes, redact or block, set per project. Off until you give a class an action.403 guardrail_blocked
3Look up the cacheExact first, then semantic. A hit returns here.Treated as a miss. Both layers fail open
4Reserve the budgetDurable, and it survives a restart.Downroute at 90%, refuse at 100%
5RouteSender override, then the project default, then straight through.400 no models available for route
6Prove the endpointThe breaker must be closed and the price inside the 3.0× guard.The next entry in the chain
7Inject the credentialServer-side only. The real key never reaches a client.502
8Record the costRead from the provider’s answer and priced per model.—
Why step 2 runs before steps 3, 4 and 7

Redaction runs first so a card number never reaches a cache embedding, a stored request body or the provider. One rewrite covers all three. A blocked request stops there too, before the budget step, so it uses no call slot.

Routing

The first rule that matches wins.
OrderRuleSourceExample
1Sender overrideaxon_route_overridesx-axon-source: batch asking for standard goes somewhere else
2Project defaultaxon_tier_defaultsEverything else asking for standard
3Passthroughthe request itselfopenrouter/qwen-3-235b is used as sent

You name the tiers. Most teams set up trivial, standard and heavy, and you can call them anything you like, because a tier is a row in your routing table. A key can point at any tier the gateway serves. Two names are reserved: auto and classifier pick a model for each request. A new key starts on trial until you move it.

If a project has no vendor key of its own and no grant on the shared pool, it can't route anywhere. The key still authenticates and the tier still resolves, then every candidate is skipped, and the call fails with 400 no models available for route.

A routing change or a revoked key reaches every replica within 30 seconds.

The response cache

Off until you set CACHE_ENABLED=true.
LayerStoreKey or matchCost of a hit
ExactDragonflysha256(tenant + model + endpoint + body)$0.00
Semanticpgvectorcosine ≥ 0.95, inside 3600s, same tenant and model$0.00

The cache is checked before the budget, so a hit makes no provider call and uses no call slot. A semantic hit is copied into the exact layer, and the next repeat is found by its hash. On a miss, the embedding from the lookup is reused for the write, so each miss is embedded once. If either layer is down, the request goes to the provider.

Caution · a mismatch here fails silently

CACHE_EMBED_DIM must equal the vector(N) column width. The shipped column is vector(1024). If they disagree, every embedding is dropped: the semantic layer stops working, the exact layer keeps hitting, and the deployment looks healthy.