Inference
Chat, Messages, Responses, embeddings, moderation and rerank calls. Each call goes through the request loop of proxium: authentication, budget, routing, failover and cost. The response cache applies to chat and Messages calls only.
Create a chat completion
Runs an OpenAI chat completion request through the full proxium loop: authentication, memory, the response cache, the budget, routing, failover and cost. A cache hit costs no call slot. The answer is the provider's body without change. With `stream: true`, the answer is server-sent events. A cache hit is sent as events too.
Create embeddings
Sends an OpenAI embeddings request through the proxium loop: authentication, the budget, routing, failover and cost.
Create a message (Anthropic)
Accepts an Anthropic Messages request and runs it through the same proxium loop as `/v1/chat/completions`: proxium translates the request to the OpenAI chat shape, routes it to any provider, and translates the answer back to the Anthropic shape. With `stream: true`, the answer is Anthropic server-sent events. Each error body is in the Anthropic shape.
Create a moderation
Sends an OpenAI moderation request through the proxium loop: authentication, the budget, routing, failover and cost.
Rerank documents
Sends a rerank request to the `/rerank` path of the provider, through the proxium loop: authentication, the budget, routing, failover and cost.
Create a model response
Sends an OpenAI Responses request through the proxium loop: authentication, memory, the budget, routing, failover and cost. With `stream: true`, the answer is server-sent events. The response cache does not apply to this route.