Use the Responses, embeddings, rerank and moderations APIs
Proxium serves four more OpenAI APIs with the same base URL and virtual key as chat:
| Route | What it does | Routing type | Stream |
|---|---|---|---|
POST /v1/embeddings | Turns text into vectors | Embeddings | No |
POST /v1/responses | The OpenAI Responses API | Text | Yes |
POST /v1/rerank | Orders documents by how well they match a query | Text | No |
POST /v1/moderations | Checks text against a content policy | Text | No |
Proxium sends your request body to the same path at the vendor, for example <base URL of the vendor>/embeddings. It sends the answer of the vendor back without change.
The vendor must serve that path. For example, a Groq vendor serves chat. A rerank call that goes to it fails, and Proxium tries the next model of the chain.
Make embeddings
- Add a vendor that serves embeddings, for example the OpenAI (embeddings) preset. See Add vendors and your own keys.
- Get the model id from
GET /v1/models, for examplet.acme.openai-embed/text-embedding-3-small. - Send the id in
model:
- Python
- curl
result = client.embeddings.create(
model="t.acme.openai-embed/text-embedding-3-small",
input=["The first text.", "The second text."],
)
print(len(result.data), len(result.data[0].embedding))
curl https://proxium.tech/v1/embeddings \
-H "Authorization: Bearer $PROXIUM_KEY" \
-H "Content-Type: application/json" \
-d '{"model": "t.acme.openai-embed/text-embedding-3-small", "input": ["The first text."]}'
Proxium does not store or search the vectors. It records the tokens and the cost of the call.
To send a tier name in place of the id, set a chain for the tier embed in Routing › Embeddings routing.
Use the Responses API
Send the call to a vendor that serves the Responses API, such as OpenAI. Set stream to true to get server-sent events.
response = client.responses.create(
model="t.acme.openai/gpt-4o-mini",
input="Write one sentence about queues.",
)
print(response.output_text)
Rerank and moderations
Send the body that your vendor documents. Proxium needs only the model field:
curl https://proxium.tech/v1/rerank \
-H "Authorization: Bearer $PROXIUM_KEY" \
-H "Content-Type: application/json" \
-d '{"model": "t.acme.my-reranker/bge-reranker-v2-m3", "query": "refund policy", "documents": ["Refunds take 5 days.", "We ship worldwide."]}'
Here my-reranker is a server that you added with A server I run myself, with the request type Text.
What is the same as chat, and what is not
| Chat and Messages | These four routes | |
|---|---|---|
| Virtual key, ceilings, cost record | Yes | Yes |
| Routing and failover | Yes | Yes |
| Row in the Attempt log | Yes | Yes |
| Response cache | Yes | No. Each call goes to a vendor |
x-proxium-memory: recall | Adds memories | No effect. /v1/responses keeps the call in memory as write |
A call without model | 400 bad_json | 400 missing_model |
Related pages
- Route requests to models: set the chain of a tier.
- API reference: the body of each route.