Skip to main content

Sensitive data scanning

A developer can paste a card number or a live API key into a prompt by mistake. Proxium is the last stop before the request goes to a vendor. So it is the last place where that value can still be stopped.

:::info State on proxium.tech A new project starts with card numbers and credentials on redact, with an alert, and every other class off. You set each class on the Data protection screen. See Protect sensitive data. :::

What it finds​

ClassFindsCheck against false matches
credit_cardCard numbersThe Luhn check digit
ibanBank account numbersThe mod-97 check
emailEmail addresses
phonePhone numbers
credentialAWS, GitHub, Slack, Stripe, OpenAI, Anthropic, Google and GitLab keys, JWTs, private key blocks, passwords in URLs such as postgres://user:pass@host, Authorization: Bearer values, *_KEY=, *_SECRET= and *_TOKEN= assignmentsThe known format of each key, or a long random value after a secret word
card_securityAn expiry date (12/27) or a security code (CVV 123)Only within 48 characters of a card number
national_idUS SSN, UK National Insurance number, Romanian CNP, Spanish DNI and NIE, Italian codice fiscaleThe check digit or check letter, or the strict structure of the number
ip_addressIPv4 and IPv6 addressesEach part in range; no version numbers such as 1.2.3.4.5
customThe words and phrases of your project's listA whole word, in any case
prompt_injectionKnown instruction-override shapes, also in base64

The scan reads the whole request body except the model field. Each matcher reads the text one time, from start to end, with no regular expressions. One automaton finds your terms in a single pass too. So a long or a hostile prompt cannot slow it down.

What it does with a match​

Each class has one of four actions:

ActionThe requestThe match
offGoes out unchangedNot looked for
flagGoes out unchangedLogged, and an alert is sent
redactGoes out with a placeholder in place of the valueLogged, and an alert if you turned it on
blockIs refused with 403 guardrail_blockedLogged, and an alert if you turned it on

Use flag to watch real traffic before you choose redact or block.

Fig. 1 · what the scan does with a card number
request"Refund card4242 4242 4242 4242"scanone passfirst403 · blockedcredit_card on blocknothing sentredact[REDACTED:credit_card:1]replaces the valueunchangedno match, or flagonly the placeholder reachesresponse cachestored promptsmemoryvendornew project: redactnew project: redactrequest"Refund card 4242 4242 4242 4242"scan · one pass, first403 · blockednothing is sentredact[REDACTED:credit_card:1]unchanged · no match, or flagonly the placeholder reachesresponse cachestored promptsmemoryvendor

Example. The prompt Refund card 4242 4242 4242 4242, please, with credit_card set to redact, goes to the vendor as Refund card [REDACTED:credit_card:1], please. The same value two times gets the same number, so the model can still tell that two mentions are one card.

With block, your app gets:

403 body
{
"error": {
"message": "blocked by egress guardrail: credit_card detected at /messages/0/content (the matched value is deliberately not echoed)",
"type": "guardrail_blocked",
"code": "guardrail_blocked"
}
}

Where the value never goes​

The scan runs first in a call, before the response cache, the stored prompts, the memory and the vendor. A redacted or blocked value is therefore never written anywhere.

Each match writes an audit record: the project, the app, the key prefix, the time, the class and the action. It also names the place in the request, such as /messages/0/content. The record never holds the matched value. The refusal message and the alert do not hold it either. A flag sends the value to the vendor unchanged, so it is in the stored prompts if your project keeps them.

Alerts​

An alert goes to the channels of Settings › Spend alerts: Slack, Discord, Telegram or a webhook. Proxium sends at most one alert for each class and app in 10 minutes. The alert says how many requests matched since the last alert of that class and app, and it never holds the value.