protect object on the source schema.
Why
Datris sends samples of your data to an AI provider when a pipeline uses an AI rule, an AI transformation, or error explanation. Field protection runs before all of them. A protected column reaches the AI provider, Live Read, and every destination only in its protected form, so the raw value of that column is never sent to a model provider. That means you do not need a business associate agreement (or a similar data-processing agreement) with the provider for the columns you protect. Columns you do not protect are sent as they are. Protection is local: no value leaves the Datris server to be protected, and no status line, error message, or log line carries a raw value.Methods
Empty values stay empty for every method.
An
encrypt value is longer than its input: enc:v<n>: followed by the base64url of the input’s UTF-8 bytes plus 28 bytes (a 12-byte IV and a 16-byte authentication tag). A 20-character email becomes 71 characters. If the destination column is a varchar(n), size it for the ciphertext, not the original value.
Config example
"protection": {"purgeSource": false} at the top level (see Source retention).
From an agent, create_pipeline takes the same policies as a protect map keyed by field name, for example {"mrn": {"method": "hmac"}, "ssn": {"method": "drop"}}.
Suggestions
Datris can propose a starting point. The configured CodeGen model receives the field names and types only (onename:type line per field) and returns, for each field, hmac, mask with an optional preserve, redact, drop, or no protection (it never suggests encrypt; choose that yourself), with a one-line reason. No row value is sent, and nothing is saved: you confirm the fields you want and then save them as protect.
{"pipeline": "<name>"} to use a saved pipeline’s source schema, or {"fields": [{"name": "...", "type": "..."}]} for a schema you have not saved yet. Each entry in the response has name, type, suggested ({"method", "preserve"}, or null for no protection), reason, and current: the field’s existing protect, so running it again on a protected pipeline never reads as an overwrite. The model field names the model that answered.
Every suggestion follows the type rules: a field that is not string is only ever suggested drop or nothing, and preserve only comes with mask. With pipeline, the saved destination is checked too: a keyFields column is only ever suggested hmac or nothing, and a field whose destination type is not string is only suggested drop or nothing. A JSON pipeline’s _json field is not sent to the model; list its top-level keys in fields instead. One call takes at most 150 fields; for a wider table, pass subsets through fields (more returns HTTP 400 before the model is called). The call needs the same permission as reading a pipeline, is recorded as a pipeline / protect-suggest event in the audit log (field names only, failed model calls included), and returns HTTP 400 with {"error": "..."} when AI is not configured.
Names can mislead, so treat the answer as a proposal. From an agent, the suggest_field_protection tool returns one line per field; the agent shows them to you and asks which to accept before passing any of them as protect to create_pipeline.
In the wizard
In the pipeline wizard’s Source Schema step, every field has a Protect select (None,hmac, mask, redact, drop, encrypt), and choosing mask adds a Keep select (mask all, last 4, email domain, year). A choice the type rules do not allow, such as hmac on an int field, switches back to None with the reason shown under the field. A column you chose as a key field is offered only None and hmac (a suggestion for any other method is declined with the reason), and a field protected with anything but hmac is left out of the destination’s Key Fields list.
Suggest protection sends the field names and types only, never data, to the suggestions endpoint and fills each empty Protect select with its suggestion, marked with the reason until you Keep it (which applies the suggestion, even over a value you had already chosen) or Clear it; nothing is saved until you save the pipeline. JSON and XML pipelines show no Protect column, so set protect on a JSON pipeline’s top-level keys in the config JSON instead. The pipeline’s detail page lists its protected fields and their methods.
Where it runs
The stage runs once per job, after the preprocessor and before data quality, transformation, Live Read, and every destination. The preprocessor still sees raw values; if your preprocessor endpoint must not receive them, protect the data upstream or do not use a preprocessor on that pipeline. Each run that protects data adds one status line such asProtected 3 fields: mrn=hmac, email=mask:domain, ssn=drop. Column-level lineage shows protected fields as derived with the method as evidence, and dropped fields as dropped.
Type rules
hmac,mask,redact, andencryptproduce a string, so the field must bestringin the source schema and, when a destination schema lists it,stringthere too.dropworks on a field of any type.- A column named in the destination’s
keyFieldsmay only usehmac:maskandredactwould give different rows the same key (so upserts would merge them),encryptgives the same value a different ciphertext on every run (so upserts would never match), anddropwould remove the key. - The source must be delimited (CSV and similar) or JSON. For JSON pipelines the source schema lists
_jsonplus one entry for each protected top-level key; nested keys are not supported yet. - Unknown methods,
preserveon anything butmask, and the reserved methodsfpeandtokenizeare rejected when the pipeline is saved. Like other validation errors, the save returns HTTP 500 with{"error": "..."}naming the field.
The key
hmac uses one key per environment, stored in Vault as the secret <environment>/field-protection, field key. The server issues it the first time a protected run needs it and records an key / issue / field-protection event in the audit log.
The key is read fail-closed: if Vault cannot be read, the run fails rather than writing data with a new key. Datris never rotates the key automatically. Rotating it yourself changes every pseudonym, so new values will no longer match or join with data already landed.
encrypt uses separate, versioned keys in the same secret: fields enc.v1, enc.v2, and so on, with encCurrent naming the version new values are encrypted with. The server issues enc.v1 the first time an encrypt run needs it (audited as key / issue / field-protection). Keys never leave the server and no user ever holds one. See Key rotation.
The field-protection secret cannot be referenced by a tap or a pipeline.
Revealing a value
encrypt is the one reversible method. To read an original value back, send the ciphertexts you read from the destination to the reveal endpoint:
- Who can call it: an API key granted
protect:reveal, or an admin session. No key template and neither the editor nor the viewer role carries it, so it must be granted explicitly in the API Keys tab (thefull-accesstemplate and legacy unscoped keys hold it through*:*). There is no MCP tool, Assistant action, or CLI command for it (the only UI path is the Search screen below, which calls this same endpoint), and the recovery agent’s key does not have it, so an agent can never reveal a value. The reveal and rotate endpoints enforce their capability themselves, soCAPABILITY_ENFORCEMENT=log-onlydoes not open them. - What is checked: the field must carry
protectwith methodencrypt(otherwise HTTP 400). Each ciphertext is bound to its pipeline and field name, so a value copied to another pipeline or column does not reveal: that slot isnullanderrorsgets{"index", "message"}, and the rest of the call still succeeds. - Limit: at most 1000 values per call (more returns HTTP 400).
- What is logged: every call is a
protect / revealevent in the audit log naming who called, the pipeline, the field, and how many values were requested, revealed, and failed (outcomewarningwhen any failed). The values themselves and the ciphertexts are never logged or recorded. A call without the capability is refused with HTTP 403 and recorded as asecuritydenied event.
enc:v2) instead of the base64. When the result set has encrypted cells, admins (or anyone, when the install uses API keys without user logins) see Reveal encrypted values in the results header; editors and viewers see the badges but no button. One click reveals every encrypted cell on screen through the endpoint above, one call per encrypted column (split at 1000 values), so the audit log gets one protect / reveal entry per column per click. Revealed cells show the plaintext with an open-lock badge; a value the server could not reveal keeps its lock and shows the server’s reason on hover. Databricks, Snowflake, and object-store queries use the pipeline selected in the form. For PostgreSQL and MongoDB the pipeline is found from the table or collection you queried; if no pipeline or more than one writes it, pick the pipeline beside the button. Without the capability the screen shows “You do not have the protect:reveal capability” and nothing is retried. Nothing is stored: plaintext stays only in the page, and running the query again shows the ciphertext.
Key rotation
To start encrypting with a new key, rotate it with a key that hasprotect:admin (or an admin session):
{"version": 2}. Runs after this write enc:v2: values. The old versions stay in the secret, so rows already landed with enc:v1: still reveal; Datris does not re-encrypt landed rows. The rotation is recorded as a key / rotate / field-protection event. The hmac key is never rotated by this endpoint.
To retire an old version, save the field-protection secret (Secrets tab or PUT /api/v1/secrets/field-protection) without its enc.v<n> field, using a key that has protect:admin. Never remove the current version’s field: encCurrent must keep naming a stored version. The highest version number cannot be removed, and at least one version must remain. After that, every value encrypted with that version can no longer be revealed; re-landing those rows from the original source is the only way back. The rotate endpoint never reuses a retired version number.
Editing the key secret through the API
The secrets API guards<environment>/field-protection:
key(the hmac key) cannot be changed or removed through the API by anyone. The request fails with409and nothing is written. Changing it changes every hmac pseudonym, so the only supported path is a deliberate edit of the secret in Vault.- Adding, changing or removing an
enc.v<n>, or changingencCurrentneedsprotect:admin. A key with onlysecret:writegets403and asecurity / deniedaudit entry; nothing is written. Eachenc.v<n>value must be 64 hex characters (a 32-byte AES-256 key) andencCurrentmust name a stored version; otherwise the request fails with400. The next run after an accepted edit encrypts with the new current version; no restart is needed. - Accepted edits are audited with the changed field names only, never the values:
key / rotate / field-protectionwhen anenc.v<n>was added or changed,key / retire / field-protectionwhen one was removed, andkey / set-current / field-protectionwhen onlyencCurrentmoved. - A save that leaves every key field as it is (fields sent back masked or blank) works with plain
secret:write, so the Secrets tab can round-trip the secret. - Deleting the secret (
DELETE /api/v1/secrets/field-protection) is refused with the same409for everyone: the next protected run would issue a new hmac key and every landed ciphertext would be orphaned. - Creating the secret through the API is refused with
409: the server issues it on the first protected run. - Field names are limited to
key,encCurrent,enc.v<n>(withnfrom 1 to 1000000, no sign or leading zero) andcreatedByKeyLabel. Any other field, including look-alikes such asKeyorenc.v05and a_typetag, fails with400for every caller.encCurrentfollows the same number rule. - Version numbers are never reused. Retiring removes older versions only: removing every
enc.v<n>, or the highest version number, fails with400. A newenc.v<n>must be numbered above the highest stored version; re-adding a retired lower number fails with400.
Source retention
Protecting the loaded data is not enough if the raw file stays on disk or in a bucket. As soon as the protected copy exists, Datris deletes:- the raw staged file, including the payload a preprocessor received as input;
- the ingest object, when the run came from a bucket drop, a Kafka temp object, or an archive drop (the archive and its extracted files both go).
POST /pipeline/upload streams the file straight into staging and writes no bucket object, so for uploads only the staged file applies.
Each purge is recorded as a pipeline / purge-source event in the audit log. A purge that fails is a warning on the run, never a failed run.
Because the raw source is gone, a later failure in the same run cannot be replayed from Datris: re-upload the file or re-run the tap.
To keep the ingest object, set "protection": {"purgeSource": false} at the top level of the pipeline. The staged file is still deleted. Pipelines without any protect field never purge anything.
What still sees raw data
Some helpers read the file you upload before any pipeline (and so before any protection) exists:- schema generation (
POST /pipeline/generate, used by the wizard and bycreate_pipeline); - data profiling (
POST /pipeline/profile); - JSON and XSD schema generation;
- files attached to an Assistant chat.
Not yet supported
- Format-preserving encryption and tokenization.
- Nested JSON keys.
- XML sources.
- Vector destinations.
