Skip to main content
Field protection rewrites sensitive columns as the very first step after ingest, so every later stage only ever sees the protected value. You opt in per field with a protect object on the source schema.

Why

Datris sends samples of your data to an AI provider when a pipeline uses an AI rule, an AI transformation, or error explanation. Field protection runs before all of them. A protected column reaches the AI provider, Live Read, and every destination only in its protected form, so the raw value of that column is never sent to a model provider. That means you do not need a business associate agreement (or a similar data-processing agreement) with the provider for the columns you protect. Columns you do not protect are sent as they are. Protection is local: no value leaves the Datris server to be protected, and no status line, error message, or log line carries a raw value.

Methods

Empty values stay empty for every method. An encrypt value is longer than its input: enc:v<n>: followed by the base64url of the input’s UTF-8 bytes plus 28 bytes (a 12-byte IV and a 16-byte authentication tag). A 20-character email becomes 71 characters. If the destination column is a varchar(n), size it for the ciphertext, not the original value.

Config example

To keep the ingest object after protection, add "protection": {"purgeSource": false} at the top level (see Source retention). From an agent, create_pipeline takes the same policies as a protect map keyed by field name, for example {"mrn": {"method": "hmac"}, "ssn": {"method": "drop"}}.

Suggestions

Datris can propose a starting point. The configured CodeGen model receives the field names and types only (one name:type line per field) and returns, for each field, hmac, mask with an optional preserve, redact, drop, or no protection (it never suggests encrypt; choose that yourself), with a one-line reason. No row value is sent, and nothing is saved: you confirm the fields you want and then save them as protect.
Pass {"pipeline": "<name>"} to use a saved pipeline’s source schema, or {"fields": [{"name": "...", "type": "..."}]} for a schema you have not saved yet. Each entry in the response has name, type, suggested ({"method", "preserve"}, or null for no protection), reason, and current: the field’s existing protect, so running it again on a protected pipeline never reads as an overwrite. The model field names the model that answered. Every suggestion follows the type rules: a field that is not string is only ever suggested drop or nothing, and preserve only comes with mask. With pipeline, the saved destination is checked too: a keyFields column is only ever suggested hmac or nothing, and a field whose destination type is not string is only suggested drop or nothing. A JSON pipeline’s _json field is not sent to the model; list its top-level keys in fields instead. One call takes at most 150 fields; for a wider table, pass subsets through fields (more returns HTTP 400 before the model is called). The call needs the same permission as reading a pipeline, is recorded as a pipeline / protect-suggest event in the audit log (field names only, failed model calls included), and returns HTTP 400 with {"error": "..."} when AI is not configured. Names can mislead, so treat the answer as a proposal. From an agent, the suggest_field_protection tool returns one line per field; the agent shows them to you and asks which to accept before passing any of them as protect to create_pipeline.

In the wizard

In the pipeline wizard’s Source Schema step, every field has a Protect select (None, hmac, mask, redact, drop, encrypt), and choosing mask adds a Keep select (mask all, last 4, email domain, year). A choice the type rules do not allow, such as hmac on an int field, switches back to None with the reason shown under the field. A column you chose as a key field is offered only None and hmac (a suggestion for any other method is declined with the reason), and a field protected with anything but hmac is left out of the destination’s Key Fields list. Suggest protection sends the field names and types only, never data, to the suggestions endpoint and fills each empty Protect select with its suggestion, marked with the reason until you Keep it (which applies the suggestion, even over a value you had already chosen) or Clear it; nothing is saved until you save the pipeline. JSON and XML pipelines show no Protect column, so set protect on a JSON pipeline’s top-level keys in the config JSON instead. The pipeline’s detail page lists its protected fields and their methods.

Where it runs

The stage runs once per job, after the preprocessor and before data quality, transformation, Live Read, and every destination. The preprocessor still sees raw values; if your preprocessor endpoint must not receive them, protect the data upstream or do not use a preprocessor on that pipeline. Each run that protects data adds one status line such as Protected 3 fields: mrn=hmac, email=mask:domain, ssn=drop. Column-level lineage shows protected fields as derived with the method as evidence, and dropped fields as dropped.

Type rules

  • hmac, mask, redact, and encrypt produce a string, so the field must be string in the source schema and, when a destination schema lists it, string there too.
  • drop works on a field of any type.
  • A column named in the destination’s keyFields may only use hmac: mask and redact would give different rows the same key (so upserts would merge them), encrypt gives the same value a different ciphertext on every run (so upserts would never match), and drop would remove the key.
  • The source must be delimited (CSV and similar) or JSON. For JSON pipelines the source schema lists _json plus one entry for each protected top-level key; nested keys are not supported yet.
  • Unknown methods, preserve on anything but mask, and the reserved methods fpe and tokenize are rejected when the pipeline is saved. Like other validation errors, the save returns HTTP 500 with {"error": "..."} naming the field.

The key

hmac uses one key per environment, stored in Vault as the secret <environment>/field-protection, field key. The server issues it the first time a protected run needs it and records an key / issue / field-protection event in the audit log. The key is read fail-closed: if Vault cannot be read, the run fails rather than writing data with a new key. Datris never rotates the key automatically. Rotating it yourself changes every pseudonym, so new values will no longer match or join with data already landed. encrypt uses separate, versioned keys in the same secret: fields enc.v1, enc.v2, and so on, with encCurrent naming the version new values are encrypted with. The server issues enc.v1 the first time an encrypt run needs it (audited as key / issue / field-protection). Keys never leave the server and no user ever holds one. See Key rotation. The field-protection secret cannot be referenced by a tap or a pipeline.

Revealing a value

encrypt is the one reversible method. To read an original value back, send the ciphertexts you read from the destination to the reveal endpoint:
The response has one slot per input value, in order:
  • Who can call it: an API key granted protect:reveal, or an admin session. No key template and neither the editor nor the viewer role carries it, so it must be granted explicitly in the API Keys tab (the full-access template and legacy unscoped keys hold it through *:*). There is no MCP tool, Assistant action, or CLI command for it (the only UI path is the Search screen below, which calls this same endpoint), and the recovery agent’s key does not have it, so an agent can never reveal a value. The reveal and rotate endpoints enforce their capability themselves, so CAPABILITY_ENFORCEMENT=log-only does not open them.
  • What is checked: the field must carry protect with method encrypt (otherwise HTTP 400). Each ciphertext is bound to its pipeline and field name, so a value copied to another pipeline or column does not reveal: that slot is null and errors gets {"index", "message"}, and the rest of the call still succeeds.
  • Limit: at most 1000 values per call (more returns HTTP 400).
  • What is logged: every call is a protect / reveal event in the audit log naming who called, the pipeline, the field, and how many values were requested, revealed, and failed (outcome warning when any failed). The values themselves and the ciphertexts are never logged or recorded. A call without the capability is refused with HTTP 403 and recorded as a security denied event.
Reveal is stateless: Datris never queries a destination to find the values, so you choose exactly which rows to reveal. From the Search screen. In Search → Traditional, a cell holding a ciphertext shows a lock badge with its key version (enc:v2) instead of the base64. When the result set has encrypted cells, admins (or anyone, when the install uses API keys without user logins) see Reveal encrypted values in the results header; editors and viewers see the badges but no button. One click reveals every encrypted cell on screen through the endpoint above, one call per encrypted column (split at 1000 values), so the audit log gets one protect / reveal entry per column per click. Revealed cells show the plaintext with an open-lock badge; a value the server could not reveal keeps its lock and shows the server’s reason on hover. Databricks, Snowflake, and object-store queries use the pipeline selected in the form. For PostgreSQL and MongoDB the pipeline is found from the table or collection you queried; if no pipeline or more than one writes it, pick the pipeline beside the button. Without the capability the screen shows “You do not have the protect:reveal capability” and nothing is retried. Nothing is stored: plaintext stays only in the page, and running the query again shows the ciphertext.

Key rotation

To start encrypting with a new key, rotate it with a key that has protect:admin (or an admin session):
The response is {"version": 2}. Runs after this write enc:v2: values. The old versions stay in the secret, so rows already landed with enc:v1: still reveal; Datris does not re-encrypt landed rows. The rotation is recorded as a key / rotate / field-protection event. The hmac key is never rotated by this endpoint. To retire an old version, save the field-protection secret (Secrets tab or PUT /api/v1/secrets/field-protection) without its enc.v<n> field, using a key that has protect:admin. Never remove the current version’s field: encCurrent must keep naming a stored version. The highest version number cannot be removed, and at least one version must remain. After that, every value encrypted with that version can no longer be revealed; re-landing those rows from the original source is the only way back. The rotate endpoint never reuses a retired version number.

Editing the key secret through the API

The secrets API guards <environment>/field-protection:
  • key (the hmac key) cannot be changed or removed through the API by anyone. The request fails with 409 and nothing is written. Changing it changes every hmac pseudonym, so the only supported path is a deliberate edit of the secret in Vault.
  • Adding, changing or removing an enc.v<n>, or changing encCurrent needs protect:admin. A key with only secret:write gets 403 and a security / denied audit entry; nothing is written. Each enc.v<n> value must be 64 hex characters (a 32-byte AES-256 key) and encCurrent must name a stored version; otherwise the request fails with 400. The next run after an accepted edit encrypts with the new current version; no restart is needed.
  • Accepted edits are audited with the changed field names only, never the values: key / rotate / field-protection when an enc.v<n> was added or changed, key / retire / field-protection when one was removed, and key / set-current / field-protection when only encCurrent moved.
  • A save that leaves every key field as it is (fields sent back masked or blank) works with plain secret:write, so the Secrets tab can round-trip the secret.
  • Deleting the secret (DELETE /api/v1/secrets/field-protection) is refused with the same 409 for everyone: the next protected run would issue a new hmac key and every landed ciphertext would be orphaned.
  • Creating the secret through the API is refused with 409: the server issues it on the first protected run.
  • Field names are limited to key, encCurrent, enc.v<n> (with n from 1 to 1000000, no sign or leading zero) and createdByKeyLabel. Any other field, including look-alikes such as Key or enc.v05 and a _type tag, fails with 400 for every caller. encCurrent follows the same number rule.
  • Version numbers are never reused. Retiring removes older versions only: removing every enc.v<n>, or the highest version number, fails with 400. A new enc.v<n> must be numbered above the highest stored version; re-adding a retired lower number fails with 400.

Source retention

Protecting the loaded data is not enough if the raw file stays on disk or in a bucket. As soon as the protected copy exists, Datris deletes:
  • the raw staged file, including the payload a preprocessor received as input;
  • the ingest object, when the run came from a bucket drop, a Kafka temp object, or an archive drop (the archive and its extracted files both go).
POST /pipeline/upload streams the file straight into staging and writes no bucket object, so for uploads only the staged file applies. Each purge is recorded as a pipeline / purge-source event in the audit log. A purge that fails is a warning on the run, never a failed run. Because the raw source is gone, a later failure in the same run cannot be replayed from Datris: re-upload the file or re-run the tap. To keep the ingest object, set "protection": {"purgeSource": false} at the top level of the pipeline. The staged file is still deleted. Pipelines without any protect field never purge anything.

What still sees raw data

Some helpers read the file you upload before any pipeline (and so before any protection) exists:
  • schema generation (POST /pipeline/generate, used by the wizard and by create_pipeline);
  • data profiling (POST /pipeline/profile);
  • JSON and XSD schema generation;
  • files attached to an Assistant chat.
Run these on synthetic rows that have the same shape as your data, or use an AI provider you have an agreement with for that data.

Not yet supported

  • Format-preserving encryption and tokenization.
  • Nested JSON keys.
  • XML sources.
  • Vector destinations.