Two meanings of “catalog”
Datris and Unity Catalog both use the word catalog, and they mean different things:- A Datris catalog is a label on taps and pipelines that groups them in the Datris UI (see Data Catalog). It lives only in Datris and never changes where data lands.
- A Unity Catalog catalog is the top level of Databricks’ three-level namespace,
catalog.schema.table. On a Databricks pipeline,destination.database.dbNameis the Unity Catalog catalog the table lands in.
finance can load into the Unity Catalog catalog main. When the metadata push is on, the Datris catalog travels with the table as the datris_catalog tag, so you can see it from Databricks too.
What Datris writes
With"unityCatalog": { "enabled": true } on a Databricks pipeline, every successful load annotates the table: a table comment naming the pipeline and source, comments on the _datris_* provenance columns, the tags datris_pipeline, datris_catalog, datris_dq_status and managed_by = datris, and datris.* table properties. See Unity Catalog metadata for the fields, the grant it needs, and GET /api/v1/pipelines/{name}/unity-catalog.
The datris_pipeline tag is what discovery uses to recognize a table Datris loaded.
Lineage in Unity Catalog
With the metadata push on, every successful load also publishes the pipeline’s lineage to Unity Catalog, so the table’s External lineage graph in Catalog Explorer shows where its data came from without opening Datris. Datris registers three external metadata objects and two lineage relationships:- The source:
datris-tap-<tap>for a tap-fed pipeline (with the tap description, its declared source, the script identity and, for HTTP taps, the endpoint URL), ordatris-upload-<pipeline>for a pipeline fed by file uploads. - The pipeline:
datris-pipeline-<pipeline>, the transform node, with the Datris catalog, the pipeline definition version and the Datris lineage API path as properties. - The table: the Delta table the pipeline loads,
catalog.schema.table. - Source to pipeline, and pipeline to table: the pipeline-to-table relationship carries the column mappings from Datris column-level lineage, one mapping per source column. Mappings Datris inferred with AI (rather than derived from the pipeline definition) are included, and they mark the whole relationship with the properties
datris.confidence = inferredanddatris.evidence. Unity Catalog keeps properties per relationship, not per column, so the flag applies to the relationship as a whole. Publishing never runs AI inference; only mappings already computed are sent. - Run properties: both relationships carry
datris.lastRunId,datris.recordCount,datris.status(passorwarn) anddatris.lastRunAt, overwritten on every run. Unity Catalog shows the latest run, not a history.
_. The objects and column mappings are only re-sent when the pipeline definition, its column lineage or the tap script changes, or when Datris recreated the table; otherwise a run updates only the run properties.
Lineage uses the workspace REST API with the same Databricks secret as the load, and needs privileges beyond the load grants:
CREATE EXTERNAL METADATAon the metastore, which only a metastore admin can grant:GRANT CREATE EXTERNAL METADATA ON METASTORE TO `<sp-application-id>`.MODIFYon the external metadata objects Datris creates (the creator owns them, so this holds unless ownership changes).MODIFYon the table, for the pipeline-to-table relationship. The setup grants on the schema already cover it.
"unityCatalog": { "enabled": true, "lineage": false }. DATRIS_UNITY_CATALOG_SYNC=false switches lineage off along with the rest of the push.
Register Iceberg tables
An object store or S3 pipeline withfileFormat: "iceberg" can register its table in Unity Catalog. After each successful write, Datris points Unity Catalog at the table’s current Iceberg metadata file, so Databricks SQL and any Iceberg REST client can find the table by name.
Databricks today. Neither mode currently keeps Databricks Unity Catalog pointing at a table Datris writes to its own S3 prefix. Databricks does not implement the Iceberg REST register call, so the default catalogMode: "register" gets a warning on every run that the catalog has no register endpoint. And when a table is created through Databricks’ Iceberg REST catalog (catalogMode: "rest"), Databricks creates a managed Iceberg table in the schema’s own storage and ignores the location Datris asks for, so Datris refuses to write there and writes by path instead (see Keep the catalog current). Support for Databricks-managed Iceberg tables, where the catalog chooses the location, is planned. Until then, use the Databricks destination for tables you want governed in Unity Catalog. Both modes work with catalogs that implement register and honour the requested location, such as the Apache Iceberg REST test fixture.
credentialsSecretnames a Databricks Platform secret with the same fields the Databricks destination uses:host, plusclientId/clientSecretortoken. It is separate from the object store’s owncredentialsSecret.catalogis the Unity Catalog catalog. It is required and has no default.schemais the Unity Catalog schema. It defaults todefault.- The table is named after the pipeline:
<catalog>.<schema>.<pipeline name>, lowercased, with dots, slashes, spaces and backticks replaced by_. "register": falsekeeps the block but skips registration.
400 when credentialsSecret or catalog is missing, or when the object store writes parquet or orc (only Iceberg tables can be registered).
Databricks prerequisites
These are the grants the planned Databricks managed-table support and the Databricks destination need; today neither Iceberg mode keeps the pointer on a Datris-owned table (see Databricks today above). Registration uses the Unity Catalog Iceberg REST catalog, which needs three things set up once:- External data access enabled on the metastore. A metastore admin turns it on in Catalog Explorer, under the metastore’s settings.
EXTERNAL USE SCHEMAon the target schema for the secret’s identity. Only the catalog owner can grant it:GRANT EXTERNAL USE SCHEMA ON SCHEMA analytics.default TO `<sp-application-id>`.- An external location that covers the table’s S3 path (the bucket and prefix above), with a storage credential that can read it.
EXTERNAL USE SCHEMA for a permission error, or the S3 path that no external location covers.
How runs behave
- The first successful write registers the table.
- Later writes leave Unity Catalog alone. When Unity Catalog still points at an older metadata file of the same table, the run gets a warning that its pointer is behind; Unity Catalog readers see the table as of the first registration until that pointer is refreshed. To keep the pointer current on every run, set
catalogMode: "rest"(see Keep the catalog current). - If a table with that name already exists and points somewhere other than this pipeline’s table, Datris refuses to touch it. It never replaces, merges or drops a table. Rename the pipeline or choose another schema.
- Any failure (a missing grant, a wrong host, a network error) is one warning on the run and an entry in the audit log; the write itself still succeeds.
- The pipeline detail page shows when the table was registered, and whether the pointer is stale or the last attempt failed.
GET /api/v1/pipelines/{name}/unity-catalogreturnsregister(off,never,registered,staleorerror),registeredMetadataLocationandlastRegisterAt. - Deleting a registered pipeline together with its data warns that Unity Catalog still has the table registered at a location that no longer exists; an admin drops it.
DATRIS_UNITY_CATALOG_SYNC=falseswitches registration off for the whole install.
Keep the catalog current (catalogMode: rest)
By default Datris writes the table by path and registers it once (catalogMode: "register", the behaviour above). With catalogMode: "rest", every write goes through the Unity Catalog Iceberg REST catalog instead: creating the table, appends, overwrites, merges and added columns are all committed through the catalog, so Unity Catalog points at the current metadata file after every run and the stale-pointer warning goes away.
catalogMode accepts register or rest (any case) and applies to object store Iceberg pipelines only. The API returns 400 for any other value, for catalogMode on a Databricks destination, and for "register": false together with "rest". rest needs the same credentialsSecret and catalog as registration. To convert an existing non-Iceberg pipeline (parquet, orc), take two steps: first switch fileFormat to iceberg with deleteBeforeWrite on and run it once, then remove deleteBeforeWrite and add catalogMode: "rest".
The first rest run looks at both the catalog and the table already at the pipeline’s path:
- Nothing at either place: the table is created through the catalog at the pipeline’s path.
- A table Datris wrote at the path, and nothing in the catalog: Datris registers the table’s current metadata file once (the run says
adopted), then commits through the catalog. This needs theregistercall, so it is not possible on Databricks: the run is refused and writes by path, and the warning says to start from a new prefix, or to setdeleteBeforeWritefor one run (allowed because the table was never committed through the catalog) so Datris can create the table through the catalog. - The catalog already points at the table’s current metadata file (for example, a table registered with
catalogMode: "register"that has not been written since): Datris commits through the catalog. - The catalog points at an older metadata file of the same table: Datris refuses to use the catalog, because moving the pointer back in line would need the catalog table dropped, and Datris never drops a table. The warning names both files. An admin drops the Unity Catalog table; the next run registers the current file and continues through the catalog.
- The catalog has a table with that name somewhere else: Datris refuses to touch it (never merge, never drop). Rename the pipeline or choose another schema.
deleteBeforeWrite works with rest until the catalog holds a table at the pipeline’s prefix. After that, a run with deleteBeforeWrite fails and names the Unity Catalog table to drop first. A catalog table with the same name somewhere else does not block it: only the pipeline’s own prefix is deleted, and the run is then refused and writes by path.
Reads follow the catalog: querying the pipeline’s object store (the query API and MCP tool that find_data points to) reads a table committed through the catalog at the catalog’s current metadata file. If the catalog cannot be reached, the query reads the table as of the last commit Datris recorded.
There are two kinds of failure:
- Before the commit (a wrong host, a missing grant, a network error while looking the table up): one warning, an audit entry, and the run writes by path as before. The pipeline page shows
refusedwith the reason. - During the commit (the catalog rejects or loses the commit after data files are written): the write fails and so does the run, the same as any failed write.
- If the catalog cannot be reached, or
DATRIS_UNITY_CATALOG_SYNC=false, the run fails instead of writing by path, because a path write would split the table’s history. Fix the catalog connection, or point the pipeline at a new prefix and have an admin drop the Unity Catalog table. deleteBeforeWritefails the run, whatever the file format, because deleting the files would leave the catalog pointing at a table that no longer exists. To start over, have an admin drop the Unity Catalog table, or use a new prefix.- Switching
catalogModefromrestback toregister(or removing it, setting"register": falseor"enabled": false, or removing theunityCatalogblock) is not supported: the next run fails and says so. Keeprest, or start a new prefix and have an admin drop the Unity Catalog table. - Deleting the pipeline together with its data removes the files but not the Unity Catalog table; an admin drops it. The delete response carries a
warningsentry naming the table (Unity Catalog still holds <catalog>.<schema>.<table>; have an admin drop it), and the audit log records it asunity-catalog/pipeline-deletewith outcomewarning. The CLI prints each warning after the success line. - If the catalog created a table for the pipeline at a location Datris refused (a Databricks managed table), deleting or resetting the pipeline warns that Unity Catalog still holds it, created by Datris but never written to; an admin drops it.
- Resetting the pipeline (deleting its data but keeping the configuration) warns that the next run will fail until an admin drops the Unity Catalog table, or the pipeline moves to a new prefix.
- A new pipeline created with the name of a deleted one starts clean: leftover Unity Catalog state from the old pipeline is cleared, unless it guards table files that still exist at the new pipeline’s prefix, or records a table the catalog created for that name at a refused location (so deleting the new pipeline still warns about it). A run that finds “committed” state while neither the catalog nor the prefix has the table ignores the state with a warning and creates the table. On startup, Datris removes Unity Catalog state left by pipelines that no longer exist, except state that guards catalog-committed table files or records a catalog-created table.
- If the pipeline’s Unity Catalog state cannot be read (the config database is unavailable), the run fails rather than risk a path write.
DATRIS_UNITY_CATALOG_SYNC=false writes by path with one line saying the sync is off.
The pipeline detail page shows a rest chip and when the table was last kept current, or refused with the reason. GET /api/v1/pipelines/{name}/unity-catalog adds catalogMode, restMetadataLocation, lastRestCommitAt, restRefusedReason and restCreatedTable, and register gains the values rest and refused.
On Databricks. Confirmed against a live workspace: the service principal’s OAuth client-credentials token works, and Databricks’ Iceberg REST catalog accepts creating a table, but it creates a MANAGED Iceberg table in the schema’s storage (under __unitystorage/) and ignores the requested location. catalogMode: "rest" therefore does not currently work against Databricks Unity Catalog: Datris checks where the catalog put the table before writing anything, refuses to write into the managed location, warns (naming both locations and the table an admin should drop) and writes by path; the pipeline page shows refused. Databricks also has no register call, so an existing path table cannot be adopted there. Supporting Databricks-managed Iceberg tables (letting the catalog choose the location) is a follow-up. catalogMode: "rest" works against catalogs that honour the requested location, such as the Apache Iceberg REST fixture. If a run ever places a table outside the pipeline’s prefix, the same refusal applies with any catalog.
Testing locally. The Iceberg REST catalog test fixture works as a stand-in for Unity Catalog: run the apache/iceberg-rest-fixture image on the Datris network, create its default namespace, and use a secret with host set to http://iceberg-rest:8181, icebergRestPath set to / and icebergRestPrefix set to - (see the next section). Point the pipeline at the built-in MinIO store for this test.
Other Iceberg REST catalogs
For testing and for non-Databricks catalogs, the secret can carry two optional fields that change where the register call goes:host may then include a scheme and port, for example http://iceberg-rest:8181, and a secret without clientId/clientSecret or token sends no credentials. Both registration and catalogMode: "rest" work this way.
Browse what a secret can see
browse_unity_catalog (MCP) and GET /api/v1/unity-catalog/browse (REST) walk Unity Catalog one level per call, using a Databricks Platform secret. No pipeline is needed.
- Find the secret name with
list_platform_secrets. The secret needshostand eitherclientId/clientSecretortoken, the same fields the Databricks destination uses. - Catalog and schema listings are cached for 5 minutes per secret. Table and column listings are always live. Each level returns at most 1000 entries.
- A stopped SQL warehouse starts on the first call. Serverless warehouses take about 5 to 20 seconds; classic warehouses can take several minutes.
- When the secret cannot read the tag views, the response has
tagsAvailable: falseand no tags. Browsing still works.
400 when secret is missing, when schema or table is given without the levels above it, or when no warehouse can be found; 404 when no Databricks Platform secret has that name, or when your API key may not read it (the same response either way, so a key cannot probe for secret names it has no access to; the audit log records the second case as denied); 502 when the warehouse connection or a metadata query fails, with the same message the Databricks destination gives. Every call is recorded in the audit log as unity-catalog / browse with the secret name.
Access: the route needs metadata:read, and your key must also be able to read the named secret, the same check list_platform_secrets applies. Keys without metadata:read do not see the MCP tool.
Which SQL warehouse is used
The secret holds the workspace host and credentials, not a warehouse. Browse picks one in this order:- the
warehousequery parameter (MCP argumentwarehouse) - an optional
warehousefield on the secret (httpPath,http_pathandDATABRICKS_WAREHOUSEalso work). The value can be the warehouse ID, its HTTP path, or a JDBC URL - the warehouse of the first Databricks pipeline whose
credentialsSecretis this secret and that setswarehouseitself (a pipeline that leaves it out uses the secret’s field, so it adds nothing here)
400 listing these three options. Existing Databricks secrets work unchanged; the warehouse field is optional.
Search warehouse tables from find_data
find_data ranks Datris pipelines by default. With include_unity_catalog: true (REST: includeUnityCatalog=true) it also searches the Unity Catalog tables visible to every Databricks Platform secret your key can read. It matches table names, table comments and column names.
- Datris hits come first and carry
source: "datris". Unity Catalog hits follow withsource: "unity-catalog",nameascatalog.schema.table, thesecretthat found it,matchedColumns,tagsand alocationof kinddatabricks. - A table is owned when a Databricks pipeline you can see loads into those exact coordinates, or when its
datris_pipelinetag names a pipeline you can see. Owned tables haveownedByPipeline, and theirhowToQuerypoints atquery_databrickswith that pipeline. Owned tables come before unowned ones. If the owning pipeline is already a Datris hit, the table is left out because the pipeline hit already covers it. - Unowned tables get a plain SQL hint (
SELECT * FROM catalog.schema.table LIMIT 100) to run in your warehouse.query_databrickscannot run it, because that tool needs a pipeline. - The response adds
unityCatalog: { searched: [...], skipped: [{ secret, reason }] }. A secret with no known warehouse (only options 2 and 3 above apply here) or a failed connection is listed underskipped; it never fails the call. With no Databricks secrets,searchedis empty and you get the Datris hits only. limitapplies to Datris hits and Unity Catalog hits separately: you get up tolimitDatris hits plus up tolimitadditional Unity Catalog hits, solimit: 5can return 10 results.countandtotalMatchesinclude the Unity Catalog hits. The optional AI rerank orders Datris hits only.- Each secret has its own time limit, 45 seconds by default (
DATRIS_UNITY_CATALOG_SEARCH_TIMEOUT_SECONDS, see the configuration reference). A slow warehouse is skipped with the reasontimed out after 45srather than blocking the call.
find_data output is unchanged: no source field and no unityCatalog block.
Grants for discovery
Browse and search run as the secret’s identity, so they see only what it is granted:USE CATALOGandUSE SCHEMAon what you want to browse, andSELECT(orBROWSE) to describe tables.- Search and tag lookups read
system.information_schema(columns,tables,table_tags). Unity Catalog filters these views to the objects the identity can access, so a table the secret cannot use never shows up. If the identity cannot read thesystemcatalog, tag lookups returntagsAvailable: falseandfind_datalists the secret underskippedwith the warehouse’s error. Ask a workspace admin to grant the identity access tosystem.information_schema.
Not covered yet
- By default the Unity Catalog pointer is set once and not kept current; set
catalogMode: "rest"to keep it current on every run.registerstays the default. - There is no Unity Catalog browser in the Datris UI yet; use the API or MCP tool.
