Skip to main content
The pipeline includes a built-in MCP (Model Context Protocol) server that lets AI agents interact with the platform natively. Any MCP-compatible agent — Claude Desktop, Claude Code, Cursor, or custom agentic frameworks — can discover database metadata, create and manage pipelines, upload files, monitor jobs, profile data, search vector databases, query structured data, answer questions with AI-powered RAG, and upload configuration files — all without custom integration code. The MCP server is a lightweight Python service that routes all operations through the pipeline’s REST API. It runs alongside the pipeline in Docker or locally for development.

Resources

The MCP server exposes resources that agents can read on demand for detailed documentation.

Available Tools

Pipeline Management

Semantic search across any of the pipeline’s supported vector databases. Each tool takes a natural language query and returns the most similar document chunks with scores and metadata.

Database Queries

Read-only queries against the pipeline’s backend databases.

Metadata Discovery

Explore the structure of PostgreSQL, MongoDB, and vector databases managed by the platform. Use these tools to understand what data is available before writing queries or running searches.

AI

Taps

For a document tap, create_tap requires the target_pipeline to have an unstructuredAttributes source and a vector-store destination (qdrant, pgvector, weaviate, milvus, or chroma). The server rejects mismatched pairings with HTTP 400.

Configuration

Monitoring Active Agents

The Datris UI includes an Agent Monitor tab that shows a live view of every agent currently connected to the MCP server and a streaming log of the tool calls each one is making. Agents are labeled using their MCP clientInfo.name — so Claude Desktop appears as claude-ai, Claude Code as claude-code, Cursor as cursor, and so on — with graceful fallbacks to tenant name, API-key name, or session id for clients that don’t supply one. See Monitoring → Agent Monitor for details.

Setup

Docker (automatic)

The MCP server starts automatically with docker-compose up in SSE mode on port 3000. No additional setup required.

Local (for Claude Desktop / Claude Code)

The MCP server is published on PyPI. Use uvx to run it directly:

Transport Modes

Configuring Claude

Client-side setup for Claude Desktop and Claude Code — including recommended and alternative transports, and example first prompts — lives on its own page: Configuring Claude.

Environment Variables

All database connections, vector search, and embedding are handled by the pipeline server. The MCP server only needs the pipeline URL.

Authentication

The MCP server has no API key of its own. Each connecting agent sends an x-api-key header per session and the MCP server forwards it as-is to the Datris REST API on every tool call. Issue and manage agent keys from Configuration → API-Keys in the UI. When the Datris platform has USE_API_KEYS=true, set REQUIRE_API_KEY=true on the MCP server too — otherwise the MCP layer accepts unauthenticated connections that then fail with 401 at the REST tier. The pairing is intentional: turn both on, or neither. Each agent should connect with its own labeled key (e.g. claude-desktop, cursor, prod-agent) so its traffic appears under a distinct identity in the request logs and can be revoked independently. Issue from the UI:
  1. Configuration → API-Keys → Issue new key
  2. Label the key for that agent (claude-desktop, cursor, etc.)
  3. Pick a capability template (read-only, rag-builder, ops) or build a custom scope
  4. Copy the value once at issue time and paste it into the agent’s MCP client config
The platform’s capability framework enforces what each key is allowed to do — a read-only agent can’t create_tap even if its system prompt tells it to try. See API Keys for the full per-client pattern.

Example Agent Workflows

The platform’s canonical workflow is delivered to every connected agent through the MCP instructions field — agents don’t need to memorize it. These examples show the same flow applied to common tasks.Core rules the agent receives on connect: check what exists before creating (list_pipelines, list_taps); keep pipeline configs simple (source + destination only — never pass profile_data output into a pipeline config); after run_tap, read persisted and persistedReason before doing anything else; verify real completion via get_pipeline_status(publisher_token=...), not by the response body; set cron_expression on the tap when the user describes a recurrence — never propose external schedulers or shell snippets; call test_tap before run_tap and before setting a cron on any newly-created or newly-updated script; and only narrate work that actually completed in the current turn — if the tool calls didn’t fire, say so honestly rather than confabulating success; a repeated request (“run it again”) is a new operation requiring a fresh tool call, never a result forecast from the previous run.For long-form tap workflow guidance — params, scheduling, validation, error handling, the Quartz CRON cookbook — agents can re-read the datris://tap-workflow-reference resource on demand at any point in a session.

Ingest a CSV file directly

  1. Check for an existing pipelinelist_pipelines. If one already fits the data, skip to step 3.
  2. Create a pipelinecreate_pipeline with the CSV as sample data. Schema is auto-detected. Keep the config simple; only pass codegen_rule / codegen_transform if the user explicitly asks for validation or transformation.
  3. Uploadupload_data with the CSV content (base64).
  4. Verify it landedget_job_status with the pipelineToken returned from upload_data. Poll until rollup.allDone is true, then read rollup.status (success, warning, error). On failure, rollup.jobs[].lastError carries the failing process and message.

Onboard a new external data source via a tap

  1. Check existing tapslist_taps. If a tap already covers this source, run or test it directly.
  2. Create credentials — if the source needs an API key, call create_tap_secret with the credential fields (they become env vars inside the script).
  3. Create the tapcreate_tap with an instruction (AI generates the script) or with your own Python fetch() function (faster and more reliable). Pass secret_name to bind the credentials from step 2. For PDFs/Word/HTML into a vector-store pipeline, pass tap_type="document".
  4. Testtest_tap. If it fails, read the error, fix the script, and call create_tap again. Repeat until the test passes.
  5. Runrun_tap. Read the response’s persisted field first:
    • persisted: true → capture publisherToken and continue to step 6.
    • persisted: false → read persistedReason (no_target_pipeline, test_mode, run_error, no_records, debounced), tell the user exactly why, and stop. (run_tap does not return records — call test_tap if you need to preview what the script produces.) debounced means the same tap was triggered within the last 5 seconds; the earlier run is still executing — do not retry, look it up via get_tap_logs and poll get_pipeline_status instead.
  6. Verify it landedget_pipeline_status(publisher_token=...) and poll until rollup.allDone is true. Then rollup.status is the outcome (success, warning, error) and rollup.jobs[].lastError carries any failure detail. The tap log only tells you the script ran; the publisher token is how you confirm records actually reached the destination.
  7. Schedule (optional)update_tap with a cron_expression. For each scheduled run, pick the corresponding entry from get_tap_logs to get its publisherToken, then verify with step 6.

Build and query a RAG knowledge base from external documents

  1. Create a vector-store pipelinecreate_pipeline with a Qdrant / Weaviate / Milvus / pgvector / Chroma destination and an unstructuredAttributes source.
  2. Create a document tapcreate_tap with tap_type="document" and an instruction describing the document source (e.g., “ingest every PDF under legal-contracts/2026/ in S3”). If the source needs credentials, call create_tap_secret first.
  3. Test and runtest_tap, then run_tap. Read persisted / persistedReason exactly as above.
  4. Verify every document landedget_pipeline_status(publisher_token=...). Document taps fan out to N jobs (one per file) under one publisher token; poll until rollup.allDone is true. rollup.jobs[] has one entry per file with its own status and lastError.
  5. Inspect the ledger (optional)get_tap_ledger shows which files were discovered and processed. On subsequent runs, unchanged files are skipped automatically; pass clear_uri or clear_all=true to force a re-scan.
  6. Searchsearch_qdrant (or the matching tool for your destination) to retrieve relevant chunks.
  7. Answerai_answer with the retrieved chunks as context and the user’s question.

Discover and query existing data

  1. List databaseslist_postgres_databases.
  2. List schemaslist_postgres_schemas for the target database.
  3. List tableslist_postgres_tables (supports a vector-only filter).
  4. Inspect columnslist_postgres_columns to understand structure.
  5. Queryquery_postgres with a SELECT, or query_natural to ask a plain-English question and have the AI generate + run the SQL.

Cross-modal analysis (structured + vector)

  1. Vector searchsearch_pgvector (or another vector tool) for relevant document chunks.
  2. Structured queryquery_postgres or query_natural for related metrics.
  3. Combine — merge the two result sets in the response.

Automated quality monitoring of scheduled taps

  1. List tapslist_taps to find the tap of interest.
  2. Read recent runsget_tap_logs returns the last 50 entries with status, record count, duration, errors, and publisherToken for each run that submitted records.
  3. Verify completion — for any run of interest, call get_pipeline_status(publisher_token=...) and read rollup.status to confirm records actually landed in the destination. The tap log tells you the script ran; the publisher token tells you the load finished.
  4. Diagnose failures — on an errored job, read rollup.jobs[].lastError (processName + description) and either fix the tap (create_tap again with a corrected script) or retarget the pipeline (update_tap).

CLI Examples

The Datris CLI connects to the MCP server and provides the same capabilities from the terminal. See CLI for the full reference.