Skip to main content
The data platform monitors the {env}-raw MinIO bucket for new files. When a file is dropped into this bucket, the data platform picks it up and processes it automatically. There are three patterns: single data file, metadata file with data file, and bulk upload.

File Naming Convention

Single Data File

Place a file in the MinIO raw bucket following this naming pattern:
Example:

Metadata File + Data File

For more control, upload the data file first, then upload a metadata file (.metadata.json): Data file: stock_price.2026-03-15.pipeline.csv Metadata file: stock_price.2026-03-15.metadata.json
The data file must be uploaded before the metadata file.

Bulk Upload

Upload multiple files to a directory, then trigger processing with a metadata file:
Important: Use a unique dataFilePath directory for each bulk upload. Do not reuse directories.

Compressed Files

Compressed archives (.zip, .gz, .tar, .jar) are automatically decompressed. The pipeline name and publisher token are parsed from the archive filename:

How It Works

  1. A file is placed in a MinIO bucket (via mc cp, API, or another system)
  2. MinIO sends an S3-compatible event notification to the ActiveMQ file-notifier queue
  3. The pipeline polls the queue on a configurable schedule (default: every 5 seconds)
  4. The pipeline parses the event, resolves the pipeline, and processes the file

Configuration

The queue name and polling interval are configured in application.yaml:

Message Deduplication

The pipeline tracks processed message IDs in MongoDB to prevent duplicate processing. Each processed message ID is stored with a TTL (default: 60 days). If the same event arrives again within the TTL window, it is acknowledged and discarded.

MinIO Webhook Setup

Configure MinIO to send bucket notifications. The Docker setup uses the MinioWebhookController endpoint:
Alternatively, ActiveMQ-based notifications can be configured (see MinIO documentation for AMQP notification targets).