vector column.
pgvector uses your existing PostgreSQL infrastructure — no separate vector database server required. Standard SQL can be used to combine vector similarity search with traditional filters.
Configuration
Supported File Types
For structured data (CSV, JSON, XML), use the standard PostgreSQL database destination instead.
Vault Secrets
PostgreSQL connection (for pgvector)
This is separate from the pipeline’s standard
oss/postgres secret so the vector store can target a different database or server.
Embedding API
The embedding secret is server-level (ai.embedding.secretName, default oss/embedding) and is seeded automatically by docker/vault-init.sh. See AI Configuration and the Qdrant docs for the full picture.
bge-m3 and vault-init.sh seeds the embedding secret to point at it — no OpenAI key required.
Chunking Strategies
Documents are split into chunks before embedding. Each chunk becomes a row in the PostgreSQL table with the document’s metadata columns pluschunk_index, filename, and source_pipeline.
chunkSize(default 500): maximum characters per chunkchunkOverlap(default 50): characters of overlap between consecutive chunks
Metadata
Static metadata is stored as dedicated columns in the PostgreSQL table:id— deterministic UUID (idempotent upserts)text— the chunk textchunk_index— position of the chunk in the documentfilename— original uploaded filenamesource_pipeline— pipeline nameembedding— vector column for similarity search
How It Works
- Upload — an unstructured file is uploaded via
POST /api/v1/pipeline/upload - Extract — text is extracted from the document
- Chunk — text is split into chunks using the configured strategy
- Embed — each chunk is sent to the embedding API to generate a vector
- Upsert — chunks are upserted into PostgreSQL with
INSERT ... ON CONFLICT DO UPDATE - Notify — a pipeline notification is published on completion
Running PostgreSQL with pgvector
The standard PostgreSQL Docker image does not include pgvector. Use the pgvector image:Verifying
Advantages Over Dedicated Vector Databases
- No separate server — uses your existing PostgreSQL infrastructure
- Standard SQL — combine vector search with traditional WHERE clauses, JOINs, aggregations
- ACID transactions — full transactional guarantees on vector data
- Familiar tooling — use psql, pgAdmin, any PostgreSQL client
- No new dependencies — uses the existing PostgreSQL JDBC driver
