Skip to content

Pipelines

Tested with: Databasezy API v1 · zb CLI 0.1

A pipeline copies tables of the project’s primary database to an analytics destination and keeps them up to date: it reads the database’s changes (logical replication), writes them in batches, and shows its state, lag and last error in Platform → Pipelines.

DestinationSettingsCredential (write-only)
ClickHouse instancea ClickHouse instance of the same project, databasenone (the instance’s own login)
ClickHouse serverHTTPS URL of a public server, user, databasepassword
BigQueryGoogle Cloud project, datasetservice account key with BigQuery Data Editor on the dataset
Analytics bucket (Iceberg)coming soon

Each source table becomes <prefix><schema>_<table> with four extra columns: _zb_op (snapshot, insert, update, delete, truncate), _zb_lsn (the change’s position), _zb_deleted and _zb_synced_at. In ClickHouse the table is a ReplacingMergeTree keyed by the primary key: SELECT … FINAL WHERE _zb_deleted = 0 is the current state. In BigQuery the table is a change log: the row with the highest _zb_lsn per key is the current state.

  1. Tables need a primary key (or REPLICA IDENTITY FULL). Tables of platform schemas (auth, storage, …) are not replicated.
  2. The pipeline creates its own publication and replication slot (zb_pipeline_<id>), copies the tables from the slot’s snapshot, then streams every change after it.
  3. Changing the tables or the destination starts a new copy.
  4. Deleting a pipeline drops its slot and publication.

A replication slot keeps the database’s write-ahead log until the pipeline has read it. Every database caps that at a tenth of its volume: a pipeline that falls too far behind (a destination down for a long time) is reset instead of filling the database’s disk, and copies the tables again when it resumes. Paused pipelines release their slot.

Creating and changing pipelines needs the pipelines:write permission and multi-factor authentication. The number of pipelines per organization depends on the plan.