Catalog backup and restore¶
TraceLake ships no catalog and no pointer store of its own. The external Iceberg REST catalog — Lakekeeper by default — holds the atomically-swapped metadata pointer that says what each table currently is. Everything else TraceLake keeps is derivable; this is not.
Lose the catalog's database with no restorable backup and you lose the map to the lake. The Parquet files are still sitting in your object store, intact and unreadable as tables: nothing records which files constitute which table at which snapshot. Reconstructing that by hand from object listings is a forensic exercise, not a recovery procedure.
This page is the HA topology and the backup and restore procedure. A backup nobody has restored is a hypothesis: run the restore before you need it.
TraceLake does not do any of this for you, on purpose¶
There is no TraceLake-side pointer-object CAS store, no shadow copy of the catalog, no "repair from object storage" command. That is deliberate: an Iceberg catalog is a second product with its own HA, auth, backup and engine-compatibility obligations, and a production-grade implementation already exists.
The consequence is a division of responsibility worth stating plainly, because it is easy to assume the resilience lives somewhere it does not: catalog durability is the catalog's problem, and by extension yours. No TraceLake alert, retry or maintenance loop compensates for an unrecoverable catalog. What TraceLake gives you is a system that degrades honestly while the catalog is gone (see the outage table below) and comes back cleanly when it returns.
HA topology¶
┌──────────────────────────────────────┐
ingestor ────────▶│ Lakekeeper Service (ClusterIP) │
compactor ───────▶│ 3 stateless replicas, PDB maxUnav 1 │
gateway ───────▶│ spread across zones │
Trino/Spark ─────▶└───────────────────┬──────────────────┘
(any type=rest) │
▼
┌────────────────────────────────────┐
│ EXTERNAL PostgreSQL │
│ HA primary + standby, PITR/WAL │
│ archiving, tested restore │
└────────────────────────────────────┘
Apply it with the chart's HA overlay, values-ha.yaml:
helm install tracelake <chart> \
-f your-values.yaml \
-f values-ha.yaml \
--set lakekeeper.externalDatabase.host_read=pg-ha.internal \
--set lakekeeper.externalDatabase.host_write=pg-ha.internal
Catalog replicas are stateless. All state is in Postgres, so replicas buy availability across node loss and rolling upgrades — and nothing else. Three replicas over one Postgres is still one Postgres: replication is not backup, and neither is a snapshot you have never restored.
The overlay turns the bundled Postgres off, and that is the point¶
The reference chart bundles Postgres so helm install produces a working system in one command.
values-ha.yaml disables it (lakekeeper.postgresql.enabled: false). The subchart's own values
file says the embedded database is not recommended for production, and the reasons are the ones that
matter here: a single StatefulSet, no failover, no continuous archiving, no PITR, and no backup
schedule. It is a convenience, not a store of record.
Point externalDatabase at either:
- a managed Postgres — Azure Database for PostgreSQL Flexible Server (zone-redundant HA + point-in-time restore), RDS Multi-AZ, or Cloud SQL HA; or
- a Postgres operator with continuous archiving — CloudNativePG or Crunchy, with WAL archiving to object storage and a scheduled base backup.
Whichever you choose, it must give you: synchronous or near-synchronous standby, WAL archiving with a retention window at least as long as your detection time for a logical error, and a restore you have run. The last one is not a formality; see below.
What happens to TraceLake while the catalog is down¶
The catalog being unavailable is a degradation, not a data-loss event, provided it comes back. Per-component posture:
| Component | Behaviour during a catalog outage | What you will see |
|---|---|---|
| Ingestor | Keeps consuming, rotating and uploading. Batched Iceberg appends fail, so ready_* files accumulate on the WAL PVC. When the volume crosses diskHighWatermarkPct, Kafka polling pauses — backpressure to the broker, no records dropped, bounded by broker retention |
TraceLakeUploadBacklog, then TraceLakeWalDiskAboveHighWatermark, then rising TraceLakeConsumerLagHigh |
| Gateway | Keeps serving from its last cached snapshot pointer. Queries succeed against data as of the last successful poll — they go stale, not down, which means new data is silently invisible rather than erroring. Stale serving is per manifest: a query whose partition range needs a manifest this pod has not loaded — an older window, or anything on a pod that restarted during the outage — is 503 CATALOG_UNAVAILABLE / 57P03, never an answer from the loaded manifests alone |
TraceLakeCatalogSnapshotStale (the reason it is critical despite queries still working); 503s on windows older than what was queried before the outage |
| Compactor | An outage at boot fails boot and the pod restarts (the catalog connect precedes the loop); an outage during a pass lets merges run but commits fail, and the daemon logs the failed pass and retries on its next tick — either way the backlog stops draining | TraceLakeCatalogCommitConflicts or TraceLakeCompactorRunsFailing |
| External engines | Trino/Spark/DuckDB resolve tables through the same catalog, so they fail the same way | Their own errors |
Recovery is automatic when the catalog returns: ingestors drain the accumulated ready_* files,
WAL recovery keeps that duplicate-free, and gateways resume polling.
Size broker retention for the outage you intend to survive. The backpressure chain terminates at Kafka: once polling pauses, the only thing preserving unconsumed events is the broker's own retention. A catalog outage longer than that is data loss no restore can undo.
Backup¶
Two things must be backed up together. Backing up only the first is the most likely way to hold a useless backup:
- The Postgres database — every namespace, table, and metadata pointer.
- The secret-encryption key —
LAKEKEEPER__PG_ENCRYPTION_KEY. Lakekeeper stores the warehouse's storage credentials encrypted in Postgres under this key.
The encryption key is the trap¶
The Helm chart generates this key at random on first install (randAlphaNum 40) and keeps it in
a Kubernetes Secret named <release>-lakekeeper-postgres-encryption, annotated
helm.sh/resource-policy: keep. It exists nowhere else. Two consequences:
- A Postgres dump restored without that key is not a recovery. A catalog pointed at a correctly restored database with a different key still resolves the
warehouse and lists tables — the plaintext rows are fine — and then fails
loadTablewith{"error":{"message":"Error fetching secret","type":"SecretFetchError","code":500}}. The catalog looks healthy right up to the moment anything tries to read data. helm templatecannot see the existing Secret. The chart preserves the key across upgrades with alookup, which returns empty during a dry-run or a rendered-manifest workflow — sohelm template | kubectl applyregenerates the key and orphans every encrypted credential. If you deploy from rendered manifests, setlakekeeper.secretBackend.postgres.encryptionKeySecretto a Secret you manage yourself, and treat it like any other root credential.
Back the key up wherever you keep break-glass secrets — not only in the cluster whose loss is the scenario you are preparing for.
What the chart takes for you, and what it does not¶
The bundled Postgres is backed up on a schedule, by default. The chart templates a
catalog-backup CronJob and its PVC whenever the bundled Postgres is enabled (catalogBackup in
the chart values). Nightly at 02:00 by default, it writes a gzipped pg_dump, then prunes — last, so a
night the dump fails is a night nothing is retired.
Retention is deliberately not age alone, because age alone has two ways to eat your backups:
| Guard | Value | The failure it closes |
|---|---|---|
retentionDays |
7 | ordinary ageing |
retentionMinCopies |
3 | a gap in the schedule. After a fortnight of failed runs, one success would delete everything older than retentionDays and leave a single backup standing. Never prune below this many, whatever their age. |
minDumpBytes |
4096 | a successful-but-worthless dump. pg_dump exiting 0 against a database a bad migration just emptied yields a small, valid, newest artifact — and age-based retention prefers the newest. A dump under the floor fails the job and prunes nothing. |
The dump is also gzip -t-verified before anything is deleted, so a truncated artifact cannot pass
as a backup either.
/backup/catalog-20260817T020000Z.sql.gz
The encryption key is not written beside it, and that is the trade this chart makes. A dump without its key is not a recovery (above), so the two must be paired — but pairing them as two files on one volume would make that volume, and every copy anyone ever takes of it, equivalent in sensitivity to the storage credentials the key protects. A backup PVC gets snapshotted, mounted by a debug pod, and rsynced to an archive bucket; a break-glass secret store does not. So the chart keeps the artifact to database contents only, and the pairing is an obligation on you:
kubectl get secret <release>-lakekeeper-postgres-encryption \
-o jsonpath='{.data.encryptionKey}' | base64 -d > catalog-encryption-key
Take that once, keep it wherever your other root credentials live, and re-take it if you ever rotate the Secret. The dump on the PVC restores against it; neither half is a backup alone.
Be precise about what that covers. The artifacts land on a PVC in the same cluster. That is database-level loss — a bad migration, a dropped database, logical corruption — and it is the most likely way you lose the catalog. It does not survive losing the cluster or the zone, which is the scenario the rest of this page is written against: the backup goes with it. An off-cluster copy is yours to arrange, and the break-glass copy of the encryption key stays a separate obligation either way.
With managed Postgres or a PG operator (values-ha.yaml), the CronJob does not render at all —
the guard is on the bundled Postgres existing, not on the flag alone. Point-in-time recovery is the
provider's job there, and a pg_dump from a pod against a database this chart does not own would be
worse than nothing: it would look like coverage.
One value to keep in step. The backup image derives from the bundled Postgres image so
pg_dump matches the server — it refuses to dump a server newer than itself. Helm cannot read a
subchart's own defaults, so the fallback is pinned to what the bundled subchart renders. If you bump the subchart and the Job starts failing on
a version mismatch, set catalogBackup.image.tag explicitly.
Running one now, rather than waiting for the schedule:
kubectl create job --from=cronjob/<release>-catalog-backup catalog-backup-manual
Two alerts watch it when prometheus.prometheusRule.enabled and kube-state-metrics are in play:
TraceLakeCatalogBackupFailing and TraceLakeCatalogBackupStale — the second matters more, because
a schedule that quietly stopped firing looks exactly like one that is working.
The runbook has the response.
Taking the backup¶
Managed Postgres and the operators handle this for you; enable PITR and confirm the retention window. The portable form, for a self-managed database:
# Logical dump — portable across major versions.
pg_dump -U "$PGUSER" -d "$PGDATABASE" --clean --if-exists > catalog-$(date +%F).sql
# And the key it is useless without:
kubectl get secret <release>-lakekeeper-postgres-encryption \
-o jsonpath='{.data.encryptionKey}' | base64 -d > catalog-encryption-key
For anything beyond a small deployment prefer continuous archiving (PITR) over periodic dumps: a nightly dump means a worst case of ~24h of table pointers lost, and the data files those snapshots referenced become orphans that orphan cleanup will eventually delete.
Restore¶
# 1. Stop the catalog. Lakekeeper holds a connection pool; it must let go of the database.
kubectl scale deploy/<release>-lakekeeper --replicas=0
# 2. Restore the database. Restore into a FRESH, empty database, and use ON_ERROR_STOP.
# Without ON_ERROR_STOP, psql prints an ERROR line, continues, and exits 0 — you get a catalog
# that resolves tables and looks healthy while missing a schema object. See below.
psql -U "$PGUSER" -d "$PGDATABASE" -v ON_ERROR_STOP=1 < catalog-2026-08-14.sql
# For a dump that the chart made:
# gunzip -c catalog-<timestamp>.sql.gz | psql -U "$PGUSER" -d "$PGDATABASE" -v ON_ERROR_STOP=1
# (or: managed-provider point-in-time restore / operator recovery from the WAL archive)
# 3. Restore the encryption key, if this is a new cluster.
kubectl create secret generic <release>-lakekeeper-postgres-encryption \
--from-file=encryptionKey=catalog-encryption-key
# 4. Bring the catalog back, to its previous replica count.
kubectl scale deploy/<release>-lakekeeper --replicas=<n>
# 5. VERIFY — do not skip this, a restore that resolves tables can still be broken (see above).
# Load a table, which forces the catalog to decrypt storage credentials:
curl -H "Authorization: Bearer $TOKEN" \
"$CATALOG/catalog/v1/$PREFIX/namespaces/<ns>/tables/<table>"
# Then read rows through an external engine, and confirm TraceLake's own client works.
# Then confirm the restored SCHEMA is the schema you dumped, not just the data:
psql -U "$PGUSER" -d "$PGDATABASE" -Atc \
"SELECT count(*) FROM information_schema.triggers WHERE trigger_schema NOT IN ('pg_catalog','information_schema')"
# Compare against the same query on the source database, run as the SAME role. Both views are
# privilege-filtered, so a different role gives a different count with no loss behind it. Row
# counts and a resolvable table do not see a missing trigger; this does.
If the restore aborts on a statement¶
ON_ERROR_STOP=1 stops the restore at the first statement Postgres refuses. That is the intended
behaviour, and it is why step 2 says fresh database: an aborted restore leaves an incomplete
database that you must drop and recreate before you retry. Do not re-run the dump on top of it.
One such statement is known. Lakekeeper releases before the fix in
lakekeeper/lakekeeper#1781 define this trigger
with a whole-row WHEN (old.* IS DISTINCT FROM new.*) clause, and the namespace table has a
generated column (depth):
ERROR: BEFORE trigger's WHEN condition cannot reference NEW generated columns
DETAIL: A whole-row reference is used and the table contains generated columns.
namespace :: set_updated_at_and_increment_version
pg_dump emits the trigger correctly. PostgreSQL 17 refuses to re-create it against the rebuilt
schema, so the statement is lost on every restore. What it protects is the optimistic-concurrency
version counter on namespace rows, not the metadata pointer — no data is lost, but the restored
catalog stops detecting lost namespace updates, permanently and silently.
The repair is to upgrade Lakekeeper. Migration 20260525120000_fix_namespace_trigger_wholerow_when
rewrites the trigger to an explicit column list. It is idempotent and it converges both fresh and
already-affected databases. Upgrade the catalog, then take a new dump; the new dump restores cleanly.
Do not patch Lakekeeper's schema by hand — a local fork of the catalog schema is a worse problem
than the one it solves.
TraceLake needs no coordination during any of this. Ingestors and gateways reconnect on their own;
the ingestors' accumulated ready_* files drain through WAL recovery without duplicates.
After a point-in-time restore, rewind-related orphans are expected¶
Restoring the catalog to an earlier point makes it forget snapshots that were committed after that
point. The data files those snapshots referenced are still in object storage and are now unreferenced
— genuine orphans. The rows that were committed after the restore point are no longer in the
table. The files are cleaned up by the normal sweep, which only deletes objects older than
maintenance.orphanMinAgeHours (default 24). Do not shorten that window to tidy up faster: the
age guard is what stops the sweep from deleting files a concurrent writer has staged but not yet
committed. Let it run on its own schedule.