Runbook¶
Every alert in the TraceLake alert rules carries a runbook_url pointing at a heading on
this page. If you arrived from an alert, the anchor already put you in the right section.
Each entry is Symptom → Check → Fix → Verify. Nothing here is automated; orphan cleanup and WAL backpressure are the only self-correcting paths in the system, and both are called out where they apply.
The commands on this page use ghcr.io/tracelake/tracelake:<tag> as an example image and
<release> for the name of the Helm release. Use the image and the names of your deployment.
Before anything else, two orientation facts that decide half of these:
- TraceLake's freshness is minutes, not seconds (target: P95 ≤ 7 min; the configured waits add up to 343 s at defaults, which is not an upper limit). A gap of a few minutes between "produced" and "queryable" is the design, not an incident. Do not diagnose against a sub-minute expectation.
- Iceberg in your own bucket is the only copy of the data. The gateway, the sidecars and the caches are disposable acceleration. Almost everything below is a latency or visibility problem; the three that can actually lose data are flagged DATA LOSS RISK, and WAL segment overfilled is the one that corrupts rather than loses.
Alert → runbook¶
| Alert | Section |
|---|---|
TraceLakeConsumerLagHigh |
Consumer lag |
TraceLakeCatalogBackupFailing |
Catalog backup failing or stale |
TraceLakeCatalogBackupStale |
Catalog backup failing or stale |
TraceLakeWalDiskAboveHighWatermark |
WAL disk high watermark |
TraceLakeUploadBacklog |
Upload backlog |
TraceLakeUploadErrors |
Upload backlog |
TraceLakeConsumeErrors |
Consume errors |
TraceLakeReadyCleanupSlow |
Ready cleanup latency |
TraceLakeFreshnessAboveTarget |
Freshness above target |
TraceLakeRecordsDropped |
Records dropped |
TraceLakeWalSegmentOverfilled |
WAL segment overfilled |
TraceLakeCatalogSnapshotStale |
Catalog snapshot stale |
TraceLakeJwksCacheStale |
JWKS staleness |
TraceLakeIndexCacheHitRatioLow |
Index cache hit ratio |
TraceLakeSidecarMissRatioHigh |
Sidecar miss ratio |
TraceLakeManifestCacheThrash |
Manifest cache thrash |
TraceLakePuffinReadAmplification |
Puffin read amplification |
TraceLakeIdxPruningIneffective |
Idx pruning ineffective |
TraceLakePostgresTableResolveFailures |
Postgres resolve failures |
TraceLakeCatalogCommitConflicts |
Commit conflicts |
TraceLakeCompactionBacklog |
Compaction backlog |
TraceLakeIndexBackfillNotDraining |
Index backfill not draining |
TraceLakeCompactorRunsFailing |
Compactor failure |
TraceLakeCompactorScrapeDown |
Compactor failure |
TraceLakeCompactorDeadlineExceeded |
Compactor deadline |
TraceLakeCompactorEphemeralHigh |
Compactor ephemeral disk |
TraceLakeOrphanDeletionSpike |
Orphan deletion spike |
TraceLakeWalBufferSaturated |
WAL buffer saturated |
TraceLakeSchemaNameDrift |
Schema name drift |
TraceLakeBackgroundTaskDied |
Background task died |
TraceLakeTaskWedged |
Background task wedged |
TraceLakeMaintenanceTaskWedged |
Background task wedged |
Scenario 1 — the catalog is down¶
The external Iceberg REST catalog (Lakekeeper by default) holds the metadata pointer that says what each table currently is. It is the only non-derivable state in a deployment.
Every component degrades differently, and only one of them errors. That asymmetry is the whole diagnosis: a catalog outage looks like "queries are fine, but new data never appears".
| Component | What it does | What you see |
|---|---|---|
| Ingestor | Keeps consuming and uploading; appends fail, so ready_* files pile up until the WAL high watermark pauses Kafka polling |
TraceLakeUploadBacklog → TraceLakeWalDiskAboveHighWatermark → TraceLakeConsumerLagHigh |
| Gateway | Serves the last cached snapshot pointer. Queries over windows it has loaded succeed, against older data; a window whose manifests it never loaded is a 503 | TraceLakeCatalogSnapshotStale |
| Compactor | Merges may run, commits fail | TraceLakeCompactorRunsFailing |
| Trino/Spark/DuckDB | Resolve through the same catalog, fail the same way | Their own errors |
Catalog backup and restore is the HA topology and the restore procedure. This section is the response.
Catalog snapshot stale¶
Symptom — tracelake_catalog_snapshot_age_seconds above 60s for 10 minutes. Queries over
recently read windows still return 200. That is why this alert is critical despite nothing
appearing broken: new data is invisible and nothing errors. Queries reaching further back than this
pod has loaded return 503 CATALOG_UNAVAILABLE (57P03) until the catalog returns — on a pod that restarted during the outage, that is every query.
Check
# 1. Is the catalog answering at all? Any authenticated GET will do.
curl -s -o /dev/null -w '%{http_code}\n' \
-H "Authorization: Bearer $TOKEN" "$CATALOG/catalog/v1/config?warehouse=$WAREHOUSE"
000 with a curl exit of 7 is the catalog being down (connection refused). 401/403 is the
catalog being up and rejecting TraceLake's credentials — a different problem, see the note below.
200 means the catalog is healthy and the staleness is on the gateway's side: check the gateway's
own logs and queryGateway.snapshotPollIntervalSeconds.
# 2. Is this a catalog outage or an identity-provider outage? They page differently but a
# dead IdP takes the catalog's auth down with it, which reads as a catalog outage.
curl -s -o /dev/null -w '%{http_code}\n' -X POST "$ISSUER/token" \
-d grant_type=client_credentials -d client_id="$CLIENT_ID" -d client_secret="$CLIENT_SECRET"
Provider up + catalog down → this section. Provider down → JWKS staleness; the gateway's query path fails independently of anything the catalog is doing.
Fix — restore the catalog. It is an external service; you are following its runbook, not
TraceLake's. Catalog backup and restore covers the database restore, including the trap that a Postgres
dump restored without LAKEKEEPER__PG_ENCRYPTION_KEY resolves tables and then fails loadTable.
TraceLake needs no coordination during any of this. Do not restart ingestors or gateways to "help" — a restarted ingestor re-runs WAL recovery against a catalog that is still down, and gets you nothing.
Verify — the age gauge falls on its own within one poll interval
(queryGateway.snapshotPollIntervalSeconds, default 8s), with no restart and no manual step. Then
confirm the write path drained: tracelake_ingestor_wal_ready_files returns to ~0 and consumer lag
falls. WAL recovery makes that drain duplicate-free — a ready_* file whose committed
offset is at or below its start range is deleted, not re-appended.
If the outage outlasts Kafka retention, that is data loss no restore undoes. The backpressure chain terminates at the broker: once polling pauses, the only thing preserving unconsumed events is the broker's own retention. Size it for the outage you intend to survive.
Scenario 2 — WAL disk backpressure and the upload path¶
The model: rotated ready_* files wait on the WAL volume until they upload and commit. If the
downstream stalls they accumulate, and at ingestion.wal.diskHighWatermarkPct (default 80) the
ingestor pauses Kafka polling rather than filling the disk.
Read these four alerts in dependency order, not in the order they fired. Polling pauses at the watermark, so a full WAL volume causes consumer lag. Chasing the lag first wastes the outage.
upload/commit stalls → ready_* files accumulate → WAL disk crosses the watermark
→ polling pauses → consumer lag rises
WAL disk high watermark¶
Symptom — tracelake_ingestor_wal_disk_bytes past the threshold for 5 minutes. Kafka polling is
paused right now. No data is being lost yet; the broker is holding it.
Check — the drain is upload → commit, so the stall is in one of them:
rate(tracelake_ingestor_upload_errors_total[15m])— object-store failures.- Catalog reachability — see Catalog snapshot stale. A catalog outage is the most common cause of a WAL that will not drain.
tracelake_ingestor_wal_ready_filesper dataset — which dataset is stuck.
Fix — clear the downstream stall. The volume drains on its own once uploads and commits succeed; polling resumes automatically at the low-water mark (the hysteresis is deliberate, so the ingestor does not flap at the boundary).
Expanding the PVC is the second resort, not the first — it buys time and does not fix the stall. If
you do expand it, remember the alert threshold is a percentage of the volume, so the chart-computed
PrometheusRule threshold must be re-rendered (ingestor.wal.size) or it now fires at the wrong point.
Never delete ready_* files to reclaim space. DATA LOSS RISK: those are rotated, uploaded-or-
not, uncommitted rows, and their filenames carry the provenance
(ready_{topic}_{partition}_{start}-{end}_{uuid}.parquet) that recovery reconciles against the
committed offset. Deleting one drops every row in it.
Verify — wal_disk_bytes falls below the watermark, wal_ready_files returns to ~0, and
consumer lag starts falling on its own.
WAL buffer saturated¶
Symptom — tracelake_ingestor_wal_buffered_bytes above 80% of ingestion.wal.maxBufferedBytes
for 30 minutes (TraceLakeWalBufferSaturated).
Nothing is failing, and nothing is at risk of being lost. The budget is doing its job: memory is
bounded, and what the pod is paying instead is flush batch size. The ingestor buffers up to 1000
decoded records per open WAL segment before binding them into one Arrow batch; the budget caps the
sum of those buffers across the dataset, so at a high segment count each flush carries proportionally
fewer rows, and per-record bind cost rises (1.4–1.7 µs at 1000 rows, 4.1–4.3 µs at one — measured on
two dataset shapes). Rows per flush at saturation ≈
maxBufferedBytes / (open segments × bytes per record), where the gauge's bytes per record measured
1,698 B and 2,557 B on those two shapes.
The gauge under-reports real memory, on purpose and in a knowable direction. It counts owned
bytes and cannot see allocator slack or hash-table overhead; measured against RSS it came to 46–57%.
So a pod at a 256 MiB budget is holding roughly 0.5–0.6 GB resident in this buffer. Use that
factor whenever you compare the gauge with container_memory_working_set_bytes.
Check — this is always about segment count, which is (Kafka partition × Iceberg partition
tuple) pairs:
tracelake_ingestor_wal_active_filesper dataset — how many segments are open. Below ~105–158 (the measured range across two dataset shapes) the budget cannot bind at the default, so a saturated gauge there means the records are unusually large rather than the segments unusually many. Dividewal_buffered_bytesbywal_active_files× 1000 to get this dataset's real bytes per record and settle which it is.- The dataset's
partitionMappingcardinality. A key with many distinct values that recur — an hour bucket at high fan-out, a tenant id, a node name — multiplies segments per Kafka partition up tomaxOpenSegmentsPerPartition(256 by default). A key that never repeats (a raw UUID) is a different problem: it leaves ~1 row per segment, so it will not saturate this at all, and what it costs shows intracelake_ingestor_wal_rowgroup_bytesand file count instead. tracelake_ingestor_wal_rowgroup_bytesbeside it. That is the other half of the same pod's ingest memory — the in-flight Parquet row groups — and it is not bounded by anything here. If an ingestor is near its memory limit, read both gauges before concluding which one to act on.
Fix, in order of preference:
- Reduce the partitionMapping's cardinality. This is the real cause almost every time, and it also fixes the small-file pressure the same config creates downstream. The guidance holds: a partition key should have a bounded, modest number of live values.
- Lower
maxOpenSegmentsPerPartitionfor the dataset. Fewer segments open at once means larger buffers each. It rotates more often — more, smallerready_*files — so weigh it against upload and compaction load. - Raise
ingestion.wal.maxBufferedBytes, if the pod has the headroom for it. Budget ~2× the value against the ingestor's memory limit (the chart ships 4 GiB), because the gauge is a floor on RSS — and remember the in-flight row groups are drawing on that same limit with no bound at all. If you raise it, the chart derives this alert's threshold from the key, so re-render the PrometheusRule or the alert now fires at the wrong point.
Do not treat this as backpressure or restart the pod. Backpressure is a WAL disk watermark and pauses polling; this budget never pauses anything, and a restart rebuilds exactly the same segments from the same config.
Verify — wal_buffered_bytes settles well below the budget, wal_active_files drops to the
expected segment count, and tracelake_ingestor_bytes_total's rate is unchanged or better.
Upload backlog¶
Symptom — tracelake_ingestor_wal_ready_files > 0 for 30 minutes
(TraceLakeUploadBacklog), or a non-zero tracelake_ingestor_upload_errors_total rate
(TraceLakeUploadErrors). Both route here because they are two views of one stall.
A brief non-zero count is normal — a file is rotated and then uploaded. Sustained means the upload or the commit is failing.
Check
- Ingestor logs for the object-store error. Credential expiry, a network policy change, and a storage-account throttle all look the same in the metric and different in the log.
config.storage.destinationUriandstorage.propertiesstill valid — a rotated key that nobody updated on the ingestor is the classic cause.- If uploads succeed but the count stays high, the commit is failing: check catalog reachability.
Fix — restore whichever leg failed. No TraceLake-side action drains a backlog faster; the ingestor retries continuously.
Verify — wal_ready_files returns to ~0 and the error rate returns to zero. Confirm freshness
recovers too: a drained backlog with a still-high P95 means the stall moved rather than cleared.
Consume errors¶
Symptom — any increase in tracelake_ingestor_consume_errors_total over an hour. This is the
broker side, not the storage side.
Check — ingestor logs for the librdkafka error: broker unreachable, SASL/TLS failure, topic
missing, or a partition rebalance loop. Check whether the topic in
config.ingestion.datasets[].topic still exists and the groupId is not being fought over by
another consumer.
Fix — broker-side. TraceLake does not compensate for a broker that rejects it.
Verify — the counter stops increasing and consumer lag falls.
Consumer lag¶
Symptom — tracelake_ingestor_consumer_lag over 1M on a partition for 30 minutes.
Check the WAL disk panel first. If polling is paused at the high watermark, the lag is a symptom of WAL disk high watermark and nothing about the consumer is wrong. Only if WAL disk is healthy is this genuine under-provisioning.
A gap in this series is not zero lag. A partition disappears from the panel whenever a refresh tick obtains no reading for it. That is deliberate — the series used to freeze at its last healthy value, which suppressed this alert — so a gap is a signal to read, not a quiet period. Which cause it is comes from the logs:
kubectl logs <pod> --since=1h | grep -E "committed\(\) failed|fetch_watermarks failed"
A hit means the broker or the group coordinator is unreachable. Treat it as an outage, not as healthy lag: real lag is still accruing behind the gap and nothing here can measure it.
If that grep is empty, the gap is one of the causes that does not warn!. In order of
likelihood:
- The partition has no committed offset yet — there is nothing to compute lag against, so it is
omitted. Normal for the first few minutes after a pod joins a group. It logs at
debug!(no committed offset yet), deliberately notwarn!, because at every pod start it would fire for every partition every 15s. A gap that persists here means the WAL/upload path has never completed a rotation on that partition — see Upload backlog; that pod is consuming and never committing, which is exactly the silent case this panel cannot show you. - The partition was revoked to another pod. Confirm on the other replica's series — a rebalance moves the line, it does not delete it fleet-wide.
- The pod holds no assignment (idle, or mid-rebalance). Not a fault.
And one that is not a gap at all, listed because it is the other way this panel misleads:
- The refresh loop is wedged. A wedge stops the loop before it can remove anything, so the line freezes at its last value rather than breaking. See Background task wedged.
Fix (genuine case) — the ingestor is a StatefulSet, one Kafka consumer group per dataset. Add replicas up to the partition count; beyond that, partitions are the ceiling and the topic needs more. Each replica gets its own WAL PVC, which is why it is a StatefulSet and not a Deployment.
Verify — lag falls steadily. Freshness P95 will lag behind the recovery by a few minutes; that is expected.
Ready cleanup latency¶
Symptom — P95 of tracelake_ingestor_ready_cleanup_seconds over 1s for 30 minutes.
This is local post-upload cleanup. Slow cleanup holds WAL disk the backpressure model assumes is released promptly, so it is an early warning for the watermark rather than a problem in itself.
Check — the WAL volume's storage class and its IOPS. A network-attached volume under an unlink storm behaves very differently from local SSD. Check for another process sharing the volume.
Fix — faster storage for ingestion.wal.localDirectory, or fewer/larger WAL files
(wal.maxFileSizeBytes, within the 16MB–256MB range).
Verify — P95 back under 1s and wal_disk_bytes stable.
Freshness above target¶
Symptom — P95 of tracelake_freshness_seconds over 420s for 30 minutes. The published
target is breached: data is taking more than 7 minutes to become queryable.
Check — the budget is a sum of four stages, and the metric is the total. Attribute it before acting:
| Stage | Config | Diagnostic |
|---|---|---|
| Rotation | wal.maxFileAgeMinutes (5) |
is the dataset low-volume, so files age out rather than fill? |
| Upload | — | upload_latency_seconds P95, upload_errors_total |
| Commit batching | ingestion.commitIntervalSeconds (20) |
catalog commit latency |
| Gateway poll | queryGateway.snapshotPollIntervalSeconds (8) |
catalog_snapshot_age_seconds |
The four stages above sum to 343 s at their defaults, so a 300s threshold would page on a healthy
trickle dataset that is doing exactly what it is configured to do. 420s is the published target
and sits 77s above the defaults sum — that headroom is why the threshold is 420s and not 300s. In one
measurement the 95th percentile at the defaults was 382s, so the real margin was about 38s. A
per-dataset wal.maxFileAgeMinutes override changes that arithmetic for that dataset only.
Fix — whichever stage the table above implicates. Lowering maxFileAgeMinutes improves freshness
by very close to the rotation delta it removes: measured P95 fell from 382s at 5 minutes to 173s at 2
and 102s at 1, tracking the arithmetic above to within the probe's resolution.
It costs files, and how many depends on which regime you are in. Both regimes are now measured:
- Under sustained load the key is inert: a segment hits
maxFileSizeByteslong before its age threshold, and 2-minute and 1-minute arms at the same size produced 132 and 120 files. If files really are piling up under load, look atmaxFileSizeBytes, not here. - On the low-volume side — the case this runbook's table points you at — it costs close to the
full arithmetic. Three hours of steady arrival onto one segment per dataset, size threshold
unreachable, identical rows in every arm: 37 files at 5 minutes, 89 at 2, 175 at 1 — 2.41× and
4.73×. Expect
segments × window / age, and budget for it before you lower the key.
Do not check whether it worked by counting live files. In that run the 1-minute arm crossed
compaction.backlogFileThreshold and the trigger merged 115 of its files into one, so it ended
with 61 live against the 2-minute arm's 89 — the arm that wrote nearly five times the files holding
the fewest. What the key costs is files written; what you see later is that minus whatever the
compactor has caught up on.
So: if the dataset is low-volume and you have a file-count budget, lower this key and watch
tracelake_compactor_partition_files_total for that dataset — bearing in mind it is a small-file
gauge refreshed on the backlog tick, so it lags and undercounts. The backlog trigger bounds the count
once it fires: in that run it cleared the 1-minute arm one tick after it crossed 100.
Below the threshold it does nothing at all — the 2-minute arm sat at 89 candidate files and
logged no compaction backlog on every tick for three hours.
Do not wait for the nightly floor pass to clear them. The size-gain guard defers any block
whose candidate files cannot sum past compaction.minFileSizeMb — those 89 files total 1.03 MB
against a 32 MB floor — and it defers on the floor pass too. The guard has two exemptions: the
backlog bar itself, and the index-backfill rule, which ignores size but only applies to a dataset
that configures tantivy.indexingColumns. So for a low-volume dataset with no indexing columns, what
clears the partition is the file count reaching backlogFileThreshold, or enough bytes arriving to
pass minFileSizeMb — nothing else, and a partition that stops receiving data keeps its files.
A manual ingestor compact --config <file> --dataset <name> --date-range <day>:<day> does not override this
either: it routes through the same block enumeration and hits the same guard. Run it with
--dry-run to see which way the guard falls — a deferred partition logs partition not compacted:
its candidates cannot sum past minFileSizeMb and are below backlogFileThreshold. This is the guard
working as specified, not a stuck compactor: merging 89 files into one 1 MB file would only produce
another small file for the same rule to re-elect.
wal.maxFileSizeBytes is the knob with the costs, and if you lower it, know them first. Lowering it to 16 MB cost ~3.4× the data files, ~2.4× the
object-store PUTs per ingested MB, and ingest throughput — 40.7–46.8 MB/s of a 50 MB/s offered rate
against 49.5 at the 64 MB default, at fewer cores, because the extra rotation and upload work
serialises in the per-dataset pipeline rather than spending CPU (so more cores do not recover it).
There is no single percentage for the throughput cost: two runs that should have agreed differed by
15%.
The other correction: files landing on the compactor is the good outcome, not the cost. At 16 MB
the backlog trigger fired and merged 132 files in 38s, leaving a partition with full Tantivy .idx
coverage; at the 64 MB default the same load left 36 files that are never small-file candidates and
carry no .idx at all — which is what the boot warning at maxFileSizeBytes >= minFileSizeMb is
telling you.
Verify — P95 back under 420s.
Records dropped¶
Symptom — a non-zero rate on tracelake_ingestor_records_dropped_total.
DATA LOSS RISK — this is the alert that means data is already gone. Payloads failed to decode and the drain policy dropped them. They are still in Kafka until retention expires, which is the window you have.
Check — ingestor logs name the decode failure. In practice this is almost always a producer-side
schema change nobody announced: a field that changed type, or a JMESPath in
datasets[].customMapping that no longer resolves.
Fix
- Find what changed on the producer.
- Update
customMapping/parquetSchemato match, and roll the ingestor. - If the events are still inside broker retention, reset the consumer group to before the change to re-ingest what was dropped. Weigh this against duplicates: re-consuming a range that partially committed produces duplicate rows, which TraceLake does not deduplicate.
Note that a field the schema does not know about is not a drop — it lands in the fallback column. A drop is a payload that could not be decoded at all.
Verify — the drop rate returns to zero and the event rate is unchanged (a drop rate that falls because the producer stopped is not a fix).
WAL segment overfilled¶
Symptom — TraceLakeWalSegmentOverfilled: any increase on
tracelake_ingestor_wal_segment_overfilled_total.
This is silent corruption in the store of record, and it has already happened. Not data loss —
nothing is missing — but duplicate rows, which no reader can distinguish from real ones. A WAL
segment covering Kafka offsets start..=end can hold at most end - start + 1 rows. Finalizing one
with more rows than that is arithmetically impossible without the same offsets being appended twice,
so the ready_* file on disk already carries duplicates — and Iceberg in your own bucket is the
only copy of the data. There is no threshold and no false positive to tune here: the check is an
exact detector, not a heuristic, and a reading below the span means nothing (offset gaps are
legitimate).
Nothing about this is fixed by restarting or rolling the pod. The rows are written. Finalization proceeds on purpose: refusing to close the file would strand rows that are real, and WAL recovery reconciles on the offset bounds, which stay truthful. A restart neither removes the duplicates nor stops the next one.
Check — the condition is logged at error with everything needed to
identify the file:
kubectl logs -l app.kubernetes.io/component=ingestor --since=2h \
| grep "holds more rows than its offset range"
# dataset, partition, start_offset, end_offset, rows, span — one line per affected segment
The affected object is the ready_* file named for those bounds (its name carries the provenance:
ready_{topic}_{partition}_{start}-{end}_{uuid}.parquet), and by the time you are reading this it
has almost certainly been uploaded and committed. Map the offset bounds to the dataset's partition to
find which Iceberg data files received them.
Fix
- Record the
dataset/partition/start_offset/end_offsetset from every log line. That is the blast radius; nothing else in the deployment narrows it. - De-duplicate downstream, against those partitions. The table is an open Iceberg table in your
own bucket, reachable by any
type=restengine — do it from Trino or Spark through the shared catalog. TraceLake itself does not deduplicate on read and ships no delete-file writer. - Treat the cause as a defect, not a configuration problem. The append path's ownership-generation
guard (
tracelake_ingestor_wal_duplicate_records_total) exists to make this impossible, so an increase here means records reached the WAL append twice by a path that guard does not cover. Capture the logs and both counters' values before the pod is replaced.
Verify — the counter stops increasing (the alert clears an hour after the last step, since
increase[1h] is what ages it out), and a row count over the affected partitions matches the
broker's own count for the same offset range.
Schema name drift¶
Symptom — tracelake_ingestor_schema_name_drift above 0 for a dataset for an hour, alongside a
boot WARN naming a config name, a live name and a field-id.
Not an outage, and not urgent. Someone renamed a column in the live table (a non-breaking change,
no rewrite) and the ingestor's parquetSchema still declares the old name. The ingestor boots on
purpose: identity is the field-id, which the rename does not change, so no written tuple can
misalign and every query keeps working. Do not "fix" this by renaming the column back.
What is actually wrong — until config catches up, every data file this pod writes carries the
old column name against the right field-id. Anything resolving by field-id (TraceLake's gateway,
Trino, Spark) is unaffected. Anything reading the raw Parquet by name — an ad-hoc pandas read, a
hand-written column projection — sees the pre-rename name on new files and the post-rename name in
the table metadata, with nothing to explain the difference.
Check — the WARN names all three facts. Confirm against the table:
# the live name, from the catalog
curl -s -H "Authorization: Bearer $TOKEN" \
"$CATALOG_URI/v1/namespaces/$NAMESPACE/tables/$DATASET" | jq '.metadata.schemas[-1].fields'
Fix — rename the column in place in the dataset's parquetSchema to match the live table, and
roll the ingestor. Do not reorder the column list while you are in there: field-ids are assigned by
declaration order, so moving a line reassigns ids and misbinds every value written after it.
Appending is the only safe structural edit — and appending is a two-step change, in this order:
add the column to the live table first (in this version that is an external Iceberg-aware engine
through the shared catalog), then append it to parquetSchema
and roll. Boot commits a safe type widening on its own but never an add, so config that declares a
column the table does not have yet fails with "is missing from the live table" — that is the two
steps done in the wrong order, not drift.
Boot refuses a reorder rather than reading it as a rename — if it names "the columns were reordered or one was inserted", put the order back.
Verify — after the roll, the gauge reads 0 for that dataset and the boot WARN is gone.
Scenario 3 — compaction is not keeping up¶
Compaction is the only hand-rolled Iceberg operation TraceLake owns — iceberg-rust has no
rewrite_data_files, so expiry and appends use upstream and the merge does not. Treat repeated
failure here as a correctness risk, not just a backlog one.
Compaction backlog¶
Symptom — tracelake_compactor_partition_files_total at or above 500 for a partition, for an
hour. This is rare by design: sustained ingest is not supposed to be able to outrun
compaction. Sustained firing means something specific is wrong, so work the three branches in order
before concluding the fleet is undersized.
Branch A — duplicate daemons. Check this first, because the fix is a scale-down and the wrong diagnosis here costs money and pages.
kubectl get deploy -l app.kubernetes.io/component=compactor -o wide # must be 1
The compactor daemon has no range assignment: every replica sweeps every configured
dataset, so a second replica re-merges the same partition blocks. It is safe — the commit is
validation-scoped and there is an idempotency pre-check — but it burns CPU and object-store traffic
for nothing and drives Commit conflicts into a phantom page. If
TraceLakeCatalogCommitConflicts is firing steadily alongside this, that is the tell.
Scale to one. The chart pins replicas: 1 with strategy: Recreate and offers no
compactor.replicaCount for exactly this reason.
Branch B — incompatible type change. If the alert fires alongside compactor logs citing
incompatible schema at column '
The four non-breaking changes (add / drop / rename / safe widen) are not this branch —
they compact normally via the union-schema read, old rows null-filled. Only a narrowing or a
cross-family change lands here, which needs --allow-incompatible-schema (or an external writer) to
have happened at all. A column the current schema requires but a file leaves optional lands here too,
and only an external writer can produce that.
More pods will not help and there is no self-heal. The data is safe and fully queryable — only compaction of those partitions is stalled, so the cost is query latency and small-file count, not correctness. Repair is an external rewrite of the affected partitions (Spark, for example); the next compaction pass then picks them up. Queries on the affected column fail for the same reason — see Restricting a dataset across a type-change migration.
Datasets ordered after this one are not affected. A pass compacts every dataset and fails at the end naming the ones that failed, so a conflict here stalls only its own partitions. A second dataset alerting is its own incident — re-run this triage from the top for it.
What a persistent conflict does and does not suppress. It stalls compaction of the affected
partitions and nothing else. The daemon stays up on a failed pass: the
compaction.schedule floor keeps firing, so every other dataset still gets its below-threshold sweep
and its .idx backfill, and the snapshot-expiry and orphan-cleanup loops keep their own cadence.
TraceLakeCompactorRunsFailing fires every hour the conflict persists and is the signal to act on;
TraceLakeCompactorScrapeDown should stay silent, and if it is also firing the pod is down for a
different reason — triage that first. A manual pass (ingestor compact --config <file> --dataset …) still exits
non-zero, so a repair attempt reports through its exit code.
Branch C — genuinely undersized. Neither of the above, and the backlog is growing across many partitions.
Fix (Branch C) — drain with parallel manual Jobs over disjoint slices. Never by raising the replica count, which is Branch A on purpose:
for svc in checkout search payments; do
kubectl create job "drain-$svc" --image=ghcr.io/tracelake/tracelake:<tag> -- \
ingestor compact --config /etc/tracelake/config.yaml \
--dataset spans --partition "service_name=$svc"
done
Mount the same config Secret into each Job. Use --dry-run first to see the candidate files and the
estimated output without merging or committing anything. --date-range YYYY-MM-DD:YYYY-MM-DD slices
by time instead, and is ANDed with --partition.
A manually scoped run is one-shot and serves no /metrics — it registers its collectors, does
the work, and exits. No scrape covers it, so watch it by its Job status and
logs, not by a dashboard.
Longer term, raise the compactor's CPU/memory or lower compaction.backlogCheckIntervalMinutes (15)
so backlog-triggered runs start sooner.
Verify — partition_files_total falls below 500 for the affected partitions, and stays there
through a full ingest cycle. A backlog that clears and returns is Branch A or C, not a transient.
Index backfill not draining¶
Symptom — tracelake_compactor_unindexed_files_total above 0 for a partition, for 36 hours.
Data files there carry no Tantivy .idx sidecar, so every text predicate over that partition falls
back to a full scan.
Results stay correct. A missing sidecar is a positive pruning candidate: the file is scanned, never skipped. This alert is about query latency and object-store read volume, not about lost or wrong rows. Do not treat it as a data incident.
Why a file has no .idx. The sidecar is built at compaction only, never at ingest (to protect the
ingest CPU budget). A data file that rotated at ingestion.wal.maxFileSizeBytes is born at
or above compaction.minFileSizeMb, so the small-file rule alone would never pick it up. The
index-backfill rule exists to pick it up anyway — but only on the compaction.schedule
floor pass, not on the backlogCheckIntervalMinutes tick. So a non-zero reading between floor
passes is normal. Only a reading that survives more than one floor pass means something is wrong.
First, separate "behind" from "cannot complete". The gauge reports how many files are
unindexed, not whether they can be indexed. A file whose .idx build fails still commits, and the
floor pass re-elects it forever. One query tells the two apart:
sum by (dataset, stage) (increase(tracelake_compactor_idx_build_failures_total[36h]))
Zero means the backfill is behind — continue with the checks below. Non-zero means the backfill
cannot complete for those files, and the manual pass at the end of this runbook will not clear them.
The stage label says where the build gave up, and each stage takes a different route:
start— the schema or the index writer. Go to step 3 below.batch— a row the index rejected. Read the warn line the failure already emits:kubectl logs -l app.kubernetes.io/component=compactor | grep "tantivy index add failed".finish— the row-count refusal: the index holds a different number of rows than the Parquet file. This is a compactor bug, not a configuration fault. Capturegrep "tantivy .idx finalize failed"from the same logs and report it; the affected files stay scannable and correct in the meantime.upload— the.idxwas built and the object store refused the put. Check the storage account for a path ACL, an object-size limit or throttling on_metadata/datasets/<dataset>/idx/:kubectl logs -l app.kubernetes.io/component=compactor | grep "idx sidecar upload failed". The data files committed, so no rows are lost; the rewrite repeats until the put succeeds.
Check, in order.
-
Are floor passes running at all? If
TraceLakeCompactorRunsFailingorTraceLakeCompactorDeadlineExceededis firing alongside this, fix that first — this alert is a downstream symptom, not the cause. See Compactor failure and Compactor deadline. -
Is the schedule what you think it is?
bash
kubectl logs -l app.kubernetes.io/component=compactor --tail=200 | grep -i floor
A compaction.schedule of, for example, 0 3 * * 0 fires weekly, and a 36-hour for: will
page every week between passes. That is a threshold mismatch, not a fault — either widen the
alert's for: past the schedule interval or run the floor more often.
- Do the configured columns actually exist as text? A
tantivy.indexingColumnsentry that names a column absent from the table schema, or present but not a string, can never produce an.idx. The compactor skips those files rather than rewriting them forever, so this alert should not fire in that case — but the log line names it:
bash
kubectl logs -l app.kubernetes.io/component=compactor | grep "not a text column"
Fix the config to name a real text column, or drop the entry.
- Are these files even TraceLake's? Files another engine (Trino, Spark) appended to the table
are never backfilled and never counted here, by design. If the count is stuck and none of the
above applies, confirm the affected files sit under
{destinationUri}/{dataset}/data/.
Fix — drain it now with a manual pass rather than waiting for the next floor occurrence. A manual run always acts on backfill candidates:
kubectl create job drain-idx --image=ghcr.io/tracelake/tracelake:<tag> -- \
ingestor compact --config /etc/tracelake/config.yaml \
--dataset spans --partition "service_name=checkout"
Run it with --dry-run first to see the candidate files without merging or committing. Mount the
same config Secret the daemon uses.
A manually scoped run is one-shot and serves no /metrics — watch it by Job status and logs.
The gauge only moves after the daemon's next pass re-reads the manifests.
Verify — tracelake_compactor_unindexed_files_total for the affected partition falls to 0 after
the next daemon pass, and the sidecar-miss ratio for that dataset drops with it:
sum by (dataset) (rate(tracelake_gateway_sidecar_miss_total{kind=~"puffin|idx"}[30m])) / sum by (dataset) (rate(tracelake_gateway_files_scanned_count[30m])),
the expression TraceLakeSidecarMissRatioHigh alerts on.
Commit conflicts¶
Symptom — tracelake_catalog_commit_conflicts_total rate above 0.05/s for 30 minutes.
A conflict is a retry, not a loss. Two writers raced on the same table and one lost the CAS on the metadata pointer; it retries. An occasional conflict is the design working.
Note this series is absent from a scrape until the first conflict ever occurs (its dataset
label is not known at registration), so "no data" on the panel means no conflict has happened.
Check — the metric's own component label says who is fighting:
compactoragainstcompactor→ duplicate daemons. Go to Compaction backlog, Branch A.ingestoragainstcompactor→ normal contention. Sustained at this rate means the commit window is too tight: check catalog commit latency, and whetheringestion.commitIntervalSeconds(20) has been lowered.- Anything against an external engine — Trino or Spark writing to the same tables — is contention TraceLake cannot see. Its writes are legitimate; the tables are open by design.
Fix — remove the duplicate writer, or accept the contention if it is legitimate and raise
compaction.commitRetryMax (5) so retries win rather than surfacing as failed runs.
Verify — the rate falls below 0.05/s. Confirm compactor_runs_total{status="failed"} is not
rising: conflicts that exhaust the retry budget become run failures.
Compactor failure¶
Symptom — either any failed run in an hour (TraceLakeCompactorRunsFailing) or the scrape target
gone for 15 minutes (TraceLakeCompactorScrapeDown). Both route here because in both cases the
backlog stops draining silently, but they are not the same event: a
failed pass keeps the daemon up and fires only the first, while the second means the pod itself is
gone. Both are critical.
Check — scrape down first. The compactor binds /metrics before connecting the catalog, by
design, so an unreachable catalog still scrapes. That makes a missing scrape a strong signal:
kubectl get pods -l app.kubernetes.io/component=compactor
kubectl logs -l app.kubernetes.io/component=compactor --tail=200
CrashLoopBackOff → the process is dying at boot: config validation, catalog auth, or storage
credentials. A failed compaction pass does not land here — the daemon logs it
and keeps ticking, so that it stays up is what keeps the compaction.schedule floor firing
for every other dataset. OOMKilled → the merge is exceeding the pod's memory. The compactor keeps the merged files of one
partition in memory until it uploads them, so the memory grows with the total size of the input
files of the largest selected partition. Raise the memory limit, or merge that partition in smaller
slices with manual runs. Pod healthy but no scrape →
metrics.port (9090) versus the ServiceMonitor/annotation, and note the Service must be named
tracelake-compactor because the alert reads up{job="tracelake-compactor"}.
Also confirm nobody is running a manually scoped recovery Job and expecting it on the scrape — those serve no metrics by design and their absence is not this alert.
Check — runs failing. The logs name the failure. The three that matter:
- incompatible schema at column ... (field-id N) → Compaction backlog, Branch B. Those partitions are skipped by design until migrated externally; data is safe.
- commit conflict after
commitRetryMaxretries → Commit conflicts. - object-store or catalog errors → the same credential/reachability checks as Upload backlog.
Fix — per cause above. A failing compactor does not corrupt anything: the merge commit is
validation-scoped, so a run that fails mid-way leaves the table on its previous snapshot and the
merged output as unreferenced objects, which the orphan sweep reclaims after
maintenance.orphanMinAgeHours.
Verify — increase(tracelake_compactor_runs_total{status="success"}[1h]) > 0 and the backlog
resumes falling. A compactor that scrapes but never runs is not recovered.
One thing the daemon's unbounded lifetime costs you. The snapshot-expiry and orphan-cleanup loops run as detached tasks that nothing restarts, so if one dies it stays dead for the life of the pod. It is not silent: both are supervised, and a panic in either fires Background task died. Go to that runbook.
Background task died¶
Symptom — TraceLakeBackgroundTaskDied: tracelake_task_exits_total is non-zero.
{{ $labels.task }} names the loop and the alert's pod/instance label names the pod — the rule is
per-series and unaggregated precisely so the remedy below has something to point at.
One of the thirteen supervised loops panicked. Nothing restarts it — the supervisor reports, it does not self-heal — so that pod is running with that loop permanently missing. The other loops in the same pod are unaffected; a panic kills one task, not the process. The alert stays firing until the pod is replaced: it reads the raw counter, not a window, so it does not go quiet on its own while the loop is still gone.
What is now not happening, by task:
task |
Binary | Stopped |
|---|---|---|
snapshot_expiry |
compactor | Snapshots accumulate. Every CAS commit grows until commits themselves start failing |
orphan_cleanup |
compactor | Unreferenced objects accumulate. Storage cost, no correctness risk |
compactor_metrics_server |
compactor | The scrape — expect TraceLakeCompactorScrapeDown behind this |
batched_commit |
ingestor | Uploaded files never commit. Freshness and catalog_snapshot_age_seconds climb |
upload_scan |
ingestor | ready_* files pile up on the WAL disk — TraceLakeUploadBacklog, then the disk fills |
kafka_poll |
ingestor | That dataset stops consuming. Lag climbs |
kafka_lag_refresh |
ingestor | The lag gauge freezes at its last value — it reads as "lag stopped changing", not as missing |
gateway_snapshot_poll |
gateway | The gateway serves an ever-staler snapshot without knowing it |
catalog_token_refresh |
either | Everything catalog-facing fails once the current token expires |
wal_drain |
ingestor | That dataset stops ingesting entirely. Its kafka_poll loop ends cleanly behind it, so this counter is the only signal |
wal_age_sweep |
ingestor | Idle partitions never rotate — maxFileAgeMinutes is silently untrue for them |
wal_backpressure |
ingestor | Nothing pauses the consumers at the high-watermark; the WAL volume can fill to ENOSPC |
sigterm_handler |
any | The pod ignores SIGTERM and is SIGKILLed at the end of the grace period — no drain on the next rollout |
Check — find the panic. The supervisor logs error! with the task name at the moment of death,
and the unwinding panic's own backtrace is the line above it (the release profile keeps symbol names
for exactly this):
# The pod is the alert's own label; there is no search step.
kubectl logs <pod> | grep -B 20 "panicked; it will not restart on its own"
The backtrace is what matters — the counter says that it died, not why. A panic from
iceberg-rust, arrow, or the object provider on one dataset's metadata will recur on the
replacement pod; a panic in our own key parsing on one malformed object may not.
Fix — delete the pod:
kubectl delete pod <pod>
A restart revives every loop and costs one interval. If the alert returns within minutes of the
replacement coming up, the panic is deterministic on some input — do not keep cycling the pod.
Capture the backtrace, and for snapshot_expiry / orphan_cleanup scope the work away from the
failing dataset (ingestor compact --config <file> --dataset <name>) while it is diagnosed. Maintenance being off for hours
is survivable; a crash loop is not a fix.
Verify — the alert clears. The old pod's series goes stale and the replacement seeds its own at
0, which is the whole clear condition; if it is still firing, you replaced the wrong pod or the panic
recurred. Then confirm the loop's own signal moves again: tracelake_snapshots_expired_total rising
for snapshot_expiry, the backlog falling for upload_scan,
tracelake_catalog_snapshot_age_seconds dropping for gateway_snapshot_poll.
What does not fire this. A rolling restart, a SIGTERM drain, and CatalogClient shutting
down its own token refresh all cancel tasks on purpose, and cancellation is not counted — only a
panic is. If this fires, something genuinely crashed.
What it does not cover. Every long-lived loop in the fleet is in the table — ten async tasks and
the ingestor's three WAL threads. What this counter cannot see is a loop that
is wedged rather than dead: a thread blocked forever in a syscall has not panicked, so nothing
increments here. That has its own signal for the loops where it is decidable —
see Background task wedged. Five loops carry a liveness stamp;
sigterm_handler is not one of them, because a task that awaits one signal for the life of the pod
is supposed never to complete and a staleness rule on it would page on every healthy pod. For it,
and for the other event-driven loops, this counter plus up remain the only signals.
One special case for sigterm_handler: the pod it fires on cannot be drained by SIGTERM any more, so
kubectl delete pod will wait out terminationGracePeriodSeconds before the kubelet SIGKILLs it.
That is expected, not a second fault — there is nothing left to drain gracefully.
One for wal_drain too: its dataset stops ingesting, and Kafka lag climbs, but no kafka_poll exit
is counted. The poll loop ends cleanly when the panicking drain thread drops its record receiver,
and a clean return is deliberately not a fault. task="wal_drain" is the whole signal.
And the inverse for wal_drain and wal_age_sweep: their handled WAL errors exit the process with status 1,
so this counter never moves. That death surfaces as
a pod restart with this counter still at zero and a fatal WAL ...; exiting for restart line in the
logs — not as this alert. The two are different failures with different responses: the exit is the
recovery posture working as designed (WAL recovery reconciles the files on restart), while a
task="wal_drain" series means the thread panicked and the pod is still up with that dataset
silently not ingesting.
Background task wedged¶
Symptom — TraceLakeTaskWedged (wal_age_sweep, kafka_lag_refresh, gateway_snapshot_poll;
10 minutes stale by the time it pages — a 5-minute threshold plus a 5-minute for:) or
TraceLakeMaintenanceTaskWedged (snapshot_expiry, orphan_cleanup; 3 h 15 m stale):
time() - tracelake_task_last_pass_timestamp_seconds is past the loop's threshold.
{{ $labels.task }} names the loop, {{ $labels.dataset }} the dataset (- for the two
process-wide maintenance loops), and the alert's pod/instance label names the pod.
The loop is alive and stuck, not dead. It did not panic, so
Background task died is silent and will stay silent — that alert and this
one are the two halves of the same failure, and only this one fires here. Typical causes: an
object-store or catalog call with no timeout, a statvfs blocked on a dead NFS mount, a lock held
across an await, or a saturated spawn_blocking pool that never hands the task a thread.
What is now not happening, by task:
task |
Binary | Stopped |
|---|---|---|
wal_age_sweep |
ingestor | That dataset's idle partitions never rotate — maxFileAgeMinutes is silently untrue for them |
kafka_lag_refresh |
ingestor | tracelake_ingestor_consumer_lag is frozen at its last reading — a wedge stops the loop before it can remove the series, so TraceLakeConsumerLagHigh evaluates a stale number and cannot fire while real lag climbs |
gateway_snapshot_poll |
gateway | Queries silently serve an ever-staler snapshot. catalog_snapshot_age_seconds is republished at the end of the poll pass, so a wedge freezes it and TraceLakeCatalogSnapshotStale cannot fire either — and if the poller wedged at boot it reads a healthy 0 forever |
snapshot_expiry |
compactor | Snapshots accumulate. Every CAS commit grows until commits themselves start failing |
orphan_cleanup |
compactor | Unreferenced objects accumulate. Storage cost, no correctness risk |
Check — confirm it is a wedge and not a threshold that no longer matches the interval:
# The pod is the alert's own label; there is no search step.
kubectl logs <pod> --since=1h | grep -E "pass complete|sweep|expiry"
The loop logs a line at the end of every pass (snapshot-expiry pass complete,
orphan-cleanup pass complete). A gap in those lines matching the gap in the stamp is a real wedge.
If the lines are still flowing, the threshold is wrong, not the loop — see below.
Fix — delete the pod:
kubectl delete pod <pod>
Same remedy as a dead loop, for the same reason: nothing recovers a wedged loop in place, and a restart revives it at the cost of one interval. If it wedges again on the replacement pod, the hang is in a downstream dependency — check the object store and catalog for a hung endpoint before cycling the pod a third time.
Verify — the alert clears. The old pod's series leaves the scrape and the replacement seeds its
own stamp at spawn, so the clear is immediate rather than one interval later. Then confirm the loop's
own output moves again: a pass complete line for the maintenance loops, consumer_lag changing for
kafka_lag_refresh.
When the threshold is the problem, not the loop. TraceLakeMaintenanceTaskWedged's 10800s is 3
missed passes at the default 60-minute interval. Both
maintenance.expireSnapshotsIntervalMinutes and maintenance.orphanCleanupIntervalMinutes accept
15–1440, so a deployment that raised either above 60 will see this fire permanently on a healthy
compactor. Re-derive the threshold from the configured interval — the rule cannot read the config
key, exactly as TraceLakeWalDiskAboveHighWatermark cannot read the volume size.
TraceLakeTaskWedged's 300s has no such key: it is derived from two fixed intervals,
the WAL age sweep (30s) and the lag refresh (15s), and only a code change moves it.
What does not fire this. A SIGTERM drain and a pod that is shutting down: the terminating
pod's series leaves the scrape, and time() - <absent> yields no samples, so there is nothing to
compare. A rolling restart is likewise silent — the replacement seeds its stamp at spawn rather than
at 0, so it is never briefly "5 minutes stale" on boot.
What it does not cover. The eight loops with no stamp, and deliberately so. Four are event-driven
(kafka_poll, wal_drain, sigterm_handler, compactor_metrics_server): no pass is owed, an idle
topic legitimately produces none for hours, and a staleness rule on them would page on a healthy pod.
The other four are periodic but already have a consequence alert that fires behind a wedge, because
the gauge that alert reads is written outside the wedging loop: TraceLakeUploadBacklog for
upload_scan (wal_ready_files is incremented on the WAL rotation path),
TraceLakeCatalogSnapshotStale for batched_commit (the age gauge is published by the gateway, not
by the committer), TraceLakeWalDiskAboveHighWatermark for wal_backpressure (wal_disk_bytes is
set elsewhere, so a monitor that never pauses lets it climb), and the token's own expiry for
catalog_token_refresh. gateway_snapshot_poll fails that test and is stamped — see the table
above. The metrics page lists the loops that carry a stamp.
One thing worth knowing for kafka_lag_refresh: the stamp catches a wedge, and the gauge itself
covers the other route to a stale reading. A refresh tick publishes exactly the
partitions it obtained a reading for and removes the rest, so a committed() timeout or a
per-partition watermark failure shows up as a hole in tracelake_ingestor_consumer_lag (with a
warn! naming the cause) rather than as a frozen healthy number. Absence there means no reading,
never lag zero — see the metrics page.
Compactor deadline¶
Symptom — any increase in tracelake_compactor_deadline_exceeded_total over 6h.
A run passed compaction.runtimeDeadlineHours (6). The deadline is an alert threshold, not a
cap — the run continues, and now competes with the next scheduled one for CPU and object-store
bandwidth.
This has its own counter because compactor_duration_seconds tops out at a 3600s bucket: a
multi-hour run lands in +Inf where no rule can read it.
Check — is this one pathological partition or everything? Look at
partition_files_total for a partition with an extreme count, which makes one merge enormous. Check
whether compaction.targetSizeMb (128) was raised, which lengthens every run.
Fix — drain the large partitions with scoped manual Jobs (see
Compaction backlog) so the daemon's scheduled pass has less to do, then let it
resume. Raising runtimeDeadlineHours silences the alert without changing anything real; do that
only once you know the long runs are legitimate for your volume.
Verify — the counter stops increasing over a full cycle, and duration P95 comes back under
runtimeDeadlineHours.
Compactor ephemeral disk¶
Symptom — tracelake_compactor_ephemeral_peak_bytes above 80% of compactor.emptyDir.sizeLimit
for 5 minutes.
A compaction pass came close to filling the pod's /tmp volume. Passing it is a kubelet eviction
taken mid-merge: the pass is discarded, nothing commits, and the next pass re-elects the same files
and repeats it. That loop is why this is a warning you act on rather than watch.
The gauge is a per-pass high-water mark, held until the next pass overwrites it — a pass that merged nothing writes 0. Two consumers make up the peak and they overlap:
| Term | Scales with | Measured (reference stack) |
|---|---|---|
| Merge spill set | the partition's re-encoded size — not its input size | 362 MiB for a 3.69 GB / 229-file partition |
| Tantivy scratch | one output file (compaction.targetSizeMb) × the dataset's text density |
1.34 GiB for one 128 MB output |
The index term is usually the larger one, and it surprises people. An n-gram index over
text-heavy columns is an order of magnitude bigger than the data it indexes (11× in
one measurement). A dataset whose tantivy.indexingColumns name large free-text
columns pays it on every output.
Check
# 1. Which dataset, and how close to the cap?
kubectl exec deploy/<release>-compactor -- \
wget -qO- localhost:9090/metrics | grep ephemeral_peak_bytes
# 2. Which term — compare the .idx sidecars this dataset writes against its data files.
# A ratio well above 1 says the index term dominates.
# sum(rate(tracelake_compactor_sidecar_bytes_sum{kind="idx"}[1h]))
# / sum(rate(tracelake_compactor_sidecar_bytes_sum{kind="data"}[1h]))
# 3. Is anything stale left over from a killed pod? The compactor sweeps these at boot
# (older than 1 hour), so a live one is a merge in flight.
kubectl exec deploy/<release>-compactor -- ls /tmp
Fix — in order of preference:
- Raise the ceiling together.
compactor.emptyDir.sizeLimit,compactor.resources.requests.ephemeral-storageand.limits.ephemeral-storageare one decision in three places; the limit must stay abovesizeLimit(it also covers logs and the writable layer), andsizeLimitabove the request. Raising only the limit leaves the volume's own cap evicting you anyway. - Lower
compaction.targetSizeMb. The index term is per output file, so halving the target roughly halves it. The cost is more files per partition, which is the backlog alert's input. - Narrow
tantivy.indexingColumns, or loweringestion.tantivy.maxIndexedValueBytes, if a large column is being indexed that no query ever filters on. Both change what future files can prune, not correctness: a file with no usable index is scanned, never skipped.
Do not reach for compaction.sortColumns here. Removing them drops the spill entirely but gives
up clustering, and the spill is the smaller term.
Verify — the next pass's peak lands below the threshold, and the partition stops being re-elected
(ingestor compact --config <file> --dry-run --dataset <ds> returns no candidates for it).
An eviction does not appear in tracelake_compactor_runs_total{status="failed"} — the process is
killed before it can increment anything, and the daemon's own error path never runs. What it looks
like instead: the pod's restart count goes up, the scrape gaps (long enough and
TraceLakeCompactorScrapeDown fires at 15m), and the same partition is elected again on the next
tick with the same file count. A failed pass — a merge that ran and errored — is the other
signature, and that one does move the counter.
Scenario 4 — the OIDC provider is down¶
The gateway validates a bearer token on every request against keys it caches from the
provider. Losing the provider does not fail queries immediately: the cached keys keep working until
queryGateway.auth.maxJwksStalenessSeconds (3600), after which the gateway answers
503 UPSTREAM_AUTH_UNAVAILABLE — never 401. That distinction is the contract: a 401 would tell
callers their token is bad when the truth is that TraceLake cannot check it.
JWKS staleness¶
Symptom — tracelake_gateway_jwks_cache_age_seconds above 2880s for 10 minutes. The threshold
sits below the 3600s cap on purpose, so the page arrives while queries still work.
Check
# 1. Is the provider reachable at all?
curl -s -o /dev/null -w '%{http_code}\n' "$ISSUER/.well-known/openid-configuration"
curl -s -o /dev/null -w '%{http_code}\n' "$JWKS_URI"
# 2. Is this an outage, or are callers simply sending bad tokens?
# Outage: jwks_cache_age rising + error_code="UPSTREAM_AUTH_UNAVAILABLE" rate rising
# + jwks fetch outcome=error (not bare status="503": a busy pod answers
# QUERY_CONCURRENCY_EXCEEDED, a catalog outage CATALOG_UNAVAILABLE)
# Bad token: cache age flat + 401 rate rising
rate(tracelake_gateway_jwks_fetch_duration_seconds_count{outcome="error"}[5m]) rising is the
provider failing the fetch. A rising 401 rate with a flat cache age is a client problem and does not
belong to this runbook.
Rule out the boring causes before declaring an outage: a rotated issuerUrl, a network policy that
now blocks egress from the gateway pods, and an expired TLS chain on the provider all present
identically.
Fix — restore the provider. TraceLake caches keys but cannot mint them.
If the provider will be down past the cap and you must keep serving, raising
maxJwksStalenessSeconds extends the window — understand what you are choosing: TraceLake will keep
accepting tokens signed by keys it can no longer confirm are current, so a key revoked during the
outage stays honoured. That is a deliberate security trade, not a config tweak.
Verify — cache age drops to near zero on the next successful fetch, the
error_code="UPSTREAM_AUTH_UNAVAILABLE" rate returns to zero,
and the 401 rate returns to its baseline (which is not necessarily zero — see
the RBAC note below for why a steady 403 baseline is normal).
Scenario 5 — orphan cleanup is deleting too much¶
Orphan cleanup is the one maintenance path that physically deletes objects, and the only
place where a bug costs data rather than latency. Its guard is age: the sweep deletes only objects
older than maintenance.orphanMinAgeHours (default 24, and required to be ≥ 24 because that is the
WAL recovery horizon).
Orphan deletion spike¶
Symptom — more than 1000 deletions in an hour on tracelake_orphan_files_deleted_total. A normal
sweep deletes a handful.
DATA LOSS RISK. Stop the sweep before diagnosing. Do not let it finish and investigate afterwards — the objects are gone, and Kafka will not redeliver what was already committed.
kubectl scale deploy/<release>-compactor --replicas=0
Check — first ask whether the spike is expected. Two causes are legitimate:
- A catalog point-in-time restore. Restoring to an earlier point makes the catalog forget snapshots committed after it; the data files those referenced are genuinely unreferenced now. This is the documented aftermath and the deletions are correct. The rows that were committed after the restore point are no longer in the table.
- A large expiry pass.
snapshotRetentionHours(24) expiring a burst of snapshots strands their exclusive data files.
If neither applies, treat it as a bug until proven otherwise. Then re-run the sweep as a pre-flight that deletes nothing:
ingestor compact --config /etc/tracelake/config.yaml --orphan-dry-run
That lists the objects a real sweep would reclaim and the Puffin/.idx sidecars that would go
with them. Read the list: if it names files under a partition that is currently being written, or
files newer than orphanMinAgeHours, stop and escalate — the age guard is what stops the sweep from
deleting files a concurrent writer has staged but not yet committed.
Fix
- Confirm
maintenance.orphanMinAgeHoursis still ≥ 24 andcompaction.snapshotRetentionHoursis ≥ it. Never shorten the age guard to tidy up faster — that guard is the data-loss protection, not a performance knob. - If the dry run looks correct, scale the compactor back to 1 and let the sweep proceed.
- If files were already deleted and they were not orphans, this is a recovery from the object store's own versioning/soft-delete, if you enabled it. TraceLake keeps no second copy.
Verify — after the sweep, --orphan-dry-run lists few or no candidates, and queries against the
affected datasets return the expected row counts. Confirm through an external engine (Trino) as well
as through the gateway, since the gateway may be serving a cached snapshot that predates the
deletion.
Scenario 6 — the query path is degraded but working¶
None of these lose data or break queries. They are the pruning and caching machinery losing effectiveness, which shows up as latency and object-store cost.
They are grouped here because the alert bundle covers query-path degradation that the five scenario-level failures above do not touch.
Index cache hit ratio¶
Symptom — windowed hit ratio below 0.8 for 30 minutes. Queries are re-reading Tantivy .idx
sidecars from object storage instead of serving them from memory.
Check — tracelake_gateway_index_cache_bytes against queryGateway.indexCache.maxBytes
(512MB default). If the cache is full, the working set outgrew it. Also check whether query patterns
changed — a shift to older data walks sidecars that were never going to be cached.
The alert deliberately computes the ratio from the hit/miss counters rather than reading
tracelake_gateway_index_cache_hit_ratio, because that gauge is lifetime-cumulative: a pod healthy
for a week will not cross the line for hours after it goes bad. Use the same windowed view when
judging recovery.
Fix — raise indexCache.maxBytes, or accept it. This is a cost/latency trade, not a fault.
Verify — windowed ratio back above 0.8 and query duration P95 improved. If the ratio recovers and latency does not, the cache was not the bottleneck.
Sidecar miss ratio¶
Symptom — more than 20% of scanned files have no Puffin/.idx companion, for 30 minutes.
A missing sidecar is scanned, never skipped. That is the rule, and it is why this costs latency and not correctness — a file with no sidecar is a positive pruning candidate.
Check — some misses are structural and permanent:
- Tantivy
.idxis built at compaction only, never at ingest (it protects the ingest CPU budget). Recently-ingested, uncompacted files legitimately have none. - Files written by an external engine (Trino, Spark) have no TraceLake sidecars at all, ever.
So the question is whether the ratio is rising. A rising ratio with a rising compaction backlog is one problem, not two: compaction is what creates sidecars.
Fix — clear the compaction backlog. If the files are external-engine writes, this is the expected steady state and the alert threshold should be raised for that dataset.
Verify — ratio falls as compaction catches up.
Manifest cache thrash¶
Symptom — a non-zero manifest cache miss rate for a full hour.
The 1h for: is load-bearing: a query misses whenever it reaches a partition range whose manifests
this pod has not loaded — every window after boot, and each older window the first time. The poller
loads each commit's new manifests, so snapshot changes do not miss. Sustained misses are the load →
evict → load loop: the windows queries keep asking for do not fit, and each costs its manifest GETs
again.
Check — tracelake_gateway_manifest_cache_bytes summed over datasets against
queryGateway.manifestCache.maxBytes (512MB). Two things fill it: each dataset's manifest list,
which is never evicted and costs ~250 B per live data file (about 300 MiB for 1.2 million files), and the manifests queries load, held
compressed a day to a segment at ~160–190 B per file in one measurement. Lists near
the cap leave no room for loaded manifests, and every query misses. A file count far above design
is often a compaction backlog.
Fix — raise manifestCache.maxBytes (and the gateway's memory request with it), or reduce the
file count by clearing the backlog.
Verify — miss rate returns to occasional (boot and first queries of an older window only).
Puffin read amplification¶
Symptom — consulted ÷ fetched Puffin bytes below 0.4 for 6h. Informational — this does not page. The gateway is fetching bloom blobs it does not use.
Check — the ratio reads ≈ 1/N with one blob consulted out of N fetched. 0.4 is measured, not
guessed: the ranged-read path only models a win at ≥3 blobs, and then only at wide fan-out. A low
ratio usually means queries filter on fewer columns than bloomFilterColumns declares.
Fix — narrow datasets[].bloomFilterColumns to the columns actually filtered on. Every extra
column is a blob fetched on every scan.
Verify — ratio rises and tracelake_gateway_puffin_bytes_total{part="fetched"} falls.
Idx pruning ineffective¶
Symptom — prune_candidates_out_total{pass="idx"} ÷ _in_total{pass="idx"} above 0.9 for an
hour. Pass 3 is fetching an .idx per candidate file and keeping almost all of them: every query on
this dataset pays a sidecar fetch that buys nothing.
This is the pass-3 sibling of Sidecar miss ratio (the sidecar that is
absent) and Puffin read amplification (the same finding one pass
earlier). The in leg only advances when the pass actually ran — a wholesale decline records
neither leg — so this never fires on a workload the index is not asked to serve.
Check — break it down before changing anything:
sum by (reason) (rate(tracelake_gateway_idx_declines_total{dataset="..."}[1h])). The reason that
dominates decides the fix, and only one of the five is a defect.
Fix — by reason:
pattern_too_short— the workload'sLIKEliterals are shorter than one gram, so no n-gram query can be built and every file is kept. Raisingingestion.tantivy.ngramSizedoes not help (a wider gram makes more patterns too short); a narrower one would, at the cost of a larger.idx. Weigh that againsttracelake_compactor_sidecar_bytes—rate(_sum{kind="idx"}) / rate(_sum{kind="data"})is what the trade costs.no_field— the.idxwas written before the column joineddatasets[].tantivy.indexingColumns, or the file was written by an external engine. Self-clearing: the next compaction of those files writes the field. If it is not clearing, the compactor is behind — see Compaction backlog.no_rowgroups— the.idxcarries no__rowgroupsentry for this file, so a matching doc cannot be mapped onto a row group. Same posture asno_field: re-compaction rewrites it. An entry that is present and does not decode is counted asunreadable, not here.capped— the column holds values longer thaningestion.tantivy.maxIndexedValueBytes, so only their head is indexed and a non-match cannot be trusted. Not a defect and not self-clearing: re-compaction caps the same values again. Either accept the scan for that column, or raise the cap — up to Tantivy's 65530-byte ceiling, above which it drops the term itself. Weigh a raise againsttracelake_compactor_sidecar_bytes{kind="idx"}, which is what the larger index costs.unreadable— a real defect. The.idxarchive is corrupt, its__rowgroupsor__cappedentry does not decode (a truncated upload reaches the gateway this way), or the file was written by an incompatible build. Which file is not in the metric — the label set isdatasetandreasononly. Raise the gateway toRUST_LOG=tracelake_gateway::prune=debugand re-run the query: each offender logsidx unreadable; file scanned wholewith itspath. Capture one object before deleting anything; the file is still correct to scan (a missing sidecar means scan, never skip), so this costs latency, not rows.
Verify — the ratio falls below 0.9 and the dominant reason series flattens.
Postgres resolve failures¶
Symptom — a non-zero rate on tracelake_gateway_pg_table_resolve_failures_total. Only exists
when queryGateway.postgres is configured; on a gateway without the PostgreSQL-wire front-end this
series is absent, and absence is not zero.
The blast radius is wider than one dataset. While one dataset fails to resolve,
information_schema.columns fails for every dataset — a single \d in psql resolves a snapshot
per configured dataset, so one failure breaks introspection for all of them.
Check — which dataset, from the metric's label. Then whether that dataset's table exists in the
catalog at all: a dataset configured in ingestion.datasets that no ingestor has started with has no
table, and resolving it fails.
There is a third cause that is not a resolve failure at all: a dataset whose snapshot carries
row-level delete files is refused by design, and the refusal is counted here because it is raised
from the same seam. The tell is SQLSTATE 55000 in the client's error and a
"snapshot carries delete files" line in the gateway log — the catalog is perfectly healthy. See
A dataset refuses every query.
Fix — remove the dataset from config if it is not real, or start an ingestor with that dataset, which creates the table. If the catalog is unreachable, this is Catalog snapshot stale wearing a different hat.
Verify — the rate returns to zero and \d works in psql for every configured dataset.
A note that belongs here because it is where operators land: on the PostgreSQL wire,
tracelake_gateway_pg_queries_total{status="denied"} is policy working — the read-only statement allowlist,
dataset RBAC, and the raw-window cap all report denied. It is non-zero on a healthy deployment and
nothing alerts on it alone. The same is true of 403s on the HTTP path: with datasetRoles configured, dataset RBAC denies by
default, so a steady 403 baseline is the guard doing its job.
Catalog backup failing or stale¶
Alerts: TraceLakeCatalogBackupFailing, TraceLakeCatalogBackupStale
Chart-only alerts, present only when catalogBackup is active (the bundled Postgres is enabled) and
only if kube-state-metrics is scraped. They are not in the main alert bundle.
Symptom¶
Either a -catalog-backup Job has failed, or none has completed successfully in 48 hours. The
second is the dangerous one: a CronJob that stopped firing looks exactly like one that is working,
and nothing else in the deployment notices.
Neither alert means data is lost. It means the catalog — the only non-derivable state in the deployment — has no fresh dump behind it. Ingest, compaction and queries are unaffected and will stay unaffected right up until the day they are not.
Check¶
kubectl get cronjob -l app.kubernetes.io/component=catalog-backup
kubectl get jobs -l app.kubernetes.io/component=catalog-backup --sort-by=.status.startTime
kubectl logs job/<most-recent-failed-job>
Then look at the volume itself — most failures here are the boring one:
kubectl exec -it <any-tracelake-pod> -- df -h /backup # if it mounts the claim
# or, more reliably, a throwaway pod on the claim:
kubectl run pvc-check --rm -it --restart=Never --image=busybox \
--overrides='{"spec":{"volumes":[{"name":"b","persistentVolumeClaim":{"claimName":"<release>-catalog-backup"}}],"containers":[{"name":"c","image":"busybox","command":["df","-h","/backup"],"volumeMounts":[{"name":"b","mountPath":"/backup"}]}]}}'
In rough order of likelihood:
| Cause | What you'll see |
|---|---|
| PVC full | No space left on device in the Job log; df at 100% |
| CronJob suspended | kubectl get cronjob shows SUSPEND: True; no recent Jobs at all |
| Postgres unreachable | could not connect to server — check the catalog is up before assuming backup-specific breakage |
pg_dump older than the server |
server version N; pg_dump version M — the subchart's Postgres was bumped past the derived image |
| Secret renamed | password authentication failed, or the pod stuck on a missing secretKeyRef |
| Dump below the floor | FATAL: dump is N bytes, below minDumpBytes= — see below, this one is not a backup problem |
If the log says below minDumpBytes, treat it as a catalog problem, not a backup problem. The
job did connect, pg_dump did succeed, and what came back was nearly empty. That means the database
it dumped is nearly empty — a migration that wiped it, a restore that landed blank, or PGDATABASE
pointing somewhere unintended. The job deliberately failed and pruned nothing, so your existing
backups are intact. Check the live catalog before doing anything else:
kubectl exec -it <release>-lakekeeper-pg-0 -- psql -U lakekeeper -d lakekeeper \
-c "select count(*) from information_schema.tables where table_schema not in ('pg_catalog','information_schema')"
If that count is low, you are in a catalog-loss incident and the restore procedure in Catalog backup and restore is what you want — not another backup run.
Fix¶
PVC full — raise catalogBackup.persistence.size (if the storage class allows expansion) or
lower catalogBackup.retentionDays. Prune runs after a successful dump, so a volume too full to
write one never prunes itself out of the hole; delete the oldest catalog-*.sql.gz by hand to break
the deadlock.
Suspended — kubectl patch cronjob <name> -p '{"spec":{"suspend":false}}'. Find out who
suspended it before assuming it was accidental.
pg_dump version — set catalogBackup.image.tag to match the server (SELECT version()), or
the chart's derivation is behind the subchart. pg_dump refuses to dump a newer server, by design.
Run one now, rather than waiting for the schedule:
kubectl create job --from=cronjob/<release>-catalog-backup catalog-backup-manual
kubectl logs -f job/catalog-backup-manual
Verify¶
A successful run leaves one file:
/backup/catalog-<timestamp>.sql.gz
The encryption key is deliberately not written beside it — see
Catalog backup and restore. A dump
without its key is not a backup: a correctly restored
database under a different LAKEKEEPER__PG_ENCRYPTION_KEY resolves the warehouse and lists tables,
then fails loadTable with SecretFetchError — healthy-looking and useless. So confirm you hold a
break-glass copy of <release>-lakekeeper-postgres-encryption as well; the dump alone will not
restore.
To prove the artifact actually restores, run the procedure in Catalog backup and restore against the newest dump, into a scratch database.
What this does not cover¶
The artifacts are on a PVC in this cluster. That covers database-level loss — a bad migration, a dropped database, logical corruption. It does not survive losing the cluster or the zone, because the backup goes with it. Keep the encryption key wherever you keep break-glass secrets, not only in the cluster whose loss is the scenario you are preparing for.
Pod stuck terminating¶
Alerts: none — this one you reach from kubectl, not from a page. There is deliberately no drain
metric (see What this page is not); the signals below are the ones an
existing family already carries.
Symptom¶
A pod sits in Terminating for most of its terminationGracePeriodSeconds (120 s for an ingestor,
60 s for a gateway) during a rollout, or is killed at the end of it. A rollout that should take
seconds takes minutes per pod.
On SIGTERM an ingestor leaves its Kafka consumer groups, which rotates every active_* WAL file to
ready_*, then uploads, commits, and exits; a gateway stops accepting on both the HTTP and PostgreSQL-wire ports
and lets in-flight responses finish. Each is bounded by its own drainTimeoutSeconds (80% of the
grace period, computed by the chart), so a healthy drain exits well inside the budget.
No data is at risk here. A drain cut short at any point leaves state recovery already handles: rotation renames and commits the offset exactly as the steady-state path does, an uploaded but uncommitted file is the case the next pod's recovery reconciles, and anything still in a consumer's channel when polling stopped was never committed, so Kafka redelivers it. The cost of a slow drain is time and re-consumed records, not rows.
Check¶
- Read the pod's last logs — the drain narrates every step:
kubectl logs <pod> --previous | tail -40
"SIGTERM received; draining" means the handler fired. If that line is absent, the signal
never reached the process; skip to the Fix note about PID 1.
"drain deadline exceeded" means the drain ran but did not finish inside drainTimeoutSeconds.
"drain complete" with a large elapsed_s means it finished, but slowly.
-
Which step was it in? The last drain log line names it: rotating a WAL, uploading
ready_*files, or committing. On a gateway,"postgres drain: ... waiting for in-flight sessions"with no following"postgres drain complete"means a SQL session would not let go. -
If the drain stalls on upload or commit, it is the same underlying condition as Upload backlog or Commit conflicts — object storage or the catalog is slow or unreachable. Check those first; the drain is a symptom, not the cause.
-
If it stalls on the gateway's PostgreSQL surface,
tracelake_gateway_pg_connectionson the terminating pod is the count still holding on. A BI tool's pooled connection running a long scan is the usual answer.
Fix¶
- Immediate, for one stuck pod: send a second
SIGTERM. The process exits immediately rather than finishing the drain — which is safe, per the paragraph above, and is why you should not needSIGKILL:
kubectl exec <pod> -- kill -TERM 1
-
Drain stalled on storage or the catalog: fix that, not the drain. Raising the grace period hides a dependency outage behind a longer rollout.
-
No
"SIGTERM received"line at all: the signal is not reaching PID 1. That is an image or workload defect, not a drain one — a shell-form entrypoint, a wrapper script, or acommand:override puts something other than the binary at PID 1, and PID 1 with no handler ignores SIGTERM. The shipped image is checked for both halves (PID 1 is the binary; SIGTERM exits it cleanly); compare it with whatever image the pod is actually using. -
Do not raise
terminationGracePeriodSecondsto make this go away, and do not add apreStophook. The grace period is the budget the drain fits inside, and the chart derivesdrainTimeoutSecondsfrom it — raising it widens both, which lengthens every rollout without fixing whatever the drain is waiting on.
Verify¶
Roll one pod and watch it:
kubectl delete pod <pod> --wait=true
It should go from Terminating to gone in seconds, not at the end of the grace period. On the
replacement pod, no "recovery: dropping stray active_* file" lines at startup means the previous
pod rotated its WAL rather than leaving it to be reconciled.
A dataset refuses every query — delete files¶
No alert points here directly, and the refusal has no metric of its own — it renders through the
ordinary error path, as tracelake_gateway_query_requests_total{status="501",
error_code="DELETE_FILES_UNSUPPORTED"} on HTTP and
tracelake_gateway_pg_queries_total{status="failed"} on the Postgres wire. If the PostgreSQL-wire
front-end is configured you may nonetheless arrive from TraceLakePostgresTableResolveFailures,
which this condition does fire: the refusal is raised from the same seam and is counted in
tracelake_gateway_pg_table_resolve_failures_total. Otherwise you got here from a user report.
Symptom — one dataset answers 501 DELETE_FILES_UNSUPPORTED (SQLSTATE 55000) to every query on
both surfaces. The gateway logs
"snapshot carries delete files; TraceLake does not apply them" with the dataset and a count — once
per snapshot refresh, not once per query.
Other datasets still answer queries, but on the Postgres wire information_schema.columns fails
while this lasts — one \d resolves a snapshot per configured dataset and the first failure fails the
query, the same ceiling as Postgres resolve failures. Driver
autocomplete and schema reflection are down until the deletes are cleared.
This is not a fault and not data loss. Someone ran a DELETE (or a MERGE/upsert) against the
Iceberg table with an external engine. Iceberg recorded it as a position- or equality-delete file
rather than by rewriting data. TraceLake does not apply delete files in this version, and the alternative
to refusing is serving rows Iceberg says are gone. The data is intact and fully readable — with
the deletes correctly applied — through Trino, Spark, or DuckDB against the same catalog. Point
affected users there while you clear it.
Check — confirm the deletes are real and find out who wrote them:
# Delete-file count on the current snapshot, straight from the table.
trino> SELECT file_path, content, record_count
FROM <catalog>.<namespace>."<table>$files"
WHERE content != 0;
# Which snapshot introduced them, and what operation it was.
trino> SELECT snapshot_id, committed_at, operation, summary
FROM <catalog>.<namespace>."<table>$snapshots"
ORDER BY committed_at DESC LIMIT 5;
content 1 is a position delete, 2 an equality delete. operation = 'delete' or 'overwrite' with
a non-TraceLake summary names the writer.
Fix — compact the deletes away with the engine that wrote them, so the next snapshot carries none:
trino> ALTER TABLE <catalog>.<namespace>.<table> EXECUTE optimize;
or, from Spark, CALL <catalog>.system.rewrite_data_files(table => '<ns>.<table>'). Either rewrites
the affected data files with the deleted rows removed, which is the state TraceLake reads.
Do not delete the delete files by hand out of the bucket — that resurrects the deleted rows and corrupts the snapshot that references them. TraceLake's own orphan sweep will not touch them: the sweep walks every manifest entry regardless of content type, so a delete file referenced by any retained snapshot is in the live set and is never a sweep candidate.
Then make sure it does not recur: TraceLake writes no deletes itself, so a recurring refusal means a pipeline is issuing them. Either stop that, schedule a compaction after it, or accept that the dataset is served through Trino rather than through TraceLake.
Verify — the $files query above returns no rows with content != 0, and a query through the
gateway answers 200. Allow up to queryGateway.snapshotPollIntervalSeconds for the gateway to pick up the new
snapshot; tracelake_catalog_snapshot_age_seconds falling confirms it did.
Restricting a dataset across a type-change migration¶
No alert points here, and this is not a fault — it is the operator procedure to use
while an incompatible type change is being migrated externally. You get here from having run the
ingestor with --allow-incompatible-schema, or from a user reporting failed queries on one column
of one dataset.
Symptom — one dataset's queries fail, but only the ones naming the column whose type changed. Queries on every other column of the same dataset answer normally, and every other dataset is unaffected. The failure is a cast refusal from the scan (Cannot cast string '…' to value of Int32 type, or the equivalent for the pair in question), raised per file as the pre-migration bytes meet the post-migration schema.
The failure does not carry a failing status code. The status line is already on the
wire by the time the first batch is read, so the failure arrives as the error trailer — a
well-formed 200 whose body carries the error code:
{"rows":[],"error":{"code":"STORAGE_INTERNAL_ERROR","message":"the query could not be completed","target":"spans","timestamp":"2026-08-27T09:14:02Z"}}
The scrape says so too: the trailer's code is carried back to the request
boundary, so the query counts as
tracelake_gateway_query_requests_total{status="200",error_code="STORAGE_INTERNAL_ERROR"} and lands
on the dashboard's errors by code panel. It stays a 200 on the by-status panel, because a 200
is what went out on the wire. On the Postgres wire the query fails and counts in
tracelake_gateway_pg_queries_total{status="failed"}.
Fence it anyway. What the metric buys you is knowing; it does not stop the failures. Every
query naming the column keeps failing for as long as the boundary is open, and a client that does
not check the error member reads the trailered body as an empty result set. The
fence turns that into an explicit 403 with a dataset_fenced gauge an operator can read while it
is in force.
This is not data loss and not wrong rows. The old bytes are never reinterpreted as the new type — the cast refuses on the data. The table is intact and fully readable through Trino, Spark, or DuckDB against the same catalog, which is where the migration itself runs.
Check — confirm which field-id is at the boundary. The ingestor named it when it was started across the change:
kubectl logs -l app.kubernetes.io/component=ingestor --since=24h \
| grep 'incompatible schema at column'
The line carries the dataset, the column, the field-id and both types. The same string appears in compactor exits — if it is also stalling compaction, that is Compaction backlog, Branch B, and the fix below does not clear it.
Fix — fence the dataset for the duration of the migration, with the deny-all default of datasetRoles. This is
the control in this version; there is deliberately no schema-version pin (see Why not a pin below).
queryGateway:
auth:
roleClaimPath: realm_access.roles
datasetRoles:
spans: [] # <- was [trace-reader]; empty list authorizes no role
Keep the entry and empty its list rather than deleting the key: deny-all fences it either way, and the entry is what reminds you the dataset is fenced when the migration ends. Roll the gateway — the policy is compiled once at boot and there is no reload path.
Then run the migration and restore the role list, rolling the gateway again. The migration is done with any Iceberg-aware engine through the shared catalog — Spark is one; DuckDB is the smallest, and is below.
Migrating the type with DuckDB¶
One option, not the procedure. Spark is the other. DuckDB is the smallest thing that does it.
Do not hand-rebuild the table. A CREATE TABLE AS SELECT produces a table with no partition
spec, no sort order, and field-ids in whatever order the SELECT listed — and the ingestor validates
all three against config at boot. Provisioning against a CTAS-shaped table fails with, verbatim:
catalog error: dataset 'spans': partition spec drift — config declares 2 partition field(s),
the live table has 0; recreate the table or restore partitionMapping to match
Restore the partition spec and sort order drift fires next, and that one has no recovery in
this version: a sort order is registered at table creation only, and a CTAS cannot express one. The table
would be unbootable and the fix would be to recreate it again.
So let the provisioning path that already gets those right build the table, and use DuckDB only to move the rows.
Step 0 — rehearse on a copy. Not optional and not a footnote: steps 2–4 rename the table holding the only copy of the data. Point a scratch config at a copied dataset and run the whole sequence there first.
Step 1 — attach the same catalog TraceLake uses. Substitute your own endpoints and credentials (storage.catalogUri, storage.catalogAuth) and keep
real secrets out of shell history.
INSTALL iceberg; LOAD iceberg;
INSTALL httpfs; LOAD httpfs;
-- Global `SET s3_*`, not a named S3 secret: the iceberg extension's storage reads do not pick up a
-- named secret and 403. This is the most common way this attach fails.
SET s3_endpoint='minio:9000';
SET s3_access_key_id='…';
SET s3_secret_access_key='…';
SET s3_use_ssl=false;
SET s3_url_style='path';
SET s3_region='us-east-1';
CREATE SECRET lk (TYPE ICEBERG, CLIENT_ID '…', CLIENT_SECRET '…',
OAUTH2_SERVER_URI 'https://…/token', OAUTH2_SCOPE 'lakekeeper');
ATTACH 'tracelake' AS ice (TYPE ICEBERG, ENDPOINT 'https://…/catalog', SECRET lk);
If the change is non-breaking, stop here — you should not be in this runbook. Add, drop, rename and the three safe widenings are metadata-only and DuckDB expresses them directly, with no rewrite, no fence and no downtime:
ALTER TABLE ice.tracelake.spans ALTER COLUMN ts TYPE BIGINT; -- int32 → int64, a safe widen
Step 2 — drain the WAL, then move the live table aside. Iceberg only promotes losslessly, so the
incompatible change (string → int) cannot be an ALTER in any engine; the data has to be
rewritten. Renaming rather than dropping is what preserves the pre-migration snapshots.
The rename is only safe against a drained WAL, and this is the step that can lose you data if you
skip it. WAL recovery reconciles each ready_* file against the table config names. Rename the table out from under it and that
lookup finds the fresh, empty spans: a file already appended to the old table is no longer "in a
retained snapshot", so recovery downgrades delete to leave, and the upload loop appends it a
second time — then step 4 copies the same rows a third time out of spans_premigration. Those files
also still hold service_name as STRING, so they either reintroduce the boundary into the new
table or fail the append outright.
Stop the ingestor cleanly and confirm the WAL is empty, per pod — it is a fleet:
kubectl scale statefulset/<release>-ingestor --replicas=0
# Then, for each WAL volume (`ingestion.wal.localDirectory`): no ready_* and no active_* may remain.
kubectl exec <pod> -- ls <ingestion.wal.localDirectory>
Only then:
ALTER TABLE ice.tracelake.spans RENAME TO spans_premigration;
Step 3 — let the ingestor provision the new table. Edit the dataset's parquetSchema to the new
type, then start the ingestor. It creates spans from config with the correct field-ids, partition
spec and sort order by construction — the three things a CTAS gets wrong.
parquetSchema: "span_id STRING, service_name INTEGER, message STRING, ts TIMESTAMPTZ, event_date LONG"
Take --allow-incompatible-schema off at the same time. You are here because it was set. After
step 2 the table is absent, so this boot creates rather than validates and the flag does nothing —
leaving it on silently disarms the guard for the next genuine incompatible change, which is the same
outliving-its-migration shape this runbook already warns about for datasetRoles.
"Booted clean" means specifically that it created the table, not that it validated one. If
spans still exists at this point — the rename did not land, or a sibling pod raced you — the
ingestor validates an existing table instead, and you get parity drift rather than a fresh table.
This is also the step that proves the new schema is one the ingestor accepts, while the old data is still untouched under its own name.
Then stop it again before the backfill. It provisioned the table; letting it keep appending live Kafka rows through step 4 makes the step 5 count check meaningless (see there). Nothing is lost by pausing — Kafka retains, and the ingestor resumes from its committed offsets — and the dataset is fenced from queries anyway.
Step 4 — take a baseline, then backfill through DuckDB. The ingestor consumed Kafka from the
moment it started in step 3, and ingestion.wal.maxFileAgeMinutes defaults to 5 — so any care
taken confirming that boot probably crossed a rollover, and spans is not empty. Record what is
already in it, or step 5's arithmetic is measuring your own confirmation time:
-- Whatever step 3 ingested while you were confirming the boot. Run this immediately before the INSERT.
SELECT count(*) AS baseline,
count(*) FILTER (WHERE service_name IS NULL) AS null_baseline
FROM ice.tracelake.spans;
If the dataset declares sortColumns, DuckDB refuses the insert. The reference spans does, so
expect this on the first attempt:
Not implemented Error: INSERT into a sorted iceberg table is not supported yet.
To bypass this guard and write without applying the table's declared sort order, run
"SET unsafe_iceberg_ignore_sort_order=true"
Taking the bypass is correct here, despite the name, and it is worth knowing why rather than
trusting the adjective. The files it writes are unclustered with sort_order_id unset — which is
exactly what the ingest path already writes; TraceLake clusters at compaction, not at write,
and the next compaction pass sorts these files like any other. The bypass
skips applying an ordering that nothing at this stage was going to apply anyway. It does not write
files that claim an ordering they lack, which is the failure the guard exists to prevent.
The guard is version-scoped, and not yet escapable. duckdb/duckdb-iceberg #1135 (merged
2026-07-03) removes it and applies the sort order on write. It is not reachable from DuckDB v1.5.5,
the version this procedure was verified on: measured 2026-08-27, the guard still fires on both the
core and core_nightly extension repos, and on a sorted-unpartitioned table too — the case #1135
explicitly handles. The nightly served to a v1.5.5 CLI is built from the v1.5 branch, which never
got that change; reaching it needs a newer CLI. Retest this on a newer DuckDB,
and drop the SET below if the insert succeeds without it.
Then backfill. BY NAME matches columns by name, so declaration order stops mattering; TRY_CAST
yields NULL on an unparseable value instead of aborting the statement, which is what makes the
check in step 5 mean something.
SET unsafe_iceberg_ignore_sort_order=true; -- only if the dataset declares sortColumns; see above
INSERT INTO ice.tracelake.spans BY NAME
SELECT span_id,
TRY_CAST(service_name AS INTEGER) AS service_name,
message,
ts,
event_date
FROM ice.tracelake.spans_premigration;
Name every column. SELECT * hides a column added since you last looked, and event_date in
particular is a partition source — dropping it silently would repartition the dataset.
Step 5 — check what the cast cost, before you rely on the result. Compare against the source's own NULL count, not against zero: the column is optional, so pre-existing NULLs are not loss.
SELECT (SELECT count(*) FROM ice.tracelake.spans_premigration) AS before,
(SELECT count(*) FROM ice.tracelake.spans) AS after,
(SELECT count(*) FROM ice.tracelake.spans_premigration WHERE service_name IS NULL) AS null_before,
(SELECT count(*) FROM ice.tracelake.spans WHERE service_name IS NULL) AS null_after;
Subtract step 4's baseline before comparing: after - baseline must equal before. The
subtraction is what makes this exact no matter how long step 3 took — the operator never has to
reason about whether a WAL rollover landed mid-procedure, which is the thing that would be got wrong
at 3am. It holds because the ingestor is stopped across step 4; started, it would move after while
the INSERT runs.
Likewise null_after - null_baseline above null_before is the count of values the cast could not
parse — decide deliberately whether that is acceptable rather than discovering it later.
If you genuinely cannot pause ingest across step 4, drop the baseline and bound the after halves at
the cutover instead (WHERE event_date <= <the event_date at step 2>), treating the result as
approximate: late-arriving rows land in old partitions and inflate it.
Then restart the ingestor, restore the role list, and roll the gateway; it picks up the new table
within one queryGateway.snapshotPollIntervalSeconds.
Three things specific to TraceLake:
service_nameis a partition source column (partitionMappingin the example). Changing the type of a partitioned-on column is the harder case: every partition tuple in the new table is derived from the new type, so the old and new tables do not share partition values even where the underlying rows are identical. A plain data column is simpler than this example.
A failed cast is also a repartitioning. TRY_CAST turns an unparseable value into NULL, and
on a partition source column that moves the row into the null partition, where it becomes
indistinguishable from rows that were always null. Observed in the verification run — the file
holding both the genuinely-null row and the cast-failed one:
s3://traces/…/data/event_date=20260828/service_name=__HIVE_DEFAULT_PARTITION__/….parquet 2 rows
That is why step 5 takes the null counts before you rely on the result: afterwards, no query
over the new table can separate a cast failure from an original null.
- Sidecars do not come across. Puffin blooms and .idx are per-data-file and keyed on the data
file's UUID; the backfilled files are new files with none. Queries stay correct — a missing
sidecar is a positive scan candidate, never a skip — and the next compaction pass rebuilds them.
Expect slower queries until it does.
- Keep spans_premigration for a while. It holds the pre-migration snapshots and is the only way
back. Drop it deliberately, after the migration is confirmed in production.
Note
This procedure was run end to end on DuckDB v1.5.5 against an S3-compatible object store;
the SET s3_* lines in step 1 are for such a store. It was not run against Azure Blob. The
ingestor's own start and WAL recovery around steps 2–3 were not part of that run. Step 0 still
stands.
While it is in force, every attempt is a defined 403 FORBIDDEN (SQLSTATE 42501 on the
Postgres wire) — a counted failure on the right dataset, instead of a 200 that reads as a success:
tracelake_gateway_query_requests_total{dataset="spans",status="403",error_code="FORBIDDEN"}
That series only moves when someone queries, and it does not exist at all until the first 403 — a failure pair is absent from the scrape rather than zero (see the metrics page), so no data there does not mean not fenced. The signal that does not depend on traffic is the gauge:
tracelake_gateway_dataset_fenced{dataset="spans"} == 1
Seeded for every configured dataset, so 0 really is 0, and it covers both fenced shapes —
the role list emptied, and the entry deleted outright. The gateway also says it once per pod start:
WARN dataset="spans" dataset fenced: no role authorized …
That line only covers the emptied shape: it walks the entries that exist, so a dataset fenced by deleting its entry raises the gauge and logs nothing. One more reason to empty the list rather than delete the key, as Fix above says.
A fence left up after the migration ends reads exactly like one still in progress, on every signal here. Treat restoring the role list as part of the migration, not as cleanup afterwards.
Verify — tracelake_gateway_dataset_fenced{dataset="spans"} is back to 0, and a query naming
the previously-boundary column answers
200 with rows. Until the external rewrite has actually replaced the pre-migration files, restoring
the roles brings the cast refusal back rather than the rows — and, per the symptom above, brings it
back uncounted: verify the rewrite first, through the engine that ran it.
Why not a pin. This version has no control that restricts queries to a single schema version:
the Iceberg library cannot reach a per-file schema id that survives compaction or
snapshot expiry, so a pin would silently serve the wrong files. datasetRoles is the control.
Ingestor refuses to boot — partition spec drift¶
Alerts: none — you reach this one from a CrashLoopBackOff, not from a page. The ingestor exits
non-zero at boot, before it consumes anything.
Symptom¶
kubectl logs -l app.kubernetes.io/component=ingestor --tail=40
dataset 'spans': partition spec drift — config's partitionMapping declares a partition spec the
live table does not have. Committing it is metadata-only and rewrites nothing, but it is
irreversible: Iceberg retains the superseded spec forever. If the change is intended, restart with
--allow-partition-evolution; if partitionMapping was edited by mistake, restore it instead
The other wording — "the table was evolved past this config … Update partitionMapping to the live spec" — is not this runbook. That one says the live table already carries a newer default spec than config describes, and the flag does not clear it: fix the config, do not evolve backwards.
Check¶
Decide whether partitionMapping was changed on purpose. Compare it to the table's live default:
-- Trino, or any type=rest engine
SELECT * FROM ice.tracelake."spans$partitions" LIMIT 1;
SHOW CREATE TABLE ice.tracelake.spans; -- the partitioning = [...] clause
# and what the pods were actually given
kubectl get secret <release>-config -o jsonpath='{.data.config\.yaml}' \
| base64 -d | grep -A5 partitionMapping
If the config's mapping is not the one you intended to ship — a renamed key, a dropped entry, a bad merge — stop here and restore it. Booting with the flag would add that mistake to the table permanently.
Fix¶
Only if the change is deliberate. Two things to know before you run it:
- It cannot be undone. TraceLake never removes a partition spec; Iceberg keeps the superseded spec in the table metadata forever. That retention is what keeps files already written under the old spec readable — it is correct, just not reversible.
- The whole fleet has to restart, not one pod. Each pod captures the spec id once, at boot, and the append rejects a data file stamped with any other id. A pod that was already running when the evolution landed keeps appending under the old id and starts failing its uploads until it is restarted.
Preferred — put the flag on the StatefulSet, let the rolling restart land it. (The ingestor is a
StatefulSet, not a Deployment: volumeClaimTemplates is what gives each pod a WAL volume that
survives rescheduling, which is what makes WAL recovery real.) The first pod to come up commits the
spec; the others find it already present and boot straight through, logging a sibling pod committed
the same partition spec first; continuing.
kubectl patch statefulset/<release>-ingestor --type=json \
-p '[{"op":"add","path":"/spec/template/spec/containers/0/args/-","value":"--allow-partition-evolution"}]'
kubectl rollout status statefulset/<release>-ingestor
A kubectl patch on a Helm-managed StatefulSet is reverted by the next helm upgrade — normally
harmless, since the flag is meant to come off anyway, but an upgrade landing mid-roll drops it and
the not-yet-restarted pods refuse to boot. Do not upgrade until the rollout completes.
Alternative — one container, by hand. Useful when you want to watch the commit land before
rolling the fleet. ingestor run is a daemon: it commits the spec during boot and then keeps
consuming, which is why this is a one-off pod and not a Job.
Never run a second
ingestor runagainst alocalDirectorya live pod is using — scale the StatefulSet to0first. There is no cross-process WAL lock, and this is not a race you might win. The new process's recovery scan deletes the running pod's in-flightactive_*segments unconditionally, and from then on both processes' upload loops scan the sameready_*files and both append them — permanent duplicate rows, for as long as both run.And give the one-off pod the fleet's WAL claim, not a fresh volume. The Kafka offset is committed at rotation, before the Iceberg append — that asymmetry is exactly what WAL recovery exists to reconcile. A rotated
ready_*left in a throwaway volume is a file whose offsets are already committed and which no later boot will ever scan: Kafka does not redeliver, nothing reports it, and the rows are simply gone. Mountingwal-<release>-ingestor-0means anything the one-off leaves behind is recovered by<release>-ingestor-0on the way back up — recovery working as designed instead of bypassed. With the StatefulSet at0replicas nothing else holds the claim.
Between them, that is also why this is a fresh container rather than kubectl exec into a live pod:
exec manufactures the co-tenancy, and there is nothing to attach to anyway while the pod is in
CrashLoopBackOff.
kubectl scale statefulset/<release>-ingestor --replicas=0
# The config Secret is <release>-config — one Secret for all three workloads — and the WAL is the
# existing claim, not an ephemeral volume. The mount path is the literal /var/lib/tracelake/wal:
# the chart hard-fails on any other `ingestion.wal.localDirectory`, so on a Helm
# deployment there is nothing to look up. Off-chart, use whatever that key is set to.
kubectl run ingestor-evolve --rm -it --restart=Never \
--image=ghcr.io/tracelake/tracelake:<tag> \
--overrides='{"spec":{"volumes":[{"name":"config","secret":{"secretName":"<release>-config"}},{"name":"wal","persistentVolumeClaim":{"claimName":"wal-<release>-ingestor-0"}}],"containers":[{"name":"ingestor-evolve","image":"ghcr.io/tracelake/tracelake:<tag>","command":["ingestor"],"args":["run","--config","/etc/tracelake/config.yaml","--allow-partition-evolution"],"volumeMounts":[{"name":"config","mountPath":"/etc/tracelake","readOnly":true},{"name":"wal","mountPath":"/var/lib/tracelake/wal"}],"stdin":true,"tty":true}]}}'
On a plain container runtime the same thing is podman run / docker run against your own
deployment's image, config and WAL volume, with ingestor run --config <path>
--allow-partition-evolution as the command.
Watch for:
INFO dataset="spans" spec_id=1 committed a partition spec evolution at boot …
Then stop it with a single Ctrl-C and let the drain finish — do not interrupt it again. A drain cut short
at any point is safe because a later boot recovers the
directory; on SIGTERM the ingestor leaves its consumer groups, which rotates what is in flight, then uploads,
commits, and exits. A second Ctrl-C, a kill -9, or exceeding ingestion.drainTimeoutSeconds leaves files for
recovery instead — which is fine here only because you mounted the real claim. Give it longer than
drainTimeoutSeconds before worrying.
Then bring the fleet back without the flag — kubectl scale statefulset/<release>-ingestor
--replicas=<n>. Like --allow-incompatible-schema, it is an acknowledgment of one change, not a
setting: left on, it silently disarms the guard for the next accidental partitionMapping edit.
Verify¶
kubectl get pods -l app.kubernetes.io/component=ingestor # all Running, no restarts
kubectl logs -l app.kubernetes.io/component=ingestor | grep -c "partition spec drift" # 0
And confirm rows are landing under the new spec — the boot commit alone proves nothing about the append path:
SELECT partition, record_count FROM ice.tracelake."spans$partitions";
A new partition tuple appearing with a rising record_count is the check. Freshness is minutes, so
give it one ingestion.wal.maxFileAgeMinutes window (default 5) plus an upload cycle before calling
it failed.
What this page is not¶
There is no automated remediation here, by design. Every step is something an operator runs deliberately.
The two paths that do self-correct are called out where they apply: WAL backpressure pauses and resumes Kafka polling on its own, and a catalog outage recovers with no TraceLake-side action once the catalog returns.