Alerts¶
TraceLake has a bundle of 30 Prometheus alert rules. Each rule has a severity, a subsystem, a for:
duration, and a runbook_url that points to a section of the runbook. The rules read
the metrics on the metrics page. The Helm chart can install the bundle as a
PrometheusRule.
Labels¶
| Label | Values | Purpose |
|---|---|---|
severity |
critical, warning, info |
Routing. critical pages; warning is a ticket; info never pages. |
subsystem |
ingestor, gateway, compactor, catalog, maintenance |
Which part of TraceLake owns the response. |
subsystem, not component, and that is deliberate. component is already a metric label on
the catalog metrics, carrying the emitter (ingestor/compactor/maintenance). A static
alert label of the same name silently overwrites it, and TraceLakeCatalogCommitConflicts would
have reported its own group instead of naming the writer that was actually conflicting.
Thresholds that mirror TraceLake config¶
These are set to the configuration defaults and are plain values in the rules file. The Helm chart derives some of them from its values. A deployment that changes the underlying key must re-derive the threshold — nothing detects a changed key for you.
| Alert | Threshold | Mirrors |
|---|---|---|
TraceLakeWalDiskAboveHighWatermark |
42949672960 bytes |
ingestion.wal.diskHighWatermarkPct (80) × the WAL volume size |
TraceLakeWalBufferSaturated |
214748364 bytes |
80% of ingestion.wal.maxBufferedBytes (268435456). Unlike the disk watermark above, the chart re-derives this one from the key, so only the static bundle carries the literal |
TraceLakeFreshnessAboveTarget |
420 s |
The published P95 target. It is a fixed target, not a mirror of the sum: the four stages (maxFileAgeMinutes + upload + commitIntervalSeconds + snapshotPollIntervalSeconds) come to 343 s at defaults, leaving 77 s of headroom. Raising any of them spends that headroom and is not re-derived here — at maxFileAgeMinutes' 60-minute ceiling the sum is far past 420 s and this rule fires permanently on a correctly-configured deployment |
TraceLakeJwksCacheStale |
2880 s |
80% of queryGateway.auth.maxJwksStalenessSeconds (3600) |
TraceLakeCatalogSnapshotStale |
60 s |
About 7 missed polls at snapshotPollIntervalSeconds: 8 |
TraceLakeCompactionBacklog |
500 files |
The files-per-partition ceiling |
TraceLakeIndexBackfillNotDraining |
any unindexed file | No config key. The for: (36h) is what carries the threshold: it must exceed one compaction.schedule interval, since the backfill merge runs on the floor pass only |
TraceLakeCompactorDeadlineExceeded |
any increase | compaction.runtimeDeadlineHours (6), applied by the compactor, not the rule |
TraceLakeCompactorEphemeralHigh |
5153960755 bytes |
80% of the chart's compactor.emptyDir.sizeLimit (6Gi). Not a TraceLake configuration key — the eviction it predicts is the kubelet's, against the volume's own cap — so the chart re-derives it from that value and only the static bundle carries the literal. Its for: (5m) must not outlive compaction.backlogCheckIntervalMinutes (5–120), since the gauge is a per-pass high-water mark the next pass overwrites. It equals that range's floor rather than clearing it: at an interval of 5 the only margin is the pass's own duration |
TraceLakeCompactorRunsFailing |
any failed run in 3h |
compaction.backlogCheckIntervalMinutes — the lookback, not the threshold, mirrors it: a permanently failing daemon stays up and steps the counter once per tick, so a window shorter than the key's 120-minute ceiling reads 0 between failures and the alert flaps. Raise the lookback past any raised interval |
TraceLakeMaintenanceTaskWedged |
10800 s |
3 missed passes at the default 60 of maintenance.expireSnapshotsIntervalMinutes and maintenance.orphanCleanupIntervalMinutes. Both keys accept 15–1440 |
TraceLakeTaskWedged |
300 s |
No config key. 10 missed WAL age sweeps (30s) and 20 missed lag refreshes (15s). Both intervals are fixed — only a code change moves this number |
TraceLakeWalDiskAboveHighWatermark is the one threshold with no default to inherit: the WAL
volume size is a chart value, not a TraceLake config key, so the shipped number is 80% of a 50Gi
reference volume. Set it per deployment or the alert is decorative.
TraceLakeWalBufferSaturated mirrors a key that does exist — ingestion.wal.maxBufferedBytes, 80%
of it — and the chart re-derives the threshold from that key rather than shipping the static number,
so raising the budget moves the alert with it. The static bundle carries 80% of the default.
TraceLakeMaintenanceTaskWedged is the second such case, for the opposite reason: its two keys
have defaults, but both accept 15–1440, so a deployment that raises either above 60 minutes will see
this fire permanently on a healthy compactor. Re-derive it from the configured interval.
Ingestion¶
| Alert | Expression, in words | for: |
Severity | Meaning |
|---|---|---|---|---|
TraceLakeConsumerLagHigh |
lag > 1M on any partition | 30m | warning | The ingestor is not keeping up. Check WAL disk first — polling pauses at the high watermark, so a full volume causes lag rather than resulting from it. → runbook |
TraceLakeWalDiskAboveHighWatermark |
WAL bytes past the watermark | 5m | critical | Kafka polling is paused. The drain is upload+commit, so check upload errors and catalog reachability. → runbook |
TraceLakeWalBufferSaturated |
buffered records > 80% of maxBufferedBytes |
30m | warning | The per-dataset record buffer is at its budget. Memory is bounded and nothing is failing — what is being paid is flush batch size, because the partitionMapping has opened enough segments that 1000 rows each no longer fit. Not backpressure. → runbook |
TraceLakeUploadBacklog |
any ready_* file pending |
30m | warning | Rotated WAL files are not reaching cloud storage, and the volume fills behind them. → runbook |
TraceLakeUploadErrors |
upload error rate > 0 | 15m | warning | Uploads retry internally, so a blip is survivable; a sustained rate means the backlog grows faster than it drains. → runbook |
TraceLakeConsumeErrors |
any increase over 1h | 5m | warning | recv() errors arrive once per state change — a deleted topic produces a single error, so a rate-based rule would never fire. → runbook |
TraceLakeReadyCleanupSlow |
cleanup P95 > 1s | 30m | warning | WAL unlinks are blocking the commit path far past the sub-millisecond norm. Read with wal_ready_files to tell a large batch from a slow volume. → runbook |
TraceLakeFreshnessAboveTarget |
freshness P95 > 420s | 30m | warning | The published target is breached. The lag is in one of four stages: rotation, upload, commit batching, or the gateway poll. → runbook |
TraceLakeRecordsDropped |
drop rate > 0 | 30m | warning | Payloads are failing to decode and are being dropped by the drain policy. Data is being lost silently — usually a producer-side schema change nobody announced. → runbook |
TraceLakeWalSegmentOverfilled |
any increase over 1h | 5m | critical | A WAL file closed holding more rows than its own Kafka offset range can contain — the same offsets were appended twice and the duplicate rows are already on disk, bound for Iceberg. No threshold to tune and no config key to mirror: a segment covering start..=end holds at most end - start + 1 rows, so every non-zero value is corruption rather than a tolerance being exceeded. increase, not rate, for TraceLakeConsumeErrors' reason — one overfilled segment is a single step. Its sibling wal_duplicate_records_total counts duplicates the append guard prevented and is deliberately not alerted on. Clearing is not repair: increase[1h] ages the alert out an hour after the last step, while the duplicate rows are permanent — resolution means no new occurrence. → runbook |
TraceLakeSchemaNameDrift |
renamed columns > 0 | 1h | warning | parquetSchema still names a column the live table renamed. Boot tolerates it (identity is the field-id), but new data files carry the stale name. Not an outage; update config and roll. → runbook |
Query gateway¶
| Alert | Expression, in words | for: |
Severity | Meaning |
|---|---|---|---|---|
TraceLakeIndexCacheHitRatioLow |
windowed hit ratio < 0.8 | 30m | warning | Queries are re-reading sidecars from object storage. Raise indexCache.maxBytes, or accept that the working set outgrew it. → runbook |
TraceLakeSidecarMissRatioHigh |
> 20% of scanned files have no puffin/idx sidecar |
30m | warning | Pruning is falling back to full scans. Sidecars are built at compaction only, so this usually means the compactor is behind. → runbook |
TraceLakeManifestCacheThrash |
miss rate > 0 on a warm pod | 1h | warning | The load→evict→load loop: the windows queries keep asking for do not fit under manifestCache.maxBytes (manifest lists + loaded manifests), and each query pays their manifest GETs again. → runbook |
TraceLakeCatalogSnapshotStale |
snapshot age > 60s | 10m | critical | A catalog outage. Queries keep serving the last good snapshot — they go stale, not down, so new data is invisible without anything erroring. The exception is a window whose manifests this pod never loaded (an older range, or any query on a pod restarted during the outage): that is a 503. → runbook |
TraceLakeJwksCacheStale |
JWKS age > 80% of the cutoff | 10m | critical | The window in which the OIDC provider can still be fixed. Past the cutoff, every query becomes a 503. → runbook |
TraceLakePostgresTableResolveFailures |
resolve failure rate > 0 | 15m | warning | While one dataset fails to resolve, information_schema.columns fails for every other dataset too — one \d resolves a snapshot per configured dataset. → runbook |
TraceLakePuffinReadAmplification |
consulted/fetched < 0.4 | 6h | info | The only alert whose remedy is a code change, not a config change — a standing review trigger, never a page. → runbook |
TraceLakeIdxPruningIneffective |
pass-3 out/in > 0.9 | 1h | warning | The .idx pass is fetching a sidecar per candidate file and keeping over 90% of them. Break down by idx_declines_total{reason} before acting: only unreadable is a defect, the rest are workload, staleness, or a value cap. → runbook |
Expression choices that look wrong and are not¶
The hit-ratio rule does not use tracelake_gateway_index_cache_hit_ratio. That gauge is a
lifetime ratio; a pod that ran hot for a week hides a cold hour inside its own average, so no
windowed alert can use it. The rule divides the two counters over a 30m window instead. On an idle
gateway that is 0/0 → NaN, and a NaN comparison is false — silence on no traffic is intended, not
a gap. The gauge remains useful on a dashboard, which is what it is for.
The manifest-cache rule alerts on the miss rate, not a hit ratio. The poller keeps every
configured dataset warm, so steady state is ~0 misses and a ratio would sit pinned at 100% until it
dipped for benign reasons. The 1h for: is load-bearing: every pod legitimately misses at boot, and again the first time a query reaches a window that the pod
has not loaded, so a shorter window would fire on every rollout.
The sidecar-miss rule selects kind, it does not sum the family. The
counter also carries kind="ndv", which is the planner estimating from Iceberg column stats instead
of a distinct-value sketch — that reads no extra byte, so folding it into a read-amplification ratio would
measure something the rule is not about, and would sit pinned above the threshold on every dataset
that configures no ndvColumns. Hence {kind=~"puffin|idx"}: those two, and only those, mean the
file is scanned.
0.4 on the Puffin rule is measured, not guessed. With one blob consulted out of N the ratio
reads ≈ 1/N, and a measurement showed that the ranged-read path only models a win at N ≥ 3
(≈ 0.33), and then only at wide fan-out. 0.4 sits between the 3-column ratio and the 2-column one
(0.5), which lands on the 20% margin — inside the error bar of the assumed RTT. The .idx is
deliberately and permanently absent from this rule: 95.2% of that container is what a query already
reads, leaving 4.8% to skip across 15 scattered ranges.
The pass-3 rule divides two counters that a declining pass never touches. The pass records
neither leg when it declines wholesale, so a workload the index is not asked to serve leaves both
series flat and the ratio undefined — silence, exactly as on the idle-gateway case above. Recording
in == out for a pass that did not run would have made every such workload fire. The counters carry
a pass label: divide them within one pass, and do not add them across passes.
Maintenance and compaction¶
| Alert | Expression, in words | for: |
Severity | Meaning |
|---|---|---|---|---|
TraceLakeCatalogCommitConflicts |
conflict rate > 0.05/s | 30m | warning | Writers are serialising against each other. A conflict is a retry, not a loss — an occasional one is the design working. → runbook |
TraceLakeCompactionBacklog |
≥ 500 files in a partition | 1h | warning | Rare by design. Sustained firing means the compactor fleet is undersized, not that the rule is noisy. → runbook |
TraceLakeIndexBackfillNotDraining |
any data file with no .idx |
36h | warning | Text predicates on the partition full-scan. Results stay correct (a missing sidecar is scanned, never skipped) — this is latency, not data loss. A non-zero reading between floor passes is normal, which is what the long for: absorbs. → runbook |
TraceLakeCompactorRunsFailing |
any failed run in 3h | 5m | critical | Compaction is the only hand-rolled Iceberg operation TraceLake owns. Treat a repeated failure as a correctness risk, not just a backlog one. → runbook |
TraceLakeCompactorScrapeDown |
up == 0 |
15m | critical | The pod itself is gone — boot failure, OOM, eviction, node loss. A failed pass does not land here: the daemon stays up and reports through TraceLakeCompactorRunsFailing. While the pod is down nothing compacts, nothing expires, and no orphans are swept. → runbook |
TraceLakeCompactorDeadlineExceeded |
any increase in 6h | 5m | warning | A run passed runtimeDeadlineHours. The deadline is an alert threshold, not a cap — the run continues, competing with the next one. → runbook |
TraceLakeCompactorEphemeralHigh |
a pass peaked > 80% of the /tmp volume's sizeLimit |
5m | warning | The next merge of that size is an eviction taken mid-merge: the pass is discarded and the next one repeats it. Peak is the merge spill set plus the Tantivy scratch of the output being written, and on a text-heavy dataset the index term is the larger of the two (measured 11× the data it indexes). → runbook |
TraceLakeOrphanDeletionSpike |
> 1000 deletions in 1h | 5m | critical | Orphan cleanup physically deletes objects and is the one maintenance path that can lose data. A normal sweep deletes a handful. → runbook |
TraceLakeBackgroundTaskDied |
a pod's task-exit counter is non-zero | 5m | critical | A supervised loop panicked. It stays dead for the life of the pod — the supervisor reports, it does not restart — so the rule fires until that pod is replaced, not for a window after the death. Counts panics only: a deliberate abort and a clean return are teardown, so a rolling restart never fires this. → runbook |
TraceLakeTaskWedged |
a wal_age_sweep, kafka_lag_refresh or gateway_snapshot_poll stamp is > 5 min stale |
5m | critical | The other half of the above: the loop is alive but stuck and never panicked, so the exit counter stays flat. A wedged wal_age_sweep makes maxFileAgeMinutes untrue for that dataset's idle partitions; a wedged kafka_lag_refresh freezes consumer_lag so TraceLakeConsumerLagHigh cannot fire; a wedged gateway_snapshot_poll freezes catalog_snapshot_age_seconds so TraceLakeCatalogSnapshotStale cannot fire either. Total wedge age at page time is threshold + for: = 10 min, which is what the annotation says. → runbook |
TraceLakeMaintenanceTaskWedged |
a snapshot_expiry or orphan_cleanup stamp is > 3 h stale |
15m | critical | Same signal at the maintenance cadence, which is three orders of magnitude longer and cannot share a threshold. Neither loop has any other signal behind a wedge. → runbook |
TraceLakeCompactorScrapeDown reads up{job="tracelake-compactor"}; the Helm chart
sets that job label. Note that the compactor binds its metrics endpoint before connecting the
catalog, precisely so an unreachable catalog still scrapes rather than surfacing here as a dead pod.
Manually scoped recovery runs (ingestor compact --dataset/--partition/--orphan-dry-run) do not
serve metrics at all — they are one-shot and exit — so no rule covers them.
TraceLakeBackgroundTaskDied is filed here because the two loops it exists for are
snapshot_expiry and orphan_cleanup — two of the four with no consequence alert behind them
(nothing alerts on snapshots_expired_total being flat, and the orphan rule catches too many
deletions, never zero); wal_age_sweep and sigterm_handler are the other two.
It covers the remaining nine too, where it fires earlier and names the dead task rather than the
symptom: today a dead upload_scan surfaces as TraceLakeUploadBacklog and a dead
gateway_snapshot_poll as a rising catalog_snapshot_age_seconds. task is the loop name, not
the dataset — see the metrics page.
It reads the raw counter, not an increase() window, and is deliberately not aggregated. The
condition is "this pod is running with a dead loop", which persists until the pod is replaced; any
window would resolve the alert with the loop still gone, which is the silence the rule exists to
remove. The counter resets on process restart, so the replacement pod clears it. Per-series because
the remedy is to delete one named pod — sum by (task) would drop the label that names it.
It covers every long-lived loop in the fleet: ten async tasks and the ingestor's three WAL
threads (wal_drain, wal_age_sweep, wal_backpressure). Two of the loops had no alert of any
kind behind them: wal_age_sweep is the only thing that finalizes an idle partition's segment, and
a dead sigterm_handler leaves the pod ignoring SIGTERM until the kubelet's SIGKILL, which nothing
observes until the next rollout.
Routing caveat: every task carries subsystem: maintenance, so an Alertmanager route keyed on
subsystem sends kafka_poll, kafka_lag_refresh, batched_commit, upload_scan,
gateway_snapshot_poll, wal_drain, wal_age_sweep, wal_backpressure and sigterm_handler to
the maintenance owner — ingestion and pod-lifecycle failures included. The task label tells the
human what actually died; it does not steer the router.
TraceLakeTaskWedged and TraceLakeMaintenanceTaskWedged are the other half of that
counter, and they take the same routing caveat. Neither reads task_exits_total — a wedged loop
never panicked, so nothing increments and the death rule is silent by construction. They read
tracelake_task_last_pass_timestamp_seconds, which only five of the thirteen loops carry:
wal_age_sweep, kafka_lag_refresh, gateway_snapshot_poll, snapshot_expiry and
orphan_cleanup. The other eight are omitted deliberately — four are event-driven, where a stale
stamp is a healthy idle pod rather than a fault, and four already have a consequence that
shows behind a wedge. The test that sorts them: is the gauge the consequence alert reads written
outside the loop that would wedge? gateway_snapshot_poll is stamped because it fails that test —
the poll republishes catalog_snapshot_age_seconds at the end of its own pass, so a wedge freezes
that gauge rather than letting it rise, and the gauge starts at 0, so a poller wedged
from boot publishes a healthy 0 forever.
Two rules rather than one because the five loops fall into two cadences three orders of magnitude apart: 30s/15s loops against 60-minute maintenance loops. A single threshold either pages on every normal maintenance gap or takes hours to notice a wedged sweeper. They share one runbook anchor, because the remedy — replace the pod — does not differ.
Series that may legitimately be absent¶
Four families can be missing from a scrape entirely. Every rule that touches them alerts on
> threshold, so absence and zero mean the same thing and no absent() guard is required — but a
rule added later must not read absence as a value.
- The four
pg_*families register only whenqueryGateway.postgresis configured. A gateway without the second door exposes none of them. tracelake_catalog_commit_conflicts_totalis never seeded: itsdatasetlabel is unknown at registration, so it appears only once a conflict has actually happened.tracelake_compactor_sidecar_bytesis deliberately not seeded either, and unlike the two above it can be absent on a healthy deployment:.idxand Puffin sidecars are written at compaction only, so the family has no series until a pass writes an object, and a manually scoped run (--dataset/--partition/--orphan-dry-run) registers the collector without ever serving/metrics. A pre-seeded histogram would publish a zero-byte ratio, which reads as a fact rather than as silence. No rule in this bundle reads it — it is a capacity signal on the compactor dashboard row, and that row being empty between runs is expected.tracelake_ingestor_consumer_lag, perdataset/partition. It is not seeded, and unlike the three above it is also removed: the refresh loop publishes exactly the partitions it obtained a reading for and drops the rest. A missing partition means no reading was obtained — acommitted()timeout, afetch_watermarksfailure, a blocking-pool stall, a partition revoked to another pod, a pod with no assignment at all, or a partition that has not committed an offset yet (routine for the first minutes after a pod joins) — and never lag zero. That is whatTraceLakeConsumerLagHighis meant to see: the previous behaviour held the last healthy value through an outage, which read as fine and suppressed the alert. Note the flip side: an intermittently failing lookup flaps the series and each gap resetsfor: 30m.