Skip to content

Alerts

TraceLake has a bundle of 30 Prometheus alert rules. Each rule has a severity, a subsystem, a for: duration, and a runbook_url that points to a section of the runbook. The rules read the metrics on the metrics page. The Helm chart can install the bundle as a PrometheusRule.

Labels

Label Values Purpose
severity critical, warning, info Routing. critical pages; warning is a ticket; info never pages.
subsystem ingestor, gateway, compactor, catalog, maintenance Which part of TraceLake owns the response.

subsystem, not component, and that is deliberate. component is already a metric label on the catalog metrics, carrying the emitter (ingestor/compactor/maintenance). A static alert label of the same name silently overwrites it, and TraceLakeCatalogCommitConflicts would have reported its own group instead of naming the writer that was actually conflicting.

Thresholds that mirror TraceLake config

These are set to the configuration defaults and are plain values in the rules file. The Helm chart derives some of them from its values. A deployment that changes the underlying key must re-derive the threshold — nothing detects a changed key for you.

Alert Threshold Mirrors
TraceLakeWalDiskAboveHighWatermark 42949672960 bytes ingestion.wal.diskHighWatermarkPct (80) × the WAL volume size
TraceLakeWalBufferSaturated 214748364 bytes 80% of ingestion.wal.maxBufferedBytes (268435456). Unlike the disk watermark above, the chart re-derives this one from the key, so only the static bundle carries the literal
TraceLakeFreshnessAboveTarget 420 s The published P95 target. It is a fixed target, not a mirror of the sum: the four stages (maxFileAgeMinutes + upload + commitIntervalSeconds + snapshotPollIntervalSeconds) come to 343 s at defaults, leaving 77 s of headroom. Raising any of them spends that headroom and is not re-derived here — at maxFileAgeMinutes' 60-minute ceiling the sum is far past 420 s and this rule fires permanently on a correctly-configured deployment
TraceLakeJwksCacheStale 2880 s 80% of queryGateway.auth.maxJwksStalenessSeconds (3600)
TraceLakeCatalogSnapshotStale 60 s About 7 missed polls at snapshotPollIntervalSeconds: 8
TraceLakeCompactionBacklog 500 files The files-per-partition ceiling
TraceLakeIndexBackfillNotDraining any unindexed file No config key. The for: (36h) is what carries the threshold: it must exceed one compaction.schedule interval, since the backfill merge runs on the floor pass only
TraceLakeCompactorDeadlineExceeded any increase compaction.runtimeDeadlineHours (6), applied by the compactor, not the rule
TraceLakeCompactorEphemeralHigh 5153960755 bytes 80% of the chart's compactor.emptyDir.sizeLimit (6Gi). Not a TraceLake configuration key — the eviction it predicts is the kubelet's, against the volume's own cap — so the chart re-derives it from that value and only the static bundle carries the literal. Its for: (5m) must not outlive compaction.backlogCheckIntervalMinutes (5–120), since the gauge is a per-pass high-water mark the next pass overwrites. It equals that range's floor rather than clearing it: at an interval of 5 the only margin is the pass's own duration
TraceLakeCompactorRunsFailing any failed run in 3h compaction.backlogCheckIntervalMinutes — the lookback, not the threshold, mirrors it: a permanently failing daemon stays up and steps the counter once per tick, so a window shorter than the key's 120-minute ceiling reads 0 between failures and the alert flaps. Raise the lookback past any raised interval
TraceLakeMaintenanceTaskWedged 10800 s 3 missed passes at the default 60 of maintenance.expireSnapshotsIntervalMinutes and maintenance.orphanCleanupIntervalMinutes. Both keys accept 15–1440
TraceLakeTaskWedged 300 s No config key. 10 missed WAL age sweeps (30s) and 20 missed lag refreshes (15s). Both intervals are fixed — only a code change moves this number

TraceLakeWalDiskAboveHighWatermark is the one threshold with no default to inherit: the WAL volume size is a chart value, not a TraceLake config key, so the shipped number is 80% of a 50Gi reference volume. Set it per deployment or the alert is decorative.

TraceLakeWalBufferSaturated mirrors a key that does exist — ingestion.wal.maxBufferedBytes, 80% of it — and the chart re-derives the threshold from that key rather than shipping the static number, so raising the budget moves the alert with it. The static bundle carries 80% of the default.

TraceLakeMaintenanceTaskWedged is the second such case, for the opposite reason: its two keys have defaults, but both accept 15–1440, so a deployment that raises either above 60 minutes will see this fire permanently on a healthy compactor. Re-derive it from the configured interval.

Ingestion

Alert Expression, in words for: Severity Meaning
TraceLakeConsumerLagHigh lag > 1M on any partition 30m warning The ingestor is not keeping up. Check WAL disk first — polling pauses at the high watermark, so a full volume causes lag rather than resulting from it. → runbook
TraceLakeWalDiskAboveHighWatermark WAL bytes past the watermark 5m critical Kafka polling is paused. The drain is upload+commit, so check upload errors and catalog reachability. → runbook
TraceLakeWalBufferSaturated buffered records > 80% of maxBufferedBytes 30m warning The per-dataset record buffer is at its budget. Memory is bounded and nothing is failing — what is being paid is flush batch size, because the partitionMapping has opened enough segments that 1000 rows each no longer fit. Not backpressure. → runbook
TraceLakeUploadBacklog any ready_* file pending 30m warning Rotated WAL files are not reaching cloud storage, and the volume fills behind them. → runbook
TraceLakeUploadErrors upload error rate > 0 15m warning Uploads retry internally, so a blip is survivable; a sustained rate means the backlog grows faster than it drains. → runbook
TraceLakeConsumeErrors any increase over 1h 5m warning recv() errors arrive once per state change — a deleted topic produces a single error, so a rate-based rule would never fire. → runbook
TraceLakeReadyCleanupSlow cleanup P95 > 1s 30m warning WAL unlinks are blocking the commit path far past the sub-millisecond norm. Read with wal_ready_files to tell a large batch from a slow volume. → runbook
TraceLakeFreshnessAboveTarget freshness P95 > 420s 30m warning The published target is breached. The lag is in one of four stages: rotation, upload, commit batching, or the gateway poll. → runbook
TraceLakeRecordsDropped drop rate > 0 30m warning Payloads are failing to decode and are being dropped by the drain policy. Data is being lost silently — usually a producer-side schema change nobody announced. → runbook
TraceLakeWalSegmentOverfilled any increase over 1h 5m critical A WAL file closed holding more rows than its own Kafka offset range can contain — the same offsets were appended twice and the duplicate rows are already on disk, bound for Iceberg. No threshold to tune and no config key to mirror: a segment covering start..=end holds at most end - start + 1 rows, so every non-zero value is corruption rather than a tolerance being exceeded. increase, not rate, for TraceLakeConsumeErrors' reason — one overfilled segment is a single step. Its sibling wal_duplicate_records_total counts duplicates the append guard prevented and is deliberately not alerted on. Clearing is not repair: increase[1h] ages the alert out an hour after the last step, while the duplicate rows are permanent — resolution means no new occurrence. → runbook
TraceLakeSchemaNameDrift renamed columns > 0 1h warning parquetSchema still names a column the live table renamed. Boot tolerates it (identity is the field-id), but new data files carry the stale name. Not an outage; update config and roll. → runbook

Query gateway

Alert Expression, in words for: Severity Meaning
TraceLakeIndexCacheHitRatioLow windowed hit ratio < 0.8 30m warning Queries are re-reading sidecars from object storage. Raise indexCache.maxBytes, or accept that the working set outgrew it. → runbook
TraceLakeSidecarMissRatioHigh > 20% of scanned files have no puffin/idx sidecar 30m warning Pruning is falling back to full scans. Sidecars are built at compaction only, so this usually means the compactor is behind. → runbook
TraceLakeManifestCacheThrash miss rate > 0 on a warm pod 1h warning The load→evict→load loop: the windows queries keep asking for do not fit under manifestCache.maxBytes (manifest lists + loaded manifests), and each query pays their manifest GETs again. → runbook
TraceLakeCatalogSnapshotStale snapshot age > 60s 10m critical A catalog outage. Queries keep serving the last good snapshot — they go stale, not down, so new data is invisible without anything erroring. The exception is a window whose manifests this pod never loaded (an older range, or any query on a pod restarted during the outage): that is a 503. → runbook
TraceLakeJwksCacheStale JWKS age > 80% of the cutoff 10m critical The window in which the OIDC provider can still be fixed. Past the cutoff, every query becomes a 503. → runbook
TraceLakePostgresTableResolveFailures resolve failure rate > 0 15m warning While one dataset fails to resolve, information_schema.columns fails for every other dataset too — one \d resolves a snapshot per configured dataset. → runbook
TraceLakePuffinReadAmplification consulted/fetched < 0.4 6h info The only alert whose remedy is a code change, not a config change — a standing review trigger, never a page. → runbook
TraceLakeIdxPruningIneffective pass-3 out/in > 0.9 1h warning The .idx pass is fetching a sidecar per candidate file and keeping over 90% of them. Break down by idx_declines_total{reason} before acting: only unreadable is a defect, the rest are workload, staleness, or a value cap. → runbook

Expression choices that look wrong and are not

The hit-ratio rule does not use tracelake_gateway_index_cache_hit_ratio. That gauge is a lifetime ratio; a pod that ran hot for a week hides a cold hour inside its own average, so no windowed alert can use it. The rule divides the two counters over a 30m window instead. On an idle gateway that is 0/0 → NaN, and a NaN comparison is false — silence on no traffic is intended, not a gap. The gauge remains useful on a dashboard, which is what it is for.

The manifest-cache rule alerts on the miss rate, not a hit ratio. The poller keeps every configured dataset warm, so steady state is ~0 misses and a ratio would sit pinned at 100% until it dipped for benign reasons. The 1h for: is load-bearing: every pod legitimately misses at boot, and again the first time a query reaches a window that the pod has not loaded, so a shorter window would fire on every rollout.

The sidecar-miss rule selects kind, it does not sum the family. The counter also carries kind="ndv", which is the planner estimating from Iceberg column stats instead of a distinct-value sketch — that reads no extra byte, so folding it into a read-amplification ratio would measure something the rule is not about, and would sit pinned above the threshold on every dataset that configures no ndvColumns. Hence {kind=~"puffin|idx"}: those two, and only those, mean the file is scanned.

0.4 on the Puffin rule is measured, not guessed. With one blob consulted out of N the ratio reads ≈ 1/N, and a measurement showed that the ranged-read path only models a win at N ≥ 3 (≈ 0.33), and then only at wide fan-out. 0.4 sits between the 3-column ratio and the 2-column one (0.5), which lands on the 20% margin — inside the error bar of the assumed RTT. The .idx is deliberately and permanently absent from this rule: 95.2% of that container is what a query already reads, leaving 4.8% to skip across 15 scattered ranges.

The pass-3 rule divides two counters that a declining pass never touches. The pass records neither leg when it declines wholesale, so a workload the index is not asked to serve leaves both series flat and the ratio undefined — silence, exactly as on the idle-gateway case above. Recording in == out for a pass that did not run would have made every such workload fire. The counters carry a pass label: divide them within one pass, and do not add them across passes.

Maintenance and compaction

Alert Expression, in words for: Severity Meaning
TraceLakeCatalogCommitConflicts conflict rate > 0.05/s 30m warning Writers are serialising against each other. A conflict is a retry, not a loss — an occasional one is the design working. → runbook
TraceLakeCompactionBacklog ≥ 500 files in a partition 1h warning Rare by design. Sustained firing means the compactor fleet is undersized, not that the rule is noisy. → runbook
TraceLakeIndexBackfillNotDraining any data file with no .idx 36h warning Text predicates on the partition full-scan. Results stay correct (a missing sidecar is scanned, never skipped) — this is latency, not data loss. A non-zero reading between floor passes is normal, which is what the long for: absorbs. → runbook
TraceLakeCompactorRunsFailing any failed run in 3h 5m critical Compaction is the only hand-rolled Iceberg operation TraceLake owns. Treat a repeated failure as a correctness risk, not just a backlog one. → runbook
TraceLakeCompactorScrapeDown up == 0 15m critical The pod itself is gone — boot failure, OOM, eviction, node loss. A failed pass does not land here: the daemon stays up and reports through TraceLakeCompactorRunsFailing. While the pod is down nothing compacts, nothing expires, and no orphans are swept. → runbook
TraceLakeCompactorDeadlineExceeded any increase in 6h 5m warning A run passed runtimeDeadlineHours. The deadline is an alert threshold, not a cap — the run continues, competing with the next one. → runbook
TraceLakeCompactorEphemeralHigh a pass peaked > 80% of the /tmp volume's sizeLimit 5m warning The next merge of that size is an eviction taken mid-merge: the pass is discarded and the next one repeats it. Peak is the merge spill set plus the Tantivy scratch of the output being written, and on a text-heavy dataset the index term is the larger of the two (measured 11× the data it indexes). → runbook
TraceLakeOrphanDeletionSpike > 1000 deletions in 1h 5m critical Orphan cleanup physically deletes objects and is the one maintenance path that can lose data. A normal sweep deletes a handful. → runbook
TraceLakeBackgroundTaskDied a pod's task-exit counter is non-zero 5m critical A supervised loop panicked. It stays dead for the life of the pod — the supervisor reports, it does not restart — so the rule fires until that pod is replaced, not for a window after the death. Counts panics only: a deliberate abort and a clean return are teardown, so a rolling restart never fires this. → runbook
TraceLakeTaskWedged a wal_age_sweep, kafka_lag_refresh or gateway_snapshot_poll stamp is > 5 min stale 5m critical The other half of the above: the loop is alive but stuck and never panicked, so the exit counter stays flat. A wedged wal_age_sweep makes maxFileAgeMinutes untrue for that dataset's idle partitions; a wedged kafka_lag_refresh freezes consumer_lag so TraceLakeConsumerLagHigh cannot fire; a wedged gateway_snapshot_poll freezes catalog_snapshot_age_seconds so TraceLakeCatalogSnapshotStale cannot fire either. Total wedge age at page time is threshold + for: = 10 min, which is what the annotation says. → runbook
TraceLakeMaintenanceTaskWedged a snapshot_expiry or orphan_cleanup stamp is > 3 h stale 15m critical Same signal at the maintenance cadence, which is three orders of magnitude longer and cannot share a threshold. Neither loop has any other signal behind a wedge. → runbook

TraceLakeCompactorScrapeDown reads up{job="tracelake-compactor"}; the Helm chart sets that job label. Note that the compactor binds its metrics endpoint before connecting the catalog, precisely so an unreachable catalog still scrapes rather than surfacing here as a dead pod. Manually scoped recovery runs (ingestor compact --dataset/--partition/--orphan-dry-run) do not serve metrics at all — they are one-shot and exit — so no rule covers them.

TraceLakeBackgroundTaskDied is filed here because the two loops it exists for are snapshot_expiry and orphan_cleanup — two of the four with no consequence alert behind them (nothing alerts on snapshots_expired_total being flat, and the orphan rule catches too many deletions, never zero); wal_age_sweep and sigterm_handler are the other two. It covers the remaining nine too, where it fires earlier and names the dead task rather than the symptom: today a dead upload_scan surfaces as TraceLakeUploadBacklog and a dead gateway_snapshot_poll as a rising catalog_snapshot_age_seconds. task is the loop name, not the dataset — see the metrics page.

It reads the raw counter, not an increase() window, and is deliberately not aggregated. The condition is "this pod is running with a dead loop", which persists until the pod is replaced; any window would resolve the alert with the loop still gone, which is the silence the rule exists to remove. The counter resets on process restart, so the replacement pod clears it. Per-series because the remedy is to delete one named pod — sum by (task) would drop the label that names it.

It covers every long-lived loop in the fleet: ten async tasks and the ingestor's three WAL threads (wal_drain, wal_age_sweep, wal_backpressure). Two of the loops had no alert of any kind behind them: wal_age_sweep is the only thing that finalizes an idle partition's segment, and a dead sigterm_handler leaves the pod ignoring SIGTERM until the kubelet's SIGKILL, which nothing observes until the next rollout.

Routing caveat: every task carries subsystem: maintenance, so an Alertmanager route keyed on subsystem sends kafka_poll, kafka_lag_refresh, batched_commit, upload_scan, gateway_snapshot_poll, wal_drain, wal_age_sweep, wal_backpressure and sigterm_handler to the maintenance owner — ingestion and pod-lifecycle failures included. The task label tells the human what actually died; it does not steer the router.

TraceLakeTaskWedged and TraceLakeMaintenanceTaskWedged are the other half of that counter, and they take the same routing caveat. Neither reads task_exits_total — a wedged loop never panicked, so nothing increments and the death rule is silent by construction. They read tracelake_task_last_pass_timestamp_seconds, which only five of the thirteen loops carry: wal_age_sweep, kafka_lag_refresh, gateway_snapshot_poll, snapshot_expiry and orphan_cleanup. The other eight are omitted deliberately — four are event-driven, where a stale stamp is a healthy idle pod rather than a fault, and four already have a consequence that shows behind a wedge. The test that sorts them: is the gauge the consequence alert reads written outside the loop that would wedge? gateway_snapshot_poll is stamped because it fails that test — the poll republishes catalog_snapshot_age_seconds at the end of its own pass, so a wedge freezes that gauge rather than letting it rise, and the gauge starts at 0, so a poller wedged from boot publishes a healthy 0 forever.

Two rules rather than one because the five loops fall into two cadences three orders of magnitude apart: 30s/15s loops against 60-minute maintenance loops. A single threshold either pages on every normal maintenance gap or takes hours to notice a wedged sweeper. They share one runbook anchor, because the remedy — replace the pod — does not differ.

Series that may legitimately be absent

Four families can be missing from a scrape entirely. Every rule that touches them alerts on > threshold, so absence and zero mean the same thing and no absent() guard is required — but a rule added later must not read absence as a value.

  • The four pg_* families register only when queryGateway.postgres is configured. A gateway without the second door exposes none of them.
  • tracelake_catalog_commit_conflicts_total is never seeded: its dataset label is unknown at registration, so it appears only once a conflict has actually happened.
  • tracelake_compactor_sidecar_bytes is deliberately not seeded either, and unlike the two above it can be absent on a healthy deployment: .idx and Puffin sidecars are written at compaction only, so the family has no series until a pass writes an object, and a manually scoped run (--dataset / --partition / --orphan-dry-run) registers the collector without ever serving /metrics. A pre-seeded histogram would publish a zero-byte ratio, which reads as a fact rather than as silence. No rule in this bundle reads it — it is a capacity signal on the compactor dashboard row, and that row being empty between runs is expected.
  • tracelake_ingestor_consumer_lag, per dataset/partition. It is not seeded, and unlike the three above it is also removed: the refresh loop publishes exactly the partitions it obtained a reading for and drops the rest. A missing partition means no reading was obtained — a committed() timeout, a fetch_watermarks failure, a blocking-pool stall, a partition revoked to another pod, a pod with no assignment at all, or a partition that has not committed an offset yet (routine for the first minutes after a pod joins) — and never lag zero. That is what TraceLakeConsumerLagHigh is meant to see: the previous behaviour held the last healthy value through an outage, which read as fine and suppressed the alert. Note the flip side: an intermittently failing lookup flaps the series and each gap resets for: 30m.