Metrics¶
Each TraceLake process serves its metrics in the Prometheus text format at GET /metrics on
metrics.port (default 9090). This page lists each metric with its type, its labels and its
meaning.
| Process | Command | Metrics |
|---|---|---|
| Ingestor | ingestor run |
Ingestor, background tasks, catalog |
| Compactor | ingestor compact |
Compactor, background tasks, catalog |
| Query gateway | gateway |
Query gateway, background tasks, catalog |
The compactor has its own pod and its own scrape target. A manual compaction run (with --dataset,
--date-range, --partition, --dry-run or --orphan-dry-run) does not serve /metrics.
How to read the tables¶
- Absent is not zero. A series that Prometheus has never seen is absent from the scrape. Most
metrics with a
datasetlabel start at zero for each configured dataset when the process starts. The tables say "not seeded" where a series appears only after the first event. An expression on such a series must tolerate its absence. - Counters start again at zero when a pod starts.
- Buckets are the upper limits of the histogram buckets, in the unit of the metric.
Ingestor¶
| Metric | Type | Labels | Meaning |
|---|---|---|---|
tracelake_ingestor_fetched_records_total |
Counter | dataset |
Records fetched from Kafka, before decode and before the WAL. Includes records that Kafka delivers again. |
tracelake_ingestor_fetched_bytes_total |
Counter | dataset |
Payload bytes of the fetched records. |
tracelake_ingestor_events_total |
Counter | dataset |
Records that the WAL accepted into a segment. This is the measure of ingested volume. |
tracelake_ingestor_bytes_total |
Counter | dataset |
Payload bytes (not compressed) of the records that the WAL accepted. |
tracelake_ingestor_consumer_lag |
Gauge | dataset, partition |
Kafka consumer lag: the end offset of the log minus the committed offset, for each assigned partition. Not seeded. A partition with no reading has no series. |
tracelake_ingestor_consume_errors_total |
Counter | dataset |
Kafka consume attempts that failed. The ingestor waits and tries again. |
tracelake_ingestor_records_dropped_total |
Counter | dataset |
Records dropped because they could not be decoded. |
tracelake_ingestor_wal_active_files |
Gauge | dataset |
Open active_* WAL files. |
tracelake_ingestor_wal_ready_files |
Gauge | dataset |
ready_* WAL files that wait for upload and commit. |
tracelake_ingestor_wal_disk_bytes |
Gauge | none | Approximate bytes of all WAL files on the WAL volume. All datasets share the volume. |
tracelake_ingestor_wal_buffered_bytes |
Gauge | dataset |
Decoded records in memory for open segments. The limit is ingestion.wal.maxBufferedBytes. The real memory use is higher than this value. |
tracelake_ingestor_wal_rowgroup_bytes |
Gauge | dataset |
Parquet row groups in memory for open segments. |
tracelake_ingestor_wal_duplicate_records_total |
Counter | dataset |
Records that the ingestor discarded because it fetched them before it released their partition. The new owner of the partition reads them again. An increase at a rebalance is normal. |
tracelake_ingestor_wal_segment_overfilled_total |
Counter | dataset |
WAL files closed with more rows than their offset range can contain. This must stay at zero. A value above zero means that a file contains duplicate rows. |
tracelake_ingestor_upload_latency_seconds |
Histogram | dataset |
Time to upload one ready_* file. Buckets: 1, 5, 10, 15, 30, 60. |
tracelake_ingestor_upload_errors_total |
Counter | dataset |
Upload attempts that failed. |
tracelake_ingestor_ready_cleanup_seconds |
Histogram | dataset |
Time to delete the local files of one committed batch. Buckets: 0.001, 0.01, 0.05, 0.25, 1, 5. |
tracelake_ingestor_drain_total |
Counter | outcome |
Shutdown drains. outcome is complete or deadline. |
tracelake_ingestor_drain_seconds |
Histogram | none | Duration of the shutdown drain. Not seeded. Buckets: 1, 5, 15, 30, 60, 96, 300, 600. |
tracelake_ingestor_schema_name_drift |
Gauge | dataset |
Columns in parquetSchema that have a different name in the table. The ingestor starts, but it writes the old names into new files. |
tracelake_freshness_seconds |
Histogram | dataset |
Time from the arrival of a record at Kafka to the table commit that contains it. Buckets: 30, 60, 120, 240, 360, 480, 600. |
Fetched and accepted¶
fetched_* counts what the pod reads from Kafka. events_total and bytes_total count what the
WAL accepts. The difference has three parts, and each is a metric or is temporary:
fetched_records_total - events_total
= wal_duplicate_records_total (fetched, then discarded at a rebalance)
+ records_dropped_total (fetched, then not decoded)
+ records in progress (between the fetch and the WAL)
fetched_*stops andconsumer_lagincreases: the Kafka side has stopped.fetched_*continues andevents_totalstops: the WAL has stopped.
Do not use events_total to find out if consumption is alive.
Ingested volume is not exact after a crash¶
events_total and bytes_total count a record when the WAL accepts it, not when the table commit
contains it. If recovery discards a WAL file and Kafka delivers its records again, the ingestor
counts those records a second time. This occurs only when recovery discards a file, not at each
rebalance. No metric shows it; the ingestor writes each discarded file to its log. For a period
where the exact number is important, compare with the row counts of the Iceberg table.
Shutdown metrics¶
drain_total and drain_seconds change one time, at the end of the drain, and then the process
exits. Prometheus will usually not collect the last value. The log line of the drain is the
reliable record.
Query gateway¶
| Metric | Type | Labels | Meaning |
|---|---|---|---|
tracelake_gateway_query_requests_total |
Counter | dataset, method, format, status, error_code |
HTTP query requests. Only the success series (status="200", empty error_code) is seeded. |
tracelake_gateway_query_duration_seconds |
Histogram | dataset, method, format |
Time from the first byte of the request to the end of the response body. Buckets: 0.05, 0.2, 0.5, 1, 3, 5, 10, 30. |
tracelake_gateway_dataset_fenced |
Gauge | dataset |
1 while queryGateway.auth.datasetRoles denies the dataset to all callers; 0 if not. |
tracelake_gateway_manifest_files_scanned |
Histogram | dataset |
Data files that remain after pruning pass 1, for each query. Buckets: 1, 4, 16, 64, 256, 1024, 4096, 16384. |
tracelake_gateway_files_scanned |
Histogram | dataset |
Data files that a query opens, after all three pruning passes. Same buckets. |
tracelake_gateway_puffin_bytes_total |
Counter | dataset, part |
Bloom-filter sidecar bytes. part is fetched (read from storage) or consulted (used by a query). |
tracelake_gateway_datafile_requests_total |
Counter | dataset, op |
Storage requests of the data-file scan, for HTTP queries only. op is get or list. list must stay at zero. |
tracelake_gateway_datafile_bytes_total |
Counter | dataset |
Bytes that the data-file scan reads from storage, for HTTP queries only. |
tracelake_gateway_sidecar_fetch_seconds |
Histogram | kind |
Time to read one sidecar from storage on a cache miss. Buckets: 0.005, 0.01, 0.025, 0.05, 0.1, 0.25, 0.5, 1, 2.5, 5. |
tracelake_gateway_sidecar_miss_total |
Counter | dataset, kind |
Sidecar lookups with no result. kind is puffin, idx or ndv. See below. |
tracelake_gateway_prune_candidates_in_total |
Counter | dataset, pass |
Candidate files that enter a pruning pass. |
tracelake_gateway_prune_candidates_out_total |
Counter | dataset, pass |
Candidate files that remain after a pruning pass. |
tracelake_gateway_idx_declines_total |
Counter | dataset, reason |
Full-text sidecars that the gateway read but that could not answer the filter. |
tracelake_gateway_bloom_unserved_total |
Counter | dataset, reason |
Filters on a bloom column that the bloom pass could not use. One for each filter in each query. |
tracelake_gateway_index_cache_hits_total |
Counter | none | Sidecar lookups answered from the sidecar cache. |
tracelake_gateway_index_cache_misses_total |
Counter | none | Sidecar lookups that read storage. |
tracelake_gateway_index_cache_bytes |
Gauge | none | Bytes in the sidecar cache. The limit is queryGateway.indexCache.maxBytes. |
tracelake_gateway_index_cache_hit_ratio |
Gauge | none | Fraction of sidecar lookups answered from the cache, 0 to 1, since the pod started. For a time window, use the two counters. |
tracelake_gateway_manifest_cache_hits_total |
Counter | dataset |
Queries answered from cached table metadata. |
tracelake_gateway_manifest_cache_misses_total |
Counter | dataset |
Queries that read manifests from storage. |
tracelake_gateway_manifest_cache_bytes |
Gauge | dataset |
Bytes of cached table metadata. The limit, queryGateway.manifestCache.maxBytes, applies to the sum of all datasets. |
tracelake_gateway_jwks_cache_age_seconds |
Gauge | none | Time since the last successful fetch of the OIDC signing keys, or since the pod started if there is none. When the value gets near queryGateway.auth.maxJwksStalenessSeconds, requests will soon get 503. |
tracelake_gateway_jwks_fetch_duration_seconds |
Histogram | outcome |
Duration of a fetch of the signing keys. outcome is ok or error. Buckets: 0.01, 0.05, 0.1, 0.25, 0.5, 1, 2, 5, 10. |
Label values are fixed sets. A caller cannot add a value.
| Label | Values |
|---|---|
method |
GET, POST, other |
format |
rows |
error_code |
Empty on success. If not, the error code of the response, for example FORBIDDEN. |
pass |
partition, manifest, bloom, idx, idx_rowgroups |
reason on idx_declines_total |
no_field, pattern_too_short, no_rowgroups, capped, unreadable |
reason on bloom_unserved_total |
type, function, not, negation, null, range, like, nested_and, sibling_ruled_out_by_bounds, sibling_unanswerable |
Errors after the response starts¶
A query that fails after the gateway sent the status 200 is counted with status="200" and an
error_code that is not empty. Write an error-rate expression on error_code!="", not on status.
Refused for concurrency¶
query_requests_total{status="503",error_code="QUERY_CONCURRENCY_EXCEEDED"} counts the queries that
the gateway refused at queryGateway.maxInFlightHttpQueries. A continuous rate means that the pod
is at its limit. Add pods, or increase the limit if the pod has the memory.
The pruning funnel¶
manifest_files_scanned and files_scanned are the two ends of the funnel: the files after pass 1,
and the files that the scan opens.
For one value of pass, prune_candidates_out_total / prune_candidates_in_total is the fraction
of files that the pass keeps. Do not add the counters of different passes together. The partition
value is recorded only when the filter of a query limits the partition column. For
pass="idx_rowgroups" the counters count row groups, not files.
Sidecar misses¶
sidecar_miss_total has two meanings:
kind="puffin"andkind="idx": a pruning pass had no sidecar to use, and the gateway scans the file. This increases the data that a query reads.kind="ndv": the planner had no distinct-value sketch for a file and used the column statistics. The query reads no more data.
An increase is not an error. Files that are not compacted yet have no full-text sidecar. A high rate with slow queries means that compaction is behind, or that a column needs an index.
idx_declines_total is a part of sidecar_miss_total{kind="idx"}: the sidecar is there, but it
cannot answer the filter.
PostgreSQL interface¶
These metrics exist only when queryGateway.postgres is configured.
| Metric | Type | Labels | Meaning |
|---|---|---|---|
tracelake_gateway_pg_connections |
Gauge | none | Open PostgreSQL connections. |
tracelake_gateway_pg_queries_total |
Counter | status |
Statements. status is ok, denied, timeout, unavailable or failed. |
tracelake_gateway_pg_query_duration_seconds |
Histogram | none | Duration of a statement. Buckets: the Prometheus defaults. |
tracelake_gateway_pg_table_resolve_failures_total |
Counter | dataset |
Statements for which the gateway could not get the snapshot of a table. |
denied includes authorization refusals, limit refusals, unknown tables and statements that do not
parse. It is above zero in a healthy deployment. failed is the value that shows a fault.
Compactor¶
| Metric | Type | Labels | Meaning |
|---|---|---|---|
tracelake_compactor_partition_files_total |
Gauge | dataset, partition |
Small files (below compaction.minFileSizeMb) in each partition of the current snapshot. |
tracelake_compactor_unindexed_files_total |
Gauge | dataset, partition |
Data files with no full-text sidecar in each partition. It counts only files that TraceLake wrote and that are not larger than compaction.targetSizeMb. Absent for a dataset with no tantivy.indexingColumns. |
tracelake_compactor_runs_total |
Counter | dataset, status |
Compaction runs. status is success or failed. |
tracelake_compactor_duration_seconds |
Histogram | dataset |
Duration of a compaction run for one dataset. Buckets: 60, 300, 900, 1800, 3600. |
tracelake_compactor_deadline_exceeded_total |
Counter | dataset |
Passes that ran longer than compaction.runtimeDeadlineHours. |
tracelake_compactor_data_files_rewritten_total |
Counter | dataset |
Data files that a merge commit replaced. |
tracelake_compactor_sidecar_bytes |
Histogram | dataset, kind |
Size of each object that a pass writes. kind is data, idx or puffin. Not seeded. Buckets: 1e4, 1e5, 1e6, 1e7, 1e8, 1e9. |
tracelake_compactor_idx_build_failures_total |
Counter | dataset, stage |
Merged files committed with no full-text sidecar because the build or the upload failed. stage is start, batch, finish or upload. |
tracelake_compactor_backfill_detect_seconds |
Histogram | dataset |
Duration of the search for files with no full-text sidecar. Not seeded. Buckets: 0.01, 0.025, 0.05, 0.1, 0.5, 1, 5, 30, 120, 300, 900. |
tracelake_compactor_ephemeral_peak_bytes |
Gauge | dataset |
Largest temporary disk use of the most recent pass. Use it to set the size of the temporary volume of the pod. |
tracelake_snapshots_expired_total |
Counter | dataset |
Snapshots that snapshot expiry removed. |
tracelake_orphan_files_deleted_total |
Counter | dataset |
Objects that orphan cleanup deleted. |
unindexed_files_total above zero between scheduled passes is normal: the compactor counts the
files at each pass, but it gives them an index only on a scheduled pass. If the value does not
decrease after the scheduled passes, look at idx_build_failures_total. An increase there means
that the index cannot be built, and more compactor capacity will not help.
The two metrics with a partition label keep one series for each partition until the pod starts
again. With one partition for each day, that is approximately 365 series for each dataset, for each
metric, for each year of retention.
Background tasks¶
All three processes have these metrics.
| Metric | Type | Labels | Meaning |
|---|---|---|---|
tracelake_task_exits_total |
Counter | task |
Background tasks that stopped because of a panic. A value above zero means that this pod runs with a dead task. Replace the pod. |
tracelake_task_last_pass_timestamp_seconds |
Gauge | task, dataset |
Unix time of the end of the last pass of a periodic task. If time() minus this value increases without limit, the task has stopped without an exit. |
A stop on purpose, such as a shutdown, is not counted in task_exits_total.
Five tasks have a time stamp: wal_age_sweep and kafka_lag_refresh in the ingestor,
snapshot_expiry and orphan_cleanup in the compactor, and gateway_snapshot_poll in the
gateway. The tasks of the compactor and the gateway have dataset="-".
Catalog¶
All processes that use the Iceberg catalog have these metrics.
| Metric | Type | Labels | Meaning |
|---|---|---|---|
tracelake_catalog_snapshot_age_seconds |
Gauge | dataset |
Time since the gateway last got the snapshot pointer from the catalog. Read it on the gateway. |
tracelake_catalog_commit_latency_seconds |
Histogram | component |
Duration of a catalog commit. Buckets: 0.1, 0.5, 1, 5, 15, 60. |
tracelake_catalog_commit_conflicts_total |
Counter | dataset, component |
Commits that the catalog refused because another writer committed first. The writer tries again. Not seeded. |
component is ingestor, compactor or maintenance. The gateway makes no commits.