Configuration¶
One YAML file configures a deployment. The ingestor, the compactor and the query gateway read the same file. To add a data source, you add a dataset to the file; you do not change code.
Rules for the file¶
- The file is examined at start. A value that is out of its range, or a combination that is not permitted, stops the process with a message that gives the full path of the key.
- A key that is not known is an error. A key with a typing error does not get its default
silently. The exceptions are
storage.catalogAuth,storage.propertiesandingestion.kafka.properties. TraceLake does not examine the names of the keys in these maps. ${VAR}in a value is replaced with the environment variable of that name. Use it for secrets. If the variable is not set, the process does not start. There is no$VARform and no default form. Write$${for a literal${. Keys and comments are not changed.- Error messages do not show the value of a variable. They show
${VAR}in its place.
The tables below give the default and the permitted range of each key. "Required" means that the key has no default.
cluster¶
| Key | Default | Meaning |
|---|---|---|
cluster.id |
Required | A name for the deployment. No process uses the value in this version. |
ingestion¶
| Key | Default | Range | Meaning |
|---|---|---|---|
ingestion.kafka.bootstrapServers |
Required | The Kafka brokers, as host:port with commas between them. |
|
ingestion.kafka.properties |
none | Properties for the Kafka client (librdkafka), for example for TLS and SASL. TraceLake sets bootstrap.servers, group.id, enable.auto.commit, auto.offset.reset (earliest) and partition.assignment.strategy itself, after these. A value that you give for one of them has no effect. |
|
ingestion.namespace |
tracelake |
The Iceberg namespace of all the tables. The name of a table is the name of its dataset. | |
ingestion.commitIntervalSeconds |
20 | 15 to 30 | The interval between table commits of the ingestor. |
ingestion.drainTimeoutSeconds |
96 | 5 to 600 | The time limit of the shutdown of the ingestor. It must be less than the termination grace period of the pod. |
ingestion.datasets |
none | The list of datasets. See Datasets. |
ingestion.wal¶
| Key | Default | Range | Meaning |
|---|---|---|---|
ingestion.wal.localDirectory |
Required | The directory of the write-ahead log, on the volume of the pod. | |
ingestion.wal.maxFileSizeBytes |
67108864 (64 MiB) | 16 MiB to 256 MiB | A WAL file rotates at this size. |
ingestion.wal.maxFileAgeMinutes |
5 | 1 to 60 | A WAL file rotates at this age. |
ingestion.wal.maxOpenSegmentsPerPartition |
256 | 1 to 65536 | The largest number of open WAL files for one Kafka partition. Above it, all the files of the partition rotate. |
ingestion.wal.maxBufferedBytes |
268435456 (256 MiB) | 16 MiB to 16 GiB | The limit of the decoded records that one dataset keeps in memory for its open WAL files. At the limit the ingestor writes them to the files. It does not pause ingestion. |
ingestion.wal.diskHighWatermarkPct |
80 | 50 to 95 | The ingestor pauses its Kafka consumers when the WAL volume is this full. It resumes them 5 percentage points below. |
A dataset can have its own value for the first four numeric keys (see Overrides for one dataset). The directory and the watermark apply to the whole ingestor, because all datasets share one volume.
Each open WAL file uses one file descriptor. At start, the ingestor compares the largest possible number of open files with the limit of the process and writes a warning to its log if the limit is too low.
ingestion.tantivy¶
These keys control the full-text index that the compactor builds.
| Key | Default | Range | Meaning |
|---|---|---|---|
ingestion.tantivy.tokenizerMemoryBudgetMb |
16 | 15 to 128 | The memory of the index writer. |
ingestion.tantivy.ngramSize |
3 | 2 to 5 | The length of the character groups that answer LIKE filters. A pattern with no literal part of this length or more cannot use the index; the gateway then scans the file. A larger value makes the index smaller. |
ingestion.tantivy.maxIndexedValueBytes |
65530 | 16 to 65530 | The longest value that goes into the index. The table keeps the full value. For a longer value, the gateway scans the file for a filter on that column. |
Datasets¶
One entry in ingestion.datasets is one Kafka topic and one Iceberg table.
| Key | Default | Range | Meaning |
|---|---|---|---|
name |
Required | The name of the dataset and of its table. | |
topic |
Required | The Kafka topic. | |
groupId |
Required | The Kafka consumer group. Use one group for each dataset. | |
parquetSchema |
none | The columns and their types. See Schema. | |
customMapping |
none | For each column, where its value comes from in the JSON payload. See Mapping. | |
partitionMapping |
none | The partition columns. See Partitions. | |
bloomFilterColumns |
none | Columns that get a bloom filter. Use them for identifiers that queries look up with an equality filter. | |
tantivy.indexingColumns |
none | Text columns that get a full-text index at compaction. | |
ndvColumns |
none | Columns that get a distinct-value sketch at compaction. The query planner uses it. | |
sortColumns |
none | The compactor sorts the rows of each merged file by these columns. | |
maxRawWindowDays |
7 | 0 to 3650 | The longest time range of a query that returns raw rows. 0 removes the limit. |
timeColumn |
none | The timestamp column on which the gateway measures that time range. Necessary when maxRawWindowDays is not 0. |
|
wal, compaction |
none | See Overrides for one dataset. | |
preset |
none | Accepted, but it has no effect in this version. |
The ingestor and the compactor do not start if bloomFilterColumns, tantivy.indexingColumns,
ndvColumns or sortColumns contain a column that is not in parquetSchema, or contain a column
two times. They also do not start if bloomFilterColumns contains a BOOLEAN column, or if
sortColumns is set and there is no parquetSchema. The gateway does not
start if a dataset with a time-range limit has no timeColumn, or if that column is not a
timestamp column.
Schema¶
parquetSchema is a list of name TYPE pairs with commas between them:
parquetSchema: "span_id STRING, service_name STRING, message STRING, ts TIMESTAMPTZ, event_date LONG"
| Type | Holds |
|---|---|
STRING |
Text. UUID is accepted and stored as STRING. |
INT, LONG |
Integers. BIGINT is the same as LONG. |
DOUBLE |
Floating-point numbers. |
BOOLEAN |
True or false. BOOL is also accepted. |
TIMESTAMP |
A date and time with no time zone. |
TIMESTAMPTZ |
An instant in UTC. Use it for event times. |
The names of the types are not case-sensitive.
The order of the columns is permanent. The position of a column gives it its identity in the Iceberg table. Add new columns at the end. Do not change the order and do not remove a column.
Each top-level field of the payload whose name is not the name of a column goes into the column
_unmapped, as JSON text. This includes a field that a customMapping expression reads for a
column with a different name. A value that does not convert to the type of its column also goes
there. The record is not lost. To keep a field out of _unmapped, give the column the same name as
the field. You cannot use the name _unmapped for a column of your own.
In the example on this page, spanId, startTimeUnixNano and the full resource object are thus
in _unmapped for each row, in addition to the columns made from them. Include this in your
estimate of storage.
With no parquetSchema, the table has only the _unmapped column.
If parquetSchema has a different type for a column than the table has, the ingestor does not
start. If the partition columns are different from those of the table, the ingestor also does not
start. Each check has a command-line flag that permits the start (--allow-incompatible-schema,
--allow-partition-evolution). Use one only for a change that you planned.
Mapping¶
customMapping gives a JMESPath expression for each column. The short
form is the expression only. The long form adds a conversion.
customMapping:
span_id: "spanId"
service_name: "resource.attributes.\"service.name\""
ts:
expr: "startTimeUnixNano"
normalize: to_micros
epochUnit: nanos
event_date:
expr: "startTimeUnixNano"
normalize: date_int
epochUnit: nanos
calendarOffset: "+02:00"
| Key | Values | Meaning |
|---|---|---|
expr |
A JMESPath expression | Where the value is in the payload. |
normalize |
strip_hyphens |
Removes the hyphens, for example from a UUID. |
to_micros |
Converts a time to microseconds since 1970 in UTC. | |
year, month, day |
The part of the date, as a number. | |
date_int |
The date as one number, yyyymmdd, for example 20230722. |
|
epochUnit |
seconds, millis, micros, nanos, datetime_string |
The form of the time in the payload. Necessary with each time conversion. TraceLake does not guess it from the value. |
datetimeFormat |
rfc3339, or a pattern such as "%Y-%m-%d %H:%M:%S%.f" |
The layout of the text. Necessary with epochUnit: datetime_string, and permitted only with it. |
defaultOffset |
A fixed offset such as "+02:00" |
The offset for a text value that has none. The default is UTC. Permitted only with epochUnit: datetime_string. It has no effect with datetimeFormat: rfc3339, because such a value always has its own offset. |
calendarOffset |
A fixed offset such as "+02:00" |
The offset of the calendar for year, month, day and date_int. The default is UTC. Not permitted with to_micros. |
Without calendarOffset, a dataset in a time zone that is not UTC has each local day in two
values of the date column. Only fixed offsets are possible; names of time zones are not.
Warning
A change to calendarOffset on a dataset that has data changes the partition of new rows. It
is a migration, not a change to a setting.
Partitions¶
Each entry of partitionMapping has a key, which is the name of a column, and an expr, which
must be the same expression that customMapping has for that column. The column must be in
customMapping, and its type cannot be DOUBLE.
partitionMapping:
- key: event_date
expr: "startTimeUnixNano"
- key: service_name
expr: "resource.attributes.\"service.name\""
A date_int column is the recommended date partition: it is one column, and one range filter
selects a range of dates. Three columns with year, month and day are also possible.
Do not partition on a column with many different values. Each value of the partition columns on each Kafka partition has its own open WAL file.
Overrides for one dataset¶
A dataset can have its own value for these keys. A key that the dataset does not set has the value from the top level. The ranges are the same.
datasets:
- name: spans
wal:
maxFileAgeMinutes: 2
compaction:
targetSizeMb: 256
| Block | Keys |
|---|---|
wal |
maxFileSizeBytes, maxFileAgeMinutes, maxOpenSegmentsPerPartition, maxBufferedBytes |
compaction |
backlogFileThreshold, minFileSizeMb, targetSizeMb, compression, compressionLevel, rowGroupSize, commitRetryMax, snapshotRetentionHours |
The compaction block of a dataset also accepts schedule and backlogCheckIntervalMinutes, but
they have no effect. The compactor uses the top-level values of these two keys for all datasets.
compaction¶
| Key | Default | Range | Meaning |
|---|---|---|---|
compaction.schedule |
0 2 * * * |
A cron expression, in UTC | The time of the scheduled pass, which selects more than a backlog tick. |
compaction.backlogCheckIntervalMinutes |
15 | 5 to 120 | The interval between passes. |
compaction.backlogFileThreshold |
100 | 1 to 100000 | A backlog tick selects a partition with more small files than this. |
compaction.minFileSizeMb |
32 | 4 to 64 | A file below this size is a small file. |
compaction.targetSizeMb |
128 | 64 to 512 | The size of a merged file. It must be larger than minFileSizeMb. |
compaction.compression |
zstd |
zstd, snappy, lz4, uncompressed |
The compression of all data files, those of the ingestor and those of the compactor. |
compaction.compressionLevel |
1 | 1 to 22 | The level for zstd. With a different compression, do not set it. |
compaction.rowGroupSize |
1048576 | 65536 to 16777216 | The number of rows in a row group of a merged file. The gateway prunes by row group. |
compaction.commitRetryMax |
5 | 1 to 20 | How many times a writer sends a commit again after the catalog refuses it. |
compaction.snapshotRetentionHours |
24 | 1 to 720 | Snapshot expiry keeps the snapshots of this period. It must not be less than maintenance.orphanMinAgeHours. |
compaction.runtimeDeadlineHours |
6 | 1 to 72 | A pass that runs longer increments a counter. The pass does not stop. |
Two combinations give a warning in the log at start, not an error:
ingestion.wal.maxFileSizeBytesis equal to or more thanminFileSizeMb. This is the case with the defaults (64 MiB and 32 MiB). A WAL file that rotates at its maximum size is then not a small file, and the compactor does not merge it for its size. If the dataset hastantivy.indexingColumns, the compactor still merges it to give it a text index.backlogFileThresholdis 1 or 2. The compactor then merges very small partitions again at each pass, with no decrease in the number of files.
maintenance¶
| Key | Default | Range | Meaning |
|---|---|---|---|
maintenance.expireSnapshotsIntervalMinutes |
60 | 15 to 1440 | The interval of snapshot expiry. |
maintenance.orphanCleanupIntervalMinutes |
60 | 15 to 1440 | The interval of orphan cleanup. |
maintenance.orphanMinAgeHours |
24 | 24 to 8760 | Orphan cleanup deletes only objects that are older than this. The minimum protects files that an ingestor has uploaded but not committed. |
storage¶
| Key | Default | Meaning |
|---|---|---|
storage.provider |
AZURE |
The only value in this version. |
storage.destinationUri |
Required | The location of the tables, for example abfss://lake@account.dfs.core.windows.net/tracelake. |
storage.catalogUri |
Required | The URL of the Iceberg REST catalog. |
storage.catalogAuth.clientId |
Required | The OAuth2 client ID for the catalog. |
storage.catalogAuth.clientSecret |
Required | The OAuth2 client secret. Use ${VAR}. |
storage.catalogAuth.scope |
none | The OAuth2 scope. |
storage.catalogAuth.oauth2ServerUri |
{catalogUri}/v1/oauth/tokens |
The token endpoint. |
storage.catalogAuth.warehouse |
none | The warehouse in the catalog. Necessary for the compactor: without it, merge commits and snapshot expiry fail. |
storage.properties |
none | Properties for the storage client, for example the storage account and its credentials. The Iceberg client that writes the table metadata gets the same properties. |
TraceLake does not connect to a catalog without credentials. It gets a token with the OAuth2 client-credentials grant and gets a new one before the token expires.
queryGateway¶
| Key | Default | Range | Meaning |
|---|---|---|---|
queryGateway.serverPort |
8080 | 1024 to 49151 | The HTTP port. |
queryGateway.memoryLimit |
none | 64 MiB to 1 TiB | The memory limit of the query engine, for example 4GB. The units are B, KB, MB, GB and TB, in multiples of 1024. With no value there is no limit. On the HTTP interface the limit applies to each query. On the PostgreSQL interface all connections share one limit. |
queryGateway.maxWorkerThreads |
16 | 1 to 64 | The number of worker threads, and the number of parallel parts of a scan. |
queryGateway.maxInFlightHttpQueries |
16 | 1 to 1024 | The number of HTTP queries in progress on one pod. The gateway refuses one more with 503. |
queryGateway.snapshotPollIntervalSeconds |
8 | 5 to 10 | The interval at which the gateway asks the catalog for the current snapshot. |
queryGateway.queryTimeoutSeconds |
30 | 5 to 300 | The time limit to plan and prune a query. |
queryGateway.drainTimeoutSeconds |
48 | 5 to 600 | The time limit of the shutdown of the gateway: it accepts no new requests and lets the responses in progress finish. |
queryGateway.indexCache.maxBytes |
536870912 (512 MiB) | 64 MiB to 8 GiB | The size of the sidecar cache. |
queryGateway.manifestCache.maxBytes |
536870912 (512 MiB) | 16 MiB to 4 GiB | The size of the cache of table metadata. It is also the limit of the table metadata that one query can hold. |
queryGateway.warmup.windowHours |
24 | 1 to 168 | At start, the gateway loads the sidecars of the files added in this period. A value above compaction.snapshotRetentionHours has no effect beyond that retention, because the gateway cannot date a file whose snapshot has expired. |
queryGateway.auth¶
| Key | Default | Range | Meaning |
|---|---|---|---|
queryGateway.auth.issuerUrl |
Required | The OIDC issuer. The gateway finds the signing keys from {issuerUrl}/.well-known/openid-configuration. |
|
queryGateway.auth.audience |
none | If set, the aud claim of a token must be equal to it. |
|
queryGateway.auth.requiredClaims |
none | Claims, as name and text value, that a token must have. | |
queryGateway.auth.roleClaimPath |
none | A JMESPath expression that gives the roles in the token. Necessary with datasetRoles. |
|
queryGateway.auth.datasetRoles |
none | For each dataset, the roles that can read it. If the key is present, a dataset that is not in it, or that has an empty list, is denied to all callers. If the key is absent, each authenticated caller can read each dataset. | |
queryGateway.auth.maxJwksStalenessSeconds |
3600 | 60 to 86400 | How long the gateway continues to use its signing keys when the provider is unavailable. After that, requests get 503. |
queryGateway.auth.negativeCacheTtlSeconds |
30 | 1 to 300 | How long the gateway remembers that a key ID is not known, or that a fetch failed, before it asks the provider again. |
audience and roleClaimPath must not be empty text if you set them.
queryGateway.postgres¶
Without this block, the gateway has no PostgreSQL port. With it, all four keys are required.
| Key | Range | Meaning |
|---|---|---|
queryGateway.postgres.port |
1024 to 49151 | The PostgreSQL port. It must be different from queryGateway.serverPort and metrics.port. |
queryGateway.postgres.tls.certPath |
The certificate chain, in PEM. The gateway does not start if it cannot read the file. | |
queryGateway.postgres.tls.keyPath |
The private key, in PKCS#8. | |
queryGateway.postgres.maxConnections |
1 to 10000 | The number of connections on one pod. |
metrics¶
| Key | Default | Range | Meaning |
|---|---|---|---|
metrics.port |
9090 | 1024 to 49151 | The port of /metrics in each process. |
Keys with no effect in this version¶
The file accepts these keys, and examines a range where the table gives one, but no process uses the values.
| Key | Default | Range |
|---|---|---|
cluster.id |
Required | |
retention.windowDays |
90 | 1 to 3650 |
queryGateway.traceGraph.maxSpans |
10000 | |
queryGateway.traceGraph.maxResponseBytes |
67108864 | |
ingestion.datasets[].preset |
none | |
ingestion.datasets[].compaction.schedule |
A cron expression | |
ingestion.datasets[].compaction.backlogCheckIntervalMinutes |
5 to 120 |
TraceLake does not delete old data. retention.windowDays does not change that.