Skip to content

Architecture

TraceLake has three processes: the ingestor, the compactor and the query gateway. They share one configuration file. All durable data is in open Apache Iceberg tables in your own Azure Blob storage. An Iceberg REST catalog that you operate records which files make up each table.

TraceLake architecture TraceLake architecture

Open the interactive diagram for pan and zoom, search, and export to PNG or SVG.

Components

Component Runs as Does
Ingestor StatefulSet Consumes JSON events from Kafka, maps them to columns, stages them in a local Parquet write-ahead log, uploads the files and commits them to the table.
Compactor Deployment Merges small files into large ones, builds a bloom-filter and a full-text index for each merged file, and expires old snapshots.
Query gateway StatefulSet Authenticates each request, prunes the files a query cannot match, and scans the rest. Serves HTTP and, optionally, the PostgreSQL wire protocol.
Iceberg REST catalog Your service, e.g. Lakekeeper Holds the pointer to each table's current snapshot. TraceLake does not ship a catalog.
Azure Blob storage Your storage account Holds the Parquet data files, the Iceberg metadata and the index sidecars.
OIDC provider Your identity provider Issues the bearer tokens. The gateway fetches its signing keys (JWKS).

Your data stays open

The tables are standard Iceberg tables. Trino, Spark, DuckDB or any other engine that supports the Iceberg REST catalog can read the same tables through the same catalog. The ingestor, compactor and gateway hold no data that is not also in the tables.