Architecture¶
TraceLake has three processes: the ingestor, the compactor and the query gateway. They share one configuration file. All durable data is in open Apache Iceberg tables in your own Azure Blob storage. An Iceberg REST catalog that you operate records which files make up each table.
Open the interactive diagram for pan and zoom, search, and export to PNG or SVG.
Components¶
| Component | Runs as | Does |
|---|---|---|
| Ingestor | StatefulSet | Consumes JSON events from Kafka, maps them to columns, stages them in a local Parquet write-ahead log, uploads the files and commits them to the table. |
| Compactor | Deployment | Merges small files into large ones, builds a bloom-filter and a full-text index for each merged file, and expires old snapshots. |
| Query gateway | StatefulSet | Authenticates each request, prunes the files a query cannot match, and scans the rest. Serves HTTP and, optionally, the PostgreSQL wire protocol. |
| Iceberg REST catalog | Your service, e.g. Lakekeeper | Holds the pointer to each table's current snapshot. TraceLake does not ship a catalog. |
| Azure Blob storage | Your storage account | Holds the Parquet data files, the Iceberg metadata and the index sidecars. |
| OIDC provider | Your identity provider | Issues the bearer tokens. The gateway fetches its signing keys (JWKS). |
Your data stays open¶
The tables are standard Iceberg tables. Trino, Spark, DuckDB or any other engine that supports the Iceberg REST catalog can read the same tables through the same catalog. The ingestor, compactor and gateway hold no data that is not also in the tables.