I. Before the Facts, There Is Raw Data

In the first post, I described the Truth Layer as a structured API that holds what is physically true about a city's buildings and energy systems. The second blog post described the Knowledge Layer, an ontological engine that turns that structured data into semantic knowledge graphs on demand for any researcher or stakeholder.

The previous posts assumed structured data already sits in the Truth Layer's API, waiting to be queried. However, in practice, the data must come from somewhere. An urban digital twin draws data from multiple sources, including relational databases, CSV files handed over by a facility manager, JSON dumps, and live telemetry streaming data from BACnet-connected sensors. These data are heterogeneous and don't arrive clean, consistent, or in the same format. 

In this post, I will talk about the layer that sits before the Truth Layer: an open-source data lakehouse built on Apache Iceberg, Nessie, MinIO, Spark, Dremio, Superset, and Kafka. This layer turns heterogeneous data from multiple sources into consistent, versioned, and queryable data for the Truth Layer. It is an installable package with a command-line interface, rather than notebooks, to support replicability. 

II. One Storage Layer, Two Ways In

Figure 1.0 (AI-generated Image) An architectural depiction of how data from heterogeneous sources is structured in a data lakehouse.

This layer's architecture (Figure 1.0)  includes a set of tables organised into three stages using a pattern known as the medallion architecture. The Bronze aspect holds raw data in its original form without modification, while Silver holds a cleaned and validated version of the data. Gold holds data curated into a shape a researcher or downstream system can query. It is important to keep all three rather than overwriting raw data as soon as it’s cleaned because when results look wrong, you need to trace it back to exactly what the data looked like at every stage, not just the final stage. 

We store these tables using Apache Iceberg, an open table format, catalogued through Nessie and physically stored on MinIO, an S3-compatible object store. With this setup, every write to an Iceberg table creates a new, independent snapshot rather than overwriting the old one. This is what Michael Armbrust, Matei Zaharia, and colleagues call the lakehouse paradigm — the flexibility of a data lake with the reliability guarantees of a database. In practice, it means a researcher can query a table as it looked last week, not just as it looks today – the same reproducibility guarantee a version-controlled schema gives a software system.

III. Ingestion as Code, Not as Notebooks

In most cases, data engineering tutorials, and a fair amount of real research tooling, live in Jupyter notebooks, a sequence of cells, run top to bottom, by one person, on one machine. Notebooks are fine for exploratory work. It is, however, a poor foundation for a pipeline that other researchers may need to run, or even extend. So the first architectural decision here was the same one our Truth Layer made with DTOs and Mappers: every source type is an ordinary Python class, not a notebook cell. The ingestors write validated data to Iceberg tables, and they differ only in how they read data from their respective sources. A small base class carries the shared part:

class BaseLakehouseIngestor:
    def load(self, df: DataFrame, target: IcebergTableTarget) -> None:
        self._ensure_namespace(target)
        writer = df.writeTo(target.full_table_name)
        if target.mode == "createOrReplace":
            writer.createOrReplace()
        elif target.mode == "append":
            writer.append()
    def run(self, source, target: IcebergTableTarget) -> None:
        df = self.extract(source)
        self.load(df, target)

An ingestor for a Postgres database, a CSV file, and a JSON file implements an additional extract() method that turns its specific source into a Spark DataFrame and inherits everything else. Adding additional source types later means writing one method, and not necessarily re-deriving the whole pipeline. This is then wrapped in a command-line interface, so the same logic a researcher runs once from a terminal is exactly what a scheduled job runs every night:

lakehouse ingest csv-to-storage \
    --path data/meters.csv \
    --namespace facility_assets --target-table meters

For JSON data sources, when ingested, an export shape like {"dataset": {...}, "assets": [...]}  wraps the records, and a naive read produces just one row. To fix this, we used Python’s explode() function, which turns each array element into a row for the ingestion pipeline.

IV. Letting Researchers Ask Their Own Questions

A Truth Layer API answers the questions it was designed to answer. It is a poor tool for exploring questions a researcher asks once, out of curiosity, while forming a hypothesis. For example, a researcher may ask: What does the distribution of meter readings actually look like this month? Forcing every such question through a bespoke API endpoint means either building an endpoint for every possible question, which never finishes, or refusing the question.

This is the role of the federation layer. We run Dremio directly over the Iceberg tables, which lets anyone write ordinary SQL against the lakehouse without copying the data anywhere or waiting for an engineer to expose a new endpoint:

SELECT building_name, AVG(power_kw) AS avg_power
FROM nessie.facility_assets.telemetry t
JOIN nessie.facility_assets.inventory i ON t.asset_id = i.asset_id
GROUP BY building_name
ORDER BY avg_power DESC;

On top of that, Superset provides the consumption layer: dashboards a research team can build and share without writing code at all. This two-layer split (we are borrowing the terminology Alex Merced uses in Architecting an Apache Iceberg Lakehouse) means the lakehouse serves two very different audiences at once: a researcher comfortable in SQL gets direct, governed access to raw and curated tables; a collaborator who is not gets a dashboard. Neither has to wait on the other.

V. The One Data Source That Never Stops Arriving

Everything described so far handles data that arrives in discrete batches. A BACnet-connected sensor does not work that way. It reports continuously, for as long as it is powered on, and the lakehouse needs a genuinely different mechanism to receive it.

A point does not necessarily publish based on its data frequency. Depending on its configuration, a point may publish only when there is a change of value (COV) that deviates or moves past a threshold since the last report. While a voltage sensor with a small threshold reports often, a slow-moving totalizer may stay quiet for long stretches. We simulate that behaviour rather than emitting readings on an arbitrary clock:

def tick(self) -> float | None:
    self.true_value += random.gauss(0, self.volatility)
    if abs(self.true_value - self.last_reported_value) >= self.point.cov_increment:
        self.last_reported_value = self.true_value
        return self.true_value
    return None

Every simulated reading is sent to Apache Kafka. This message queue separates the entity generating the readings (currently our simulator, but eventually a real BACnet gateway) from the entity that processes them. On the receiving end, Spark Structured Streaming continuously retrieves data from Kafka and writes it into the same Iceberg tables described in Section II, at brief, consistent intervals.

df.writeStream.format("iceberg")
    .outputMode("append")
    .option("checkpointLocation", checkpoint_location)
    .trigger(processingTime="10 seconds")
    .toTable(target.full_table_name)

Streaming does not necessarily mean immediate real-time processing as the data arrives. Spark handles Kafka messages in short batches based on a timer rather than processing each message individually. At the same time, true record-by-record processing is possible (typically using Apache Flink); it adds significant operational complexity that most research pipelines don’t require. Additionally, every Iceberg write generates a new table snapshot independent of the engine that created it. Hence, the storage layer itself sets a minimum latency threshold that resembles batching, regardless of what sits in front of it. For sensor data, which BACnet is designed to report in intervals of seconds to minutes, a ten-second trigger interval is negligible and not a major concern.

VI. What is Missing

​​In the spirit of the transparency our second post aimed to demonstrate by leaving its final assertion unbacked until this post, we assert that this lakehouse is currently not connected to the Truth Layer. The two platforms are operating independently for now, without integration, and we wish to express that clearly instead of suggesting otherwise. 

We are considering two different design alternatives to resolving this problem. For the first approach, we plan to use a scheduled job to extract data from the gold layer into the Truth Layer’s database at predetermined intervals. The second option allows the Truth Layer to retrieve data by querying the federation layer directly on demand, eliminating the need to maintain a separate copy.

Both approaches have significant drawbacks: a push-based job introduces synchronisation latency and creates data redundancy. In contrast, a pull-based approach creates a real-time dependency between two systems that currently function independently. We believe we should not make this decision in isolation.

VIII. Conclusion: The Layer Before the Layer

In our first blog post in this series, we established what a city's data is. Our second blog post established how that data can be made to mean something, on demand, to anyone who asks. This post stays one level below the first and second posts; it is about how raw, inconsistent, and continuously arriving data becomes trustworthy enough to be used or to mean anything at all. The foundational architecture of data ingestion does not produce a knowledge graph or answer a research question on its own. It makes the layers that do so possible, and makes them possible in a way that another researcher can inspect, re-run, and trust. That is the kind of foundation worth building underneath a digital twin.

References

  • Armbrust, M., Ghodsi, A., Xin, R., & Zaharia, M. (2021, January). Lakehouse: a new generation of open platforms that unify data warehousing and advanced analytics. In Proceedings of CIDR (Vol. 8, No. 1, p. 28). sn.
  • https://iceberg.apache.org/
  • Merced, A. (2026). Architecting an Apache Iceberg Lakehouse. Manning.
  • https://kafka.apache.org/
  • https://polarsoft.com/bex/papers/Change%20of%20Value.pdf
  • Carry over the FAIR principles citation from Post 2 for continuity
  • Wilkinson et al. — The FAIR Guiding Principles for scientific data management and stewardship — Scientific Data, 2016