This study aims to develop a generalizable machine learning pipeline that uses only two-dimensional (2D) data to predict building characteristics, specifically construction year and floor space. This addresses the data gaps in building stock models, which are crucial for developing localized greenhouse gas emissions mitigation strategies.
Using a novel, national-level building registry dataset from Sweden, we trained machine learning classification models to predict construction year and regression models to predict floor space. The models were developed using only 2D building attributes, avoiding the need for 3D data, which is often unavailable at large geographic scales.
The result shows that the best-performing classification model achieves a precision measured as Area Under the Precision-Recall Curve (AUPRC) of 0.823 and the best-performing regression model achieves an R2 of 0.789. These results demonstrate that a 2D-based approach is sufficient for accurately predicting building characteristics.
This study reveals that imputing missing building attribute data does not require height data. By relying exclusively on widely available 2D data, the proposed machine learning pipeline could overcome the data limitations of previous studies. By demonstrating the effectiveness of this approach on a national-level dataset, this study improves the generalizability of building stock models and provides a scalable solution for estimating building characteristics at larger geographic scales.
1. Introduction
In 2021, the global building and construction sector accounted for over 34% of the energy demand and around 37% of the energy and process-related CO2 emissions globally (UNEP, 2022). Rapid changes are required to align the building sector with global climate targets, and the coming decade will be crucial for implementing the necessary emissions abatement measures (IEA, 2022). In the European Union (EU), decarbonizing the building stock is a key challenge to achieving the goal of climate neutrality by 2050 (European Climate Foundation, 2022). To align with the 1.5 °C scenarios outlined by the Intergovernmental Panel on Climate Change (IPCC), even more ambitious decarbonization targets must be set for the European building stock (Camarasa et al., 2022).
Building stock decarbonization is an urgent challenge that requires policymakers to develop policies and strategies grounded in a quantitative analysis of the building stock (Nägeli et al., 2020). The two types of CO2 emissions from buildings are: (1) operational emissions from energy consumption for heating, lighting and use of appliance and (2) embodied emissions associated with the materials production and construction processes (Ibn-Mohammed et al., 2013). Building energy models that investigate operational emissions reduction potentials rely on descriptions of the building stock to estimate the baseline energy demand (Kavgic et al., 2010), as well as evaluations of energy-efficient retrofitting methods (Deb and Schlueter, 2021). Similarly, building stock models are the basis for material flow and stock analyses (MFSA) to estimate reductions in embodied CO2 emissions from construction materials (Göswein et al., 2019).
To accelerate the implementation of emissions mitigation measures, more detailed building stock models are needed to develop localized solutions. Key variables, such as the floor space and construction year of individual buildings, are critical for understanding the energy and material demands. However, statistical data on these factors are often incomplete or unavailable. Machine learning (ML) has emerged as a valuable tool for addressing these data gaps, although most existing studies rely on 3D data, which limits their geographic scope due to a lack of data. To enhance the generalizability of ML models for predicting building characteristics across larger regions, we propose a pipeline that uses two-dimensional (2D) data exclusively. Leveraging a novel Swedish national building registry dataset, we train ML classification and regression models to predict the construction year and floor space, respectively.
This study aims to develop ML models that predict both the building construction year and usable floor space at the national level without using height data, with the primary objective of improving bottom-up building stock models. The goal of this work is to improve the generalizability of ML models as a tool to fill in the gaps in the building inventory datasets for larger geographic scales. We use residential buildings in Sweden as a case study and test urban form features, as proposed previously (Nachtigall et al., 2023). The novelty of this work is the usage of 2D urban form data as features without the need for three-dimensional (3D) height data. The advantages of using 2D data are that 2D data are highly accessible compared to 3D height data, which are scarce.
This article is structured as follows. In Section 2, we provide a brief overview of current building stock models and ML approaches, motivating the present work. Section 3 presents the scope and method of this study. In Section 4, we present our findings. Finally, in Section 5, we draw conclusions from the present work and outline key areas for further investigation.
2. Overview of building stock models and machine learning approaches
2.1 Building stock models
In general, the approaches adopted to develop building stock models (both for energy and materials) can be top-down or bottom-up (Kavgic et al., 2010; Tanikawa et al., 2015). Top-down models do not consider the differences among individual buildings. For building energy models, the top-down approach estimates energy consumption at the building sector level (Li et al., 2017). For building material stock models, the top-down approach uses aggregated data on material use (e.g. material trade, consumption statistics) and applies lifetime estimates to derive the quantities of material accumulated in the building stock and material outflows (Müller et al., 2014). In contrast, the bottom-up approach models a building stock “piece by piece,” i.e. at the building level. This type of modeling is particularly important in decarbonization research because bottom-up approaches can be integrated with geospatial data in the Geographic Information System (GIS), so as to gain a better understanding of the physical composition of the building stock (Lanau et al., 2019). For energy models, the energy demand of each building (or group of buildings) is modeled by taking into account the physics of the buildings (Kamel, 2022). Material stock models take into account those building characteristics that may impact their material content, such as age, use, and/or construction type. Regardless of their focus (energy or materials), modeling building stocks using a bottom-up approach requires substantial amounts of data (Kong et al., 2023; Lanau et al., 2019) or need to be based on a condensed description of the buildings, typically involving approximation of the stock through the use of a limited number of type buildings of archetype buildings (Lanau et al., 2019).
However, most of the currently available building inventory datasets are incomplete (Wang et al., 2022; Lanau et al., 2019). Critical to these models are geometric data that describe the physical dimensions of buildings (e.g. building footprints, building heights, number of stories) (Monteiro et al., 2018). To this end, GIS datasets of building inventories are commonly used, as they provide location-specific, high-granularity spatial data for both the building material stock (Rajaratnam et al., 2023) and building energy modeling (Wong et al., 2021).
The highest-resolution building inventory data should contain geo-referenced attributes, such as building footprints, heights, floor space, building type and construction year (Milojevic-Dupon et al., 2023). However, most national-level inventory datasets do not contain comprehensive information on these attributes. For example, a harmonized European building inventory dataset using governmental data and OpenStreetMap (OSM) has shown that for building height, construction year, and building type, only 73%, 24%, and 46%, respectively, of the data are available (Milojevic-Dupon et al., 2023). In Europe, only Spain, the Netherlands and France have national-level databases that contain building footprints and attributes (Milojevic-Dupon et al., 2023). Two tools have been developed by the EU to collect building inventory data, each with its own limitations. The Copernicus Reference Data Access (CORDA) node only contains data from countries with high-quality databases (European Environment Agency, 2023). In contrast, the EU Building Stock Observatory only provides aggregated data at the country level for all of the EU Member States (European Comission, 2023). The Microsoft Corporation has developed an open-source building footprint dataset that uses a combination of neural networks and aerial images, although building height coverage in the dataset is limited to a few selected regions (Microsoft, 2023). This lack of data is a key obstacle to developing high-resolution building stock models.
2.2 Machine learning approaches
Machine learning-based approaches have emerged as a method to fill in the gaps in incomplete datasets so as give higher spatial resolution and accuracy (Milojevic-Dupont and Creutzig, 2021). Machine learning or statistical learning represents algorithms that can learn from the attributes of an input dataset to generate a prediction based on the learning process (for an introduction to ML, see (Hastie et al., 2009)). In the context of building stock modeling, ML-based approaches are often used to predict building attributes such as age or height, in order to complete or enrich building inventory datasets. As shown in Table 1, existing work using ML-based approaches for building inventory development has been applied at various spatial scales. Most of the existing work was performed at the urban scale (neighborhood or city), with some studies carried out at the national scale (Milojevic-Dupont et al., 2020b; Arehart et al., 2021; Nachtigall et al., 2023).
Comparisons of the different machine learning approaches used for building inventory modeling, sorted based on spatial scale
| Spatial scale | Geographic location | 3D building geometry | Predicted attributes | Reference |
|---|---|---|---|---|
| Neighborhood | Vancouver, Canada | Yes | Age | Tooke et al. (2014) |
| Neighborhood | Trento and Bologna, Italy | Yes | Height | Farella et al. (2021) |
| City | Rotterdam, the Netherlands | Yes | Age | Biljecki and Sindram (2017) |
| City | Nottingham, UK | Yes | Age | Rosser et al. (2019) |
| City | Shenzhen, China | Yes | Age | Mao et al. (2022) |
| City | North-Rhine Westphalia, Germany | Yes | Age | Garbasevschi et al. (2021) |
| City | Rotterdam, The Netherlands | Yes | Height | Biljecki et al. (2017) |
| City and National | The Netherlands, Germany, Italy, France | Partial | Height | Milojevic-Dupont et al. (2020b) |
| National | Canada and USA | No | Height and floor space | Arehart et al. (2021) |
| National | The Netherlands, France, Spain | Yes | Age | Nachtigall et al. (2023) |
| Spatial scale | Geographic location | 3D building geometry | Predicted attributes | Reference |
|---|---|---|---|---|
| Neighborhood | Vancouver, Canada | Yes | Age | |
| Neighborhood | Trento and Bologna, Italy | Yes | Height | |
| City | Rotterdam, the Netherlands | Yes | Age | |
| City | Nottingham, UK | Yes | Age | |
| City | Shenzhen, China | Yes | Age | |
| City | North-Rhine Westphalia, Germany | Yes | Age | |
| City | Rotterdam, The Netherlands | Yes | Height | |
| City and National | The Netherlands, Germany, Italy, France | Partial | Height | |
| National | Canada and USA | No | Height and floor space | |
| National | The Netherlands, France, Spain | Yes | Age |
The limitation with regards to spatial resolution is mainly due to the inherently high demand for data in ML-based approaches. This is especially true for studies that use 3D building geometry data. A 3D building dataset is usually derived using geoinformatics techniques or inspections of building plans, and depending on its level of detail, it contains building information such as height and roof angle (Biljecki et al., 2015). These 3D building data are informative for modeling a building stock, although official statistics 3D data are rarely available.
As demonstrated previously (Arehart et al., 2021), ML models are capable of providing good predictions of usable floor space in the absence of 3D building attribute data and with the benefit of a large geographic scope at the national level. However, Arehart and colleagues (Arehart et al., 2021) have utilized satellite remote sensing imagery, which introduces additional uncertainty to the predictions and limits the generalizability of the approach. Furthermore, no previous studies have predicted building age at the national level without the usage of building height data. Nachtigall and coworkers (Nachtigall et al., 2023) have predicted the age of residential buildings in the Netherlands, France and Spain using mainly urban form features with a mix of real height data and previously predicted height data from (Milojevic-Dupont et al., 2020b). The results indicate that building and building neighborhood-level features have the greatest predictive power in terms of feature importance (Nachtigall et al., 2023).
3. Methods and materials
In this study, we train two supervised ML models to predict the construction year and usable floor space of residential buildings using an administrative building registry dataset without height data. Geometric data, such as building footprints and streets, were used to supplement the registry dataset, so as to generate urban morphology features for ML. This section presents the case study, data, feature engineering process and the ML models used. For an overview of the methodology, see Figure 1.
The flow diagram presents a left-to-right process with two parallel modeling streams connected by directional arrows. On the far left, a rounded rectangle labeled “Building registry dataset with missing data” points rightward to a parallelogram labeled “Preprocessing and feature engineering”. From “Preprocessing and feature engineering”, two thin arrows branch. The upper branch points upward-right to a rounded rectangle labeled “Create building construction year training dataset (without usable floor space)”. A rightward arrow leads to “Training and testing different resampling methods for classification model”, followed by another rightward arrow to “Assign construction year using best performing classification model”. The lower branch from “Preprocessing and feature engineering” points downward-right to a rounded rectangle labeled “Create building usable floor space training dataset (with construction year)”. A rightward arrow leads to “Training and testing different regression algorithms”, followed by another rightward arrow to “Assign usable floor space using best performing regression model”. From both assignment boxes, thin horizontal connector lines extend rightward and converge into a single right-pointing arrow leading to the final rounded rectangle labeled “M L completed building registry dataset”.Methodological workflow that is used in this article
The flow diagram presents a left-to-right process with two parallel modeling streams connected by directional arrows. On the far left, a rounded rectangle labeled “Building registry dataset with missing data” points rightward to a parallelogram labeled “Preprocessing and feature engineering”. From “Preprocessing and feature engineering”, two thin arrows branch. The upper branch points upward-right to a rounded rectangle labeled “Create building construction year training dataset (without usable floor space)”. A rightward arrow leads to “Training and testing different resampling methods for classification model”, followed by another rightward arrow to “Assign construction year using best performing classification model”. The lower branch from “Preprocessing and feature engineering” points downward-right to a rounded rectangle labeled “Create building usable floor space training dataset (with construction year)”. A rightward arrow leads to “Training and testing different regression algorithms”, followed by another rightward arrow to “Assign usable floor space using best performing regression model”. From both assignment boxes, thin horizontal connector lines extend rightward and converge into a single right-pointing arrow leading to the final rounded rectangle labeled “M L completed building registry dataset”.Methodological workflow that is used in this article
The two target variables in this study are construction year and usable floor space of buildings. Rather than training one multioutput model, we train two separate models due to the differences between the tasks. Indeed, the construction year prediction is treated as a classification task, and the usable floor space prediction is considered as a regression task. The rationale for treating age predictions as a classification task is because classification models are more flexible as the construction years can be binned into categories or age cohort of, for example, per decade. For more details, refer to S1.2 in Supplementary Information. In both cases, existing data are used as the ground truth to train the 2 ML models, respectively.
3.1 Feature engineering
In ML, a feature serves as a numeric representation of a specific aspect of the raw data. Since ML models require an adequate number of features to capture and reflect the underlying data's characteristics effectively, feature engineering is essential, so it is the first step in the ML workflow (Dong and Liu, 2018). This process involves extracting features from the raw data and converting them into formats that are suitable for utilization by the ML model (Zheng and Casari, 2018). For the descriptions of the preprocessing steps of the data set, please see the Supplementary Information (S1.1). This section focuses on the feature generation process, more specifically on the generation of urban morphology features.
Previous research has shown urban morphology data to be highly predictive for estimating the height and age of buildings. Milojevic-Dupont et al. (2020a) used building footprints and street networks to generate urban morphology features and predict the heights of buildings in Europe. They concluded that the ML models trained using urban morphology features could predict the heights of buildings with an average error of less than 2.5 m. Using the same set of features, Nachtigall et al. (2023) expanded the approach to predict the construction year with limited data availability. Their resulting model showed good prediction precision, with a mean absolute error of 15.3–19.9 years for unseen data. Based on these results, we hypothesize that urban morphology features capture the underlying correlations between the usable floor space of a building and its surrounding urban fabrics. Thus, we generate 16 features, applied to three elements of urban morphology (building level, building plot level and building block level), for a total of 48 urban features. The following section provides the rationale for each element of urban morphology and how the features generated based on each element are related to the floor space and construction year.
3.1.1 Urban morphology elements
The quantitative analysis of urban morphology distinguishes between three key building elements: footprints, plots and street-based blocks (Berghauser Pont et al., 2019), as illustrated in Figure 2, where tessellations enclosed by streets are used as a proxy for building plots. Each element captures the unique attributes of the corresponding building, and the features are generated based on the elements.
The figure consists of three side-by-side map panels labeled “(a) Buildings”, “(b) Tessellations”, and “(c) Blocks”. The spatial pattern across panels shows detailed building-level classification in (a), polygon-based tessellations in (b), and fully aggregated gray block units in (c). A legend below indicates housing categories: “Detached houses” in green, “Apartments” in yellow, “Row Houses” in purple, “Terraced houses” in blue-green, and “Blocks” in purple. In panel “(a) Buildings”, green “Detached houses” dominate the left and central areas of the map. Yellow “Apartments” are concentrated mainly in the right and lower-right portions. Purple “Row Houses” appear in smaller clusters scattered across the central and upper-central areas. Blue-green “Terraced houses” appear in limited clusters, primarily near the left-central region. In panel “(b) Tessellations”, green “Detached houses” cover large contiguous areas across the left and central regions. Yellow “Apartments” form concentrated clusters on the right side. Purple “Row Houses” appear in distinct patches primarily in the central area. Blue-green “Terraced houses” are visible in smaller grouped sections mainly toward the left-central and mid-lower regions. In panel “(c) Blocks”, all areas are aggregated and displayed in purple, labeled as “Blocks”, covering the entire mapped region without housing-type color differentiation.Illustration of the three elements of urban morphology using an area of Gothenburg (Sweden) as an example: (a) building footprints; (b) tessellations enclosed by streets, used as a proxy for building plots and (c) building blocks generated using tessellations. Note that no color coding is used here because the blocks include a variety of building uses
The figure consists of three side-by-side map panels labeled “(a) Buildings”, “(b) Tessellations”, and “(c) Blocks”. The spatial pattern across panels shows detailed building-level classification in (a), polygon-based tessellations in (b), and fully aggregated gray block units in (c). A legend below indicates housing categories: “Detached houses” in green, “Apartments” in yellow, “Row Houses” in purple, “Terraced houses” in blue-green, and “Blocks” in purple. In panel “(a) Buildings”, green “Detached houses” dominate the left and central areas of the map. Yellow “Apartments” are concentrated mainly in the right and lower-right portions. Purple “Row Houses” appear in smaller clusters scattered across the central and upper-central areas. Blue-green “Terraced houses” appear in limited clusters, primarily near the left-central region. In panel “(b) Tessellations”, green “Detached houses” cover large contiguous areas across the left and central regions. Yellow “Apartments” form concentrated clusters on the right side. Purple “Row Houses” appear in distinct patches primarily in the central area. Blue-green “Terraced houses” are visible in smaller grouped sections mainly toward the left-central and mid-lower regions. In panel “(c) Blocks”, all areas are aggregated and displayed in purple, labeled as “Blocks”, covering the entire mapped region without housing-type color differentiation.Illustration of the three elements of urban morphology using an area of Gothenburg (Sweden) as an example: (a) building footprints; (b) tessellations enclosed by streets, used as a proxy for building plots and (c) building blocks generated using tessellations. Note that no color coding is used here because the blocks include a variety of building uses
Building elements refer to the 2D geometry of the footprints of buildings, which contain useful information that can be used for predicting construction age (Rosser et al., 2019) and floor space (Arehart et al., 2021) through ML models.
Plots, as a concept, remain ambiguous because they can be represented in different ways across different geographies (Fleischmann et al., 2020). In this article, the term plot refers to the physical form that is “a land-use unit defined by boundaries on the ground” (Conzen, 1960; Kropf, 2018). Since our cadaster dataset did not include plot information, we derived it using the Morphological Tessellation (MT) method presented by Fleischmann et al. (2020). Morphological Tessellation is a method for deriving a reliable, universal, and meaningful plot-scale spatial unit of analysis, which we use to generate morphological cells that are proxies for plots. In the tessellation generation process, we use streets as enclosures, to capture realistically the building's surrounding geography. We also include the footprints of nonresidential buildings to ensure that the effect of their presence on residential plots is considered. The resulting morphological cells represent the theoretical plots of each building, which extend up to either streets or the plots of neighboring buildings.
Block refers to buildings in the same neighborhood, which tend to have similar attributes. For example, buildings in the same blocks tend to have similar heights (Milojevic-Dupont et al., 2020b) and construction years (Nachtigall et al., 2023). To capture these underlying spatial autocorrelations, we generate building blocks based on tessellations and street networks. Similar to building plots, blocks can be defined as statistical areas, although they can be allocated differently across different geographies. We generate building blocks using the Momepy package (Fleischmann, 2019), based on buildings, tessellations of buildings, and street networks. The tessellations are grouped based on the enclosing street network and the geometries of the grouped tessellations are dissolved (i.e. union) into a single block geometry. This process also generates a unique block ID, which we use to link the blocks to the buildings and tessellations.
We hypothesize that the descriptive statistics of features, such as the mean and standard deviation of all buildings within the same block, capture the overall characteristics of the block and, thus, have predictive power for the ML models. We generate the means and standard deviations for all the building levels and the tessellation features for each block.
3.1.2 Generated features
For each urban morphology element, we generate 16 features based on their geometry (see Table 2). In the case of blocks, each building in the same block shares the same block-level features. In addition to the generated features, some information from the dataset is also used as features. These additional features include the building type, county and municipality to which each building belongs, X and Y coordinates and the construction year (both ground truth and predicted value). In total, 55 features are used in the model training process.
List of generated features, their descriptions and formulas
| Feature | Description | Formula |
|---|---|---|
| Area (m2) | Area of each given object | – |
| Perimeter (m) | Perimeter of each given object | – |
| Longest axis length (m) | The length of the longest axis of each given object; the axis is defined as the diameter of a minimal circumscribed circle around the convex hull | |
| Circular compactness | Area of each given object divided by the corresponding area of the enclosing circle | |
| Convexity | Area of each given object divided by the corresponding area of the convex hull | |
| Elongation | The elongation of each given object is regarded as elongation of its minimum bounding rectangle, where a is the area of the object and p is its perimeter | |
| Equivalent rectangular index | The equivalent rectangular index of each object, where BC is the bounding rectangle and BR is the bounding rectangle | |
| Fractal dimension | The fractal dimension of each given object | |
| Rectangularity | The rectangularity of each given object, where RRA is the rotated rectangle area | |
| Shape index | The shape index of each given object | |
| Square compactness | The square compactness of each given object | |
| Orientation | An orientation of the longest axis of the bounding rectangle in the range of 0–45° | – |
| Corners | The number of corners of each given object | |
| Centroid corners | The number of centroid corners of each given object | |
| Centroid corners mean | The mean distance of the centroid corners, where n is the number of centroid corners | |
| Centroid corners standard deviation | The standard deviation of the distance of the centroid corners, where n is the number of centroid corners and x is the distance |
| Feature | Description | Formula |
|---|---|---|
| Area (m2) | Area of each given object | – |
| Perimeter (m) | Perimeter of each given object | – |
| Longest axis length (m) | The length of the longest axis of each given object; the axis is defined as the diameter of a minimal circumscribed circle around the convex hull | |
| Circular compactness | Area of each given object divided by the corresponding area of the enclosing circle | |
| Convexity | Area of each given object divided by the corresponding area of the convex hull | |
| Elongation | The elongation of each given object is regarded as elongation of its minimum bounding rectangle, where a is the area of the object and p is its perimeter | |
| Equivalent rectangular index | The equivalent rectangular index of each object, where | |
| Fractal dimension | The fractal dimension of each given object | |
| Rectangularity | The rectangularity of each given object, where | |
| Shape index | The shape index of each given object | |
| Square compactness | The square compactness of each given object | |
| Orientation | An orientation of the longest axis of the bounding rectangle in the range of 0–45° | – |
| Corners | The number of corners of each given object | |
| Centroid corners | The number of centroid corners of each given object | |
| Centroid corners mean | The mean distance of the centroid corners, where n is the number of centroid corners | |
| Centroid corners standard deviation | The standard deviation of the distance of the centroid corners, where n is the number of centroid corners and x is the distance |
Note(s): Each of the features are calculated for each building's footprint, tessellation and corresponding block. In total, 48 features are generated and used as the training dataset. All features are calculated using the momepy package
3.2 Classification model to predict years
In the application of building inventory data, construction years are often used as age cohorts in applications, for example the material intensities of building stocks are frequently presented in 10-year cohorts (Gontia et al., 2018; Kaasalainen et al., 2023). Therefore, we treat the construction year prediction as a classification problem with 12 10-year cohorts from 1900 to 2020 and an additional cohort of 2021–2023, as shown in Figure 3b. The binning of construction years into cohorts has the additional benefit of reducing the complexity of the classification task, as well as the computational demands (Lorena et al., 2019). We refer the reader to the Supplementary Information (SI 1.2) for further justification.
The four-panel figure presents housing characteristics across panels labeled “(a)”, “(b)”, “(c)”, and “(d)”. Panel “(a)” is a pie chart showing the distribution of building types. The legend lists five categories with percentages shown in the top left. The categories in clockwise order are as follows: “Detached houses” at 87.80 percent, “Apartments” at 5.20 percent, “Row houses” at 4.40 percent, “Terraced houses” at 1.30 percent, and “Small houses” at 1.30 percent. The text below the pie chart states “Total number of buildings equals 2,833,517”. Panel “(b)” is a histogram labeled with the horizontal axis “Construction year” and the vertical axis “Count”. The horizontal axis ranges from 1900 to 2025 in increments of 25 years. The vertical axis ranges from 0 to 400000 in increments of 50000 units. The highest bar appears around the 1975 to 1985 period, with a count slightly above 400000. Elevated counts also appear around 1955 to 1965 at approximately 340000. Earlier decades, around 1900 to 1910, show counts near 250000. Counts decline after 2000, with values near 100000 to 150000. Panel “(c)” is a histogram labeled with the horizontal axis “Usable floor space” and the vertical axis “Count”. The horizontal axis ranges from 0 to 500 in increments of 100 units. The vertical axis ranges from 0 to 200000 in increments of 50000 units. The distribution is right-skewed. The highest frequency occurs between 100 and 120 square meters with counts slightly above 200000. Frequencies decline gradually beyond 150 square meters and taper off after 300 square meters. Panel “(d)” is a histogram labeled with the horizontal axis “Usable floor space” and the vertical axis “Count”. The horizontal axis ranges from 0 to 10000 in increments of 2000 units. The vertical axis ranges from 0 to 1000 in increments of 200 units. The distribution is heavily right-skewed. The highest concentration occurs below 500 square meters with counts approaching 1000. Counts decrease sharply beyond 2000 square meters and approach zero near 10000 square meters. Note: All numerical data values are approximated.Distribution of existing data in the building registry dataset: (a) shares of each type of building (by count); (b) construction years of buildings (all types) with 10-year intervals, (c) usable floor space of houses (detached, terraced, small and row) with 10 m2 intervals and (d) usable floor space of apartment buildings with intervals of 10 m2. Missing data is not reflected in the figure
The four-panel figure presents housing characteristics across panels labeled “(a)”, “(b)”, “(c)”, and “(d)”. Panel “(a)” is a pie chart showing the distribution of building types. The legend lists five categories with percentages shown in the top left. The categories in clockwise order are as follows: “Detached houses” at 87.80 percent, “Apartments” at 5.20 percent, “Row houses” at 4.40 percent, “Terraced houses” at 1.30 percent, and “Small houses” at 1.30 percent. The text below the pie chart states “Total number of buildings equals 2,833,517”. Panel “(b)” is a histogram labeled with the horizontal axis “Construction year” and the vertical axis “Count”. The horizontal axis ranges from 1900 to 2025 in increments of 25 years. The vertical axis ranges from 0 to 400000 in increments of 50000 units. The highest bar appears around the 1975 to 1985 period, with a count slightly above 400000. Elevated counts also appear around 1955 to 1965 at approximately 340000. Earlier decades, around 1900 to 1910, show counts near 250000. Counts decline after 2000, with values near 100000 to 150000. Panel “(c)” is a histogram labeled with the horizontal axis “Usable floor space” and the vertical axis “Count”. The horizontal axis ranges from 0 to 500 in increments of 100 units. The vertical axis ranges from 0 to 200000 in increments of 50000 units. The distribution is right-skewed. The highest frequency occurs between 100 and 120 square meters with counts slightly above 200000. Frequencies decline gradually beyond 150 square meters and taper off after 300 square meters. Panel “(d)” is a histogram labeled with the horizontal axis “Usable floor space” and the vertical axis “Count”. The horizontal axis ranges from 0 to 10000 in increments of 2000 units. The vertical axis ranges from 0 to 1000 in increments of 200 units. The distribution is heavily right-skewed. The highest concentration occurs below 500 square meters with counts approaching 1000. Counts decrease sharply beyond 2000 square meters and approach zero near 10000 square meters. Note: All numerical data values are approximated.Distribution of existing data in the building registry dataset: (a) shares of each type of building (by count); (b) construction years of buildings (all types) with 10-year intervals, (c) usable floor space of houses (detached, terraced, small and row) with 10 m2 intervals and (d) usable floor space of apartment buildings with intervals of 10 m2. Missing data is not reflected in the figure
We used the XGBoost algorithm for the classification task, as it has been successfully applied to predict various attributes of the building stock with high levels of precision (Milojevic-Dupont et al., 2020b; Nachtigall et al., 2023). We tested the XGBoost algorithm with different resampling methods to address class imbalance issues in the dataset.
3.2.1 Class imbalance
The construction years of the buildings in the registry dataset display an unbalanced and positively skewed distribution, as shown in Figure 3b. For more information on the dataset used, refer to Section 3.5. The number of buildings constructed between 1960 and 1980 is significantly higher than the numbers in the other construction-year cohorts. This reflects the Swedish “Million Program” in which more than 1 million buildings were built between 1965 and 1974 (Boverket, 2024). Class imbalance is a problem for classification models when the underlying data are skewed and certain classes are under-represented, making the ML model's ability to predict minority classes less-effective (Japkowicz, 2000). Thus, before using the gradient boosting algorithms – which can mitigate some of the challenges of class imbalance problem albeit not always in a satisfactory manner – a data preprocessing step was required to counteract the class imbalance and maximize the ML model's performance (Zhang et al., 2022). We focused on under-sampling techniques to tackle the imbalance problems. We refer the reader to Supplementary Information (SI 1.3) for a justification regarding the selection of under-sampling techniques rather than over-sampling or hybrid techniques.
We used the open-source Python package “Imbalanced-learn” (LemaÃŽtre et al., 2017) to test four under-sampling techniques to identify that with the optimal prediction performance. First, we tested the Random Under-Sampler (RUS) technique, which randomly under-samples the majority classes without replacement. Second, we tested the Near Miss (NM) method, which selects a subset of the majority class samples that are closest to the minority class samples (Mani and Zhang, 2003). Third, we tested the One-sided Selection (OSS) technique, which initially finds the observations that are hard to classify and then removes noisy samples (Kubat and Matwin, 1997). Fourth, we tested the Neighborhood Cleaning Rule (NCR) method, which uses a combination of edited-nearest-neighbor and a k-nearest-neighbor to remove noisy samples from the dataset (Laurikkala, 2001).
3.2.2 Classification model training and evaluation metrics
The multiclass classification models with different under-sampling techniques were implemented in an identical manner using the XGBoost algorithm in Python. The learning objective of the model was to predict the probability of each data-point belonging to each class. In the model training process, we used a five-fold cross-validation method to prevent the model from overfitting the training data. Cross-validation is a process used to test the ML model on unseen data to prevent bias in results, and a five-fold cross-validation means the model is split into 5 groups during the cross-validation process. The hyperparameters of the models were optimized using the Optuna Python package (Akiba et al., 2019). Optuna is an open-source hyperparameter optimization framework, and for the details on the package refer to the documentation (Akiba et al., 2019).
Several evaluation metrics were used, the purposes and equations for which are summarized in Table 3.
Evaluation metrics for classification models
| Evaluation metric | Purpose | Equation | Description |
|---|---|---|---|
| Precision | Single number summary of the precision recall curve | True positives (TP) are the data-points that are predicted to be the positive class and it is correct; False positives (FP) are the data-points that are predicted to be the positive class and it is false | |
| Recall | Ratio of true-positive predictions to the total actual positives; it measures the accuracy of positive predictions | False negatives (FN) are the data-points that are the data-points that are predicted to be the negative class and it is false | |
| MAUC | The averages of the AUC values of all pairs of classes in the dataset | C is the number of classes in the dataset, and is the AUC score between class and class . The MAUC score ranges from 0 to 1, and a perfect classifier has a MAUC score of 1 | |
| G-mean | A single value summary of a classifier's ability to identify correctly positive instances and negative instances | C is the number of classes in the dataset | |
| Mathews Correlation Coefficient (MCC) | The MCC score considers the confusion matrix values of all classes in the dataset | The MMC score ranges from −1 to 1, and a perfect classifier has an MMC score of 1 | |
| Mean Mathews Correlation Coefficient (MMCC) | An average of all the MCC values of all pairs of classes | C is the number of classes in the dataset, and is the MCC score between class and class |
| Evaluation metric | Purpose | Equation | Description |
|---|---|---|---|
| Precision | Single number summary of the precision recall curve | True positives ( | |
| Recall | Ratio of true-positive predictions to the total actual positives; it measures the accuracy of positive predictions | False negatives ( | |
| The averages of the AUC values of all pairs of classes in the dataset | C is the number of classes in the dataset, and | ||
| G-mean | A single value summary of a classifier's ability to identify correctly positive instances and negative instances | C is the number of classes in the dataset | |
| Mathews Correlation Coefficient ( | The | The | |
| Mean Mathews Correlation Coefficient ( | An average of all the | C is the number of classes in the dataset, and |
All the evaluation metrics are calculated for each trained and validated classification model using the different under-sampling techniques to determine the best-performing model. The predicted construction years from the best-performing model are added to the training dataset and used as a feature for the subsequent regression models.
3.3 Regression model to predict floor space
Next, we performed a regression task in which we compared the abilities of the following three algorithms to predict the usable floor space of buildings: Random forests (Breiman, 2001), XGBoost (Chen and Guestrin, 2016) and LightGBM (Ke et al., 2017).
In contrast to age cohorts, usable floor space is a continuous value and, thus, is more suitably modeled through regression. The prediction of usable floor space is treated as a single-target variable regression problem. In addition to XGBoost, we tested 3 ML algorithms that are more commonly used for regression problems. Random forest (RF) is an ensemble learning algorithm that builds a collection of decision trees and combines their predictions to create a more accurate and robust model (Breiman, 2001). The RF algorithm is implemented using the scikit-learn package in Python (Pedregosa et al., 2011). LightGBM is another gradient-boosting decision tree algorithm developed by Microsoft that has higher training efficiency and lower memory usage (Ke et al., 2017). In addition, CatBoost is a gradient-boosting algorithm that has been developed to offer more support for categorical features (Prokhorenkova et al., 2018). In the same way as for the classification model training process, we use a five-fold cross-validation method to prevent the model from overfitting to the training data. The hyper-parameters of the models were optimized using the Optuna Python package (Akiba et al., 2019).
To ensure consistency of the tests, we use the Mean squared error (MSE) as the evaluation metric during the model training process. MSE measures the average squared differences between the predicted and actual values and is defined as:
where n is the total number of data-points, is the actual value of the -th observation and is the predicted value of the -th observation.
We use two additional evaluation metrics to compare the different ML algorithms for regressions: Mean absolute error (MAE) and R-squared score (R2). MAE measures the average absolute differences between the predicted and actual values, defined as:
where n is the total number of data-points, is the actual value of the -th observation, and is the predicted value of the -th observation. MAE penalizes errors to a lesser degree than the MSE, but it is a more intuitive evaluation metric.
The R-squared score measures the proportion of the variance in the dependent variable that is predicted from the independent variables and is defined as:
where n is the total number of data-points, is the actual value of the -th observation, is the predicted value of the -th observation and is the mean of the target variable. All the evaluation metrics are calculated for each trained and validated classification model using the different algorithms to determine the best-performing model.
3.4 Sweden as a case study
The Swedish residential building stock was chosen as a case study to demonstrate the applicability of the method. Located in Northern Europe, Sweden has a total area of 447,430 km2 and a population of 10.5 million in 2023 (Statistikmyndigheten, 2023). It is divided into 21 administrative regions and 290 municipalities. Sweden is one of the most forested countries in Europe, with 69% of the land area covered by forests, and the built-up area accounts for only 3.1% of the land area or 1.3 million hectares (Statistikmyndigheten, 2023). Residential land use accounts for 36% of the total built-up area, or approximately 1.1% of the total land area. Sweden is chosen as a case study due to the good availability of GIS data relevant to describing the building stock, although the method is generalizable to other geographic areas.
3.5 Data description
This section provides an overview of the real-world datasets (referred to as ground truth data) used to calibrate the ML models in this study.
The ground truth dataset comprises the single-family and multi-family buildings from the building registry of the Swedish Land Survey (Lantmäteriet in Swedish) as of November 2022. This official database on land ownership is used primarily for property tax assessment purposes.
The registry contains approximately 8 million buildings, of which 2.8 million are residential buildings. Attributes obtained from the registry include building type, construction year and usable floor space (here defined as the area of the building intended for occupancy, excluding basements and attics) (Lantmäteriet, 2023b).
The key descriptive statistics of the dataset are summarized in Table 4. The degree of incompleteness of the dataset with regard to construction year and usable floor space varies across building types, and is likely due to the fact that the data are collected and reported by individual municipalities and that there is a lack of national-level reporting requirements (Lantmäteriet, 2023a). Statistics on the existing data show apartment buildings have significantly larger usable floor space per building than other building types. This is due to apartment buildings being registered as entire buildings, i.e. their usable floor area is the sum of the area of each apartment in the building. Apartment buildings and small houses exhibit the widest Interquartile Range (IQR) values of 828 and 179, respectively. They also have the highest proportion of missing data for usable floor space (63% and 86%, respectively) and construction year (20% and 27%, respectively). Figure 3 provides a visualization of the distribution of existing data in the building registry dataset.
Key summary statistics for the dataset used in this study
| Usable floor space/building (m2) | Construction year | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| Building type | n | Mean | Std. | Q1 | Q3 | Missing | Mean | Std. | Q1 | Q3 | Missing |
| Detached houses | 1,982,870 | 115.4 | 50.2 | 81 | 143 | 15% | 1,955 | 35 | 1,930 | 1,977 | 18% |
| Terraced houses | 112,356 | 126.0 | 32.6 | 110 | 141 | 6% | 1,974 | 17 | 1,969 | 1,979 | 5% |
| Small houses | 4,499 | 348.1 | 846.3 | 145 | 324 | 86% | 1,993 | 38 | 1,977 | 2,013 | 27% |
| Row houses | 30,000 | 127.7 | 138.9 | 104 | 132 | 15% | 1,978 | 19 | 1,970 | 1,983 | 9% |
| Apartment buildings | 50,389 | 1075.2 | 1927.6 | 292 | 1,120 | 63% | 1,954 | 27 | 1,936 | 1,968 | 20% |
| Usable floor space/building (m2) | Construction year | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| Building type | n | Mean | Std. | Q1 | Q3 | Missing | Mean | Std. | Q1 | Q3 | Missing |
| Detached houses | 1,982,870 | 115.4 | 50.2 | 81 | 143 | 15% | 1,955 | 35 | 1,930 | 1,977 | 18% |
| Terraced houses | 112,356 | 126.0 | 32.6 | 110 | 141 | 6% | 1,974 | 17 | 1,969 | 1,979 | 5% |
| Small houses | 4,499 | 348.1 | 846.3 | 145 | 324 | 86% | 1,993 | 38 | 1,977 | 2,013 | 27% |
| Row houses | 30,000 | 127.7 | 138.9 | 104 | 132 | 15% | 1,978 | 19 | 1,970 | 1,983 | 9% |
| Apartment buildings | 50,389 | 1075.2 | 1927.6 | 292 | 1,120 | 63% | 1,954 | 27 | 1,936 | 1,968 | 20% |
Note(s): The mean, standard deviation (std), 25% quantile (Q1), 75% quantile (Q3) and percentage of missing data for the usable floor space (m2) and construction year of all residential buildings based on building type are shown
4. Results
4.1 Classification model results
To train the classification model, both the features and construction year cohorts of each resampled dataset are randomly split into two parts, with 90% assigned to the training dataset and 10% to the validation dataset.
The performance levels of the trained models are shown in Table 5. The results show that the choice of under-sampling method has a relatively potent impact on the performance level of the classification model. The difference between the MMCCs of the best-performing NCR under-sampling method and the worst-performing sampling method is 70.1%. The reason for this disparity in performance level is likely due to the random nature of RUS in removing valuable and informative data-points from the sample. Based on all the evaluation metrics, the NCR produces the best results. Thus, the results obtained from the classification model that uses NCR for resampling are used for the subsequent regression model training.
Results of the evaluation metrics for each under-sampling method
| Under-sampling method | AUPRC | MAUC | G-mean | MMCC |
|---|---|---|---|---|
| RUS | 0.631 | 0.783 | 0.532 | 0.243 |
| Near Miss | 0.663 | 0.922 | 0.760 | 0.564 |
| OSS | 0.671 | 0.912 | 0.777 | 0.571 |
| NCR | 0.823 | 0.974 | 0.911 | 0.814 |
| Under-sampling method | G-mean | |||
|---|---|---|---|---|
| RUS | 0.631 | 0.783 | 0.532 | 0.243 |
| Near Miss | 0.663 | 0.922 | 0.760 | 0.564 |
| 0.671 | 0.912 | 0.777 | 0.571 | |
| 0.823 | 0.974 | 0.911 | 0.814 |
Note(s): AUPRC, Area Under the Precision-Recall Curve; MAUC, Mean Area Under the Curve; MMCC, Mean Mathews Correlation Coefficient; RUS, Random Under-sampler; OSS, One-sided Selection; NCR, Neighborhood Cleaning Rule. For each of the evaluation metrics, a higher score (closer to 1) represents a better prediction performance
To understand more clearly the classification model's performance for each class, the normalized confusion matrix for the best-performing model is presented in Figure 4. The confusion matrix is row-wise normalized to increase the readability of the plot. In a multiclass confusion matrix, the diagonal line represents the true predictions. In the confusion matrix, each row represents the true age cohort, while each column represents the age cohort predicted by the model. As shown in the confusion matrix, the two age cohorts for which the model performs poorly are 1910–19 and 2020–23. The 1910–19 cohort has significantly fewer samples compared to the previous cohort (1900–09), and the two cohorts likely have very similar features and, thus, are more difficult for the model to predict. Similarly, the 2020–23 cohort has a relatively low number of samples, although the model shows the worst performance for this cohort because around 16% of the wrong predictions are for the 1900–09 cohort.
The heatmap confusion matrix titled “Normalized Confusion Matrix” displays classification performance for 13 classes. The horizontal axis is labeled “Predicted Label” and ranges from 0 to 12 in increments of 1. The vertical axis is labeled “True Label” and ranges from 0 to 12 in increments of 1 from top to bottom. Each cell shows a normalized proportion between 0.00 and 0.94 and is color-coded using a gradient scale. A vertical color scale is shown on the right side of the heatmap, ranging from dark purple at the bottom to bright yellow at the top. The scale represents normalized values from 0.0 to 0.9, with darker colors indicating lower values and lighter yellow tones indicating higher values. Row for True Label 0 shows values 0.91, 0.00, 0.01, 0.01, 0.00, 0.01, 0.02, 0.03, 0.00, 0.00, 0.00, 0.00, and 0.00. Row for True Label 1 shows values 0.37, 0.29, 0.07, 0.07, 0.03, 0.03, 0.06, 0.07, 0.01, 0.00, 0.00, 0.00, and 0.00. Row for True Label 2 shows values 0.21, 0.00, 0.56, 0.06, 0.03, 0.03, 0.05, 0.04, 0.01, 0.00, 0.00, 0.00, and 0.00. Row for True Label 3 shows values 0.11, 0.00, 0.03, 0.67, 0.04, 0.03, 0.05, 0.05, 0.01, 0.00, 0.00, 0.00, and 0.00. Row for True Label 4 shows values 0.07, 0.00, 0.01, 0.05, 0.70, 0.05, 0.06, 0.05, 0.01, 0.00, 0.00, 0.00, and 0.00. Row for True Label 5 shows values 0.04, 0.00, 0.00, 0.01, 0.02, 0.76, 0.09, 0.06, 0.01, 0.00, 0.00, 0.00, and 0.00. Row for True Label 6 shows values 0.02, 0.00, 0.00, 0.00, 0.00, 0.02, 0.87, 0.07, 0.00, 0.00, 0.00, 0.00, and 0.00. Row for True Label 7 shows values 0.02, 0.00, 0.00, 0.00, 0.00, 0.00, 0.02, 0.94, 0.01, 0.00, 0.00, 0.00, and 0.00. Row for True Label 8 shows values 0.02, 0.00, 0.00, 0.00, 0.00, 0.01, 0.02, 0.08, 0.84, 0.01, 0.00, 0.00, and 0.00. Row for True Label 9 shows values 0.02, 0.00, 0.00, 0.00, 0.01, 0.01, 0.03, 0.05, 0.04, 0.82, 0.02, 0.00, and 0.00. Row for True Label 10 shows values 0.02, 0.00, 0.00, 0.01, 0.00, 0.01, 0.03, 0.04, 0.01, 0.01, 0.87, 0.01, and 0.00. Row for True Label 11 shows values 0.02, 0.00, 0.00, 0.00, 0.00, 0.00, 0.02, 0.03, 0.01, 0.00, 0.02, 0.87, and 0.01. Row for True Label 12 shows values 0.15, 0.00, 0.02, 0.03, 0.01, 0.04, 0.11, 0.11, 0.03, 0.01, 0.03, 0.08, and 0.39. The highest diagonal values are 0.94 for label 7, 0.91 for label 0, and 0.87 for labels 6, 10, and 11, shown in lighter yellow tones. The lowest diagonal value is 0.29 for label 1, shown in a darker green to blue tone.Row-wise normalized confusion matrix of prediction accuracy based on the model trained on the neighborhood cleaning rule for the sampled dataset
The heatmap confusion matrix titled “Normalized Confusion Matrix” displays classification performance for 13 classes. The horizontal axis is labeled “Predicted Label” and ranges from 0 to 12 in increments of 1. The vertical axis is labeled “True Label” and ranges from 0 to 12 in increments of 1 from top to bottom. Each cell shows a normalized proportion between 0.00 and 0.94 and is color-coded using a gradient scale. A vertical color scale is shown on the right side of the heatmap, ranging from dark purple at the bottom to bright yellow at the top. The scale represents normalized values from 0.0 to 0.9, with darker colors indicating lower values and lighter yellow tones indicating higher values. Row for True Label 0 shows values 0.91, 0.00, 0.01, 0.01, 0.00, 0.01, 0.02, 0.03, 0.00, 0.00, 0.00, 0.00, and 0.00. Row for True Label 1 shows values 0.37, 0.29, 0.07, 0.07, 0.03, 0.03, 0.06, 0.07, 0.01, 0.00, 0.00, 0.00, and 0.00. Row for True Label 2 shows values 0.21, 0.00, 0.56, 0.06, 0.03, 0.03, 0.05, 0.04, 0.01, 0.00, 0.00, 0.00, and 0.00. Row for True Label 3 shows values 0.11, 0.00, 0.03, 0.67, 0.04, 0.03, 0.05, 0.05, 0.01, 0.00, 0.00, 0.00, and 0.00. Row for True Label 4 shows values 0.07, 0.00, 0.01, 0.05, 0.70, 0.05, 0.06, 0.05, 0.01, 0.00, 0.00, 0.00, and 0.00. Row for True Label 5 shows values 0.04, 0.00, 0.00, 0.01, 0.02, 0.76, 0.09, 0.06, 0.01, 0.00, 0.00, 0.00, and 0.00. Row for True Label 6 shows values 0.02, 0.00, 0.00, 0.00, 0.00, 0.02, 0.87, 0.07, 0.00, 0.00, 0.00, 0.00, and 0.00. Row for True Label 7 shows values 0.02, 0.00, 0.00, 0.00, 0.00, 0.00, 0.02, 0.94, 0.01, 0.00, 0.00, 0.00, and 0.00. Row for True Label 8 shows values 0.02, 0.00, 0.00, 0.00, 0.00, 0.01, 0.02, 0.08, 0.84, 0.01, 0.00, 0.00, and 0.00. Row for True Label 9 shows values 0.02, 0.00, 0.00, 0.00, 0.01, 0.01, 0.03, 0.05, 0.04, 0.82, 0.02, 0.00, and 0.00. Row for True Label 10 shows values 0.02, 0.00, 0.00, 0.01, 0.00, 0.01, 0.03, 0.04, 0.01, 0.01, 0.87, 0.01, and 0.00. Row for True Label 11 shows values 0.02, 0.00, 0.00, 0.00, 0.00, 0.00, 0.02, 0.03, 0.01, 0.00, 0.02, 0.87, and 0.01. Row for True Label 12 shows values 0.15, 0.00, 0.02, 0.03, 0.01, 0.04, 0.11, 0.11, 0.03, 0.01, 0.03, 0.08, and 0.39. The highest diagonal values are 0.94 for label 7, 0.91 for label 0, and 0.87 for labels 6, 10, and 11, shown in lighter yellow tones. The lowest diagonal value is 0.29 for label 1, shown in a darker green to blue tone.Row-wise normalized confusion matrix of prediction accuracy based on the model trained on the neighborhood cleaning rule for the sampled dataset
The errors of the classification model are comparable to a similar study by Dabrock et al. (2025) that predicted construction years of residential buildings in Germany with morphological features, but with height data as input data (Dabrock et al., 2025). The results of Dabrock et al. (2025) show that their model achieved an F1-score (a similar evaluation metric) of 73.93%, and this accuracy is comparable to the results from our models.
4.2 Regression model results
The results for each regression model are presented in Table 6.
Results of the evaluation metrics for each regression machine learning algorithm
| ML models | MAE | RMSE | R2 |
|---|---|---|---|
| XGBoost | 28.74 | 101.04 | 0.789 |
| LightGBM | 29.43 | 101.43 | 0.788 |
| CatBoost | 31.19 | 102.92 | 0.781 |
| Random Forest | 32.16 | 108.81 | 0.756 |
| ML models | R2 | ||
|---|---|---|---|
| XGBoost | 28.74 | 101.04 | 0.789 |
| LightGBM | 29.43 | 101.43 | 0.788 |
| CatBoost | 31.19 | 102.92 | 0.781 |
| Random Forest | 32.16 | 108.81 | 0.756 |
Note(s): MAE, Mean absolute error; RMSE, Root mean-square error. For MAE and RMSE, the lower the absolute value, the better the performance, and for R2, the closer to 1, the better the results
The best overall best-performing model is XGBoost, while LightGBM achieves similar results. Random forest is the worst-performing model by a relatively wide margin. This is an expected result, since the other three algorithms are improvements over RF in their implementation. An R2 value of 0.789 can be considered as good performance. For comparison, in the study of Milojevic-Dupont and coworkers, the best-performing model for predicting building heights achieved an R2 of 0.72 (Milojevic-Dupont et al., 2020b).
To understand the performance of the model at a more granular level, the average prediction errors for floor space in all municipalities in Sweden are visualized in Figure 5. Prediction errors or residuals are defined as the ground truth floor space minus the predicted floor space. The average prediction errors for apartments are visualized separately from other types of residential buildings because the range of floor space values is much broader for apartments. The average prediction errors show large variations, with the minimum average error being 38.6 m2 and the maximum average error being 2,644.4 m2, with a standard deviation of 361.7 m2. The three municipalities with average prediction errors >2,000 m2 are all located in the Greater Stockholm area (Järfälla, Huddinge, Nacka). This is likely due to the large amounts of missing data in these municipalities, with 92.4%, 80.9% and 82.3%, respectively, of the data being unavailable. The average percentage of missing apartment floor space data for the eleven municipalities in the greater Stockholm area is 76.4%, which likely accounts for the high average error observed for these municipalities. The median prediction error is 383.4 m2 and the 75th percentile of the average errors is 544.2 m2, indicating that the model performs relatively well, apart from the municipalities with lots of missing data.
The two-panel choropleth map presents Sweden divided into municipal regions with vertical color bars indicating average error in square meters. Panel “(a) Apartments” displays regions shaded in a red gradient. The vertical legend on the right of panel (a) is labeled “Average error (in square meter)” and ranges from 0 at the bottom to 2500 at the top, with darker red indicating higher error values. Most regions across Sweden are light peach to light red, indicating lower error values, while darker red concentrations appear around the Stockholm region and its nearby coastal areas, where the highest values are shown. Panel “(b) Other buildings” displays regions shaded in a diverging blue-to-red gradient. The vertical legend on the right of panel (b) is labeled “Average error (in square meter)” and ranges from negative 25 at the bottom through 0 at the midpoint to 15 at the top with increments of 5. Dark blue represents larger negative errors, white represents values near 0, and red represents positive errors. In panel (b), northern and central regions are predominantly blue, indicating negative average errors, while scattered southern and central regions show light red tones indicating positive average errors. The Stockholm area inset in both panels highlights detailed regional variation, with panel (a) showing dark red clusters and panel (b) showing a mix of blue and light red areas.Average errors of the predictions of floor space in all municipalities in Sweden for: (a) apartment buildings and (b) all other types of residential buildings (in m2)
The two-panel choropleth map presents Sweden divided into municipal regions with vertical color bars indicating average error in square meters. Panel “(a) Apartments” displays regions shaded in a red gradient. The vertical legend on the right of panel (a) is labeled “Average error (in square meter)” and ranges from 0 at the bottom to 2500 at the top, with darker red indicating higher error values. Most regions across Sweden are light peach to light red, indicating lower error values, while darker red concentrations appear around the Stockholm region and its nearby coastal areas, where the highest values are shown. Panel “(b) Other buildings” displays regions shaded in a diverging blue-to-red gradient. The vertical legend on the right of panel (b) is labeled “Average error (in square meter)” and ranges from negative 25 at the bottom through 0 at the midpoint to 15 at the top with increments of 5. Dark blue represents larger negative errors, white represents values near 0, and red represents positive errors. In panel (b), northern and central regions are predominantly blue, indicating negative average errors, while scattered southern and central regions show light red tones indicating positive average errors. The Stockholm area inset in both panels highlights detailed regional variation, with panel (a) showing dark red clusters and panel (b) showing a mix of blue and light red areas.Average errors of the predictions of floor space in all municipalities in Sweden for: (a) apartment buildings and (b) all other types of residential buildings (in m2)
The average prediction errors for other types of buildings (excluding apartments) show a much narrower range of variations, with minimum average error of −29.3 m2 and maximum average error of 7.99 m2 and a standard deviation of 6.05 m2. The model on average underpredicts the floor space areas of other types of buildings, with a mean of −7.39 m2 across all municipalities. There are also weaker spatial correlations for the prediction errors, unlike the case for apartment buildings.
The total usable floor space, average construction year and the floor space per capita of the final completed dataset (combination of the ground truth values and the predicted values) are visualized for each municipality in Figure 6. The total usable floor space in km2 is presented in Figure 6a. The municipality with the largest total usable floor space is Stockholm, with 13.7 km2.
The three-panel choropleth map displays “Sweden” divided into “municipal regions”, with separate vertical color bars for each panel. Panel “(a) Total usable floor space” shows municipalities shaded using a purple-to-yellow gradient. The vertical legend on the right is labeled “Usable floor space square kilometer” and ranges from 0 at the bottom to 12 at the top with increments of 2. Dark purple indicates values near 0, while yellow indicates values near 12. Most northern municipalities appear dark purple, indicating low usable floor space, while higher values shown in green to yellow tones are concentrated around “southern Sweden” and particularly around the “Stockholm region”, which is highlighted by a circular inset. Panel “(b) Floor space per capita” uses a similar purple-to-yellow gradient. The vertical legend is labeled “square meter per capita” and ranges from 25 at the bottom to 225 at the top with increments of 25. Lower per-capita values appear in dark purple and blue shades, mainly in “southern” and “metropolitan areas”, while higher values in green to yellow tones appear in several “central” and “northern municipalities”. The “Stockholm inset” shows mixed values with predominantly lower to mid-range tones. Panel “(c) Average construction year” uses the same gradient style. The vertical legend is labeled “Year” and ranges from 1935 at the bottom to 1975 at the top with increments of 5. Dark purple corresponds to earlier construction years around 1935, while yellow corresponds to later years around 1975. Earlier average construction years dominate many “southern” and “central municipalities”, while later construction years in green to yellow tones appear in parts of “northern Sweden” and within the “Stockholm inset”.Total usable floor space, construction year and floor space values visualized by municipality: (a) total usable floor space (km2); (b) floor space per capita (m2) and (c) average construction year. The results for the Greater Stockholm area are zoomed in to increase visibility
The three-panel choropleth map displays “Sweden” divided into “municipal regions”, with separate vertical color bars for each panel. Panel “(a) Total usable floor space” shows municipalities shaded using a purple-to-yellow gradient. The vertical legend on the right is labeled “Usable floor space square kilometer” and ranges from 0 at the bottom to 12 at the top with increments of 2. Dark purple indicates values near 0, while yellow indicates values near 12. Most northern municipalities appear dark purple, indicating low usable floor space, while higher values shown in green to yellow tones are concentrated around “southern Sweden” and particularly around the “Stockholm region”, which is highlighted by a circular inset. Panel “(b) Floor space per capita” uses a similar purple-to-yellow gradient. The vertical legend is labeled “square meter per capita” and ranges from 25 at the bottom to 225 at the top with increments of 25. Lower per-capita values appear in dark purple and blue shades, mainly in “southern” and “metropolitan areas”, while higher values in green to yellow tones appear in several “central” and “northern municipalities”. The “Stockholm inset” shows mixed values with predominantly lower to mid-range tones. Panel “(c) Average construction year” uses the same gradient style. The vertical legend is labeled “Year” and ranges from 1935 at the bottom to 1975 at the top with increments of 5. Dark purple corresponds to earlier construction years around 1935, while yellow corresponds to later years around 1975. Earlier average construction years dominate many “southern” and “central municipalities”, while later construction years in green to yellow tones appear in parts of “northern Sweden” and within the “Stockholm inset”.Total usable floor space, construction year and floor space values visualized by municipality: (a) total usable floor space (km2); (b) floor space per capita (m2) and (c) average construction year. The results for the Greater Stockholm area are zoomed in to increase visibility
The floor space per capita for each municipality is presented in Figure 6b. The mean floor space per capita for all municipalities is 66.5 m2, with the minimum being 11.6 m2 per capita in a suburb of Stockholm (Sundbyberg) and the maximum being 229.5 m2 for a rural municipality with a population of only 2,387 people (Sorsele). To validate the results, we compare the results to the average useful floor space per person from Statistics Sweden (SCB). According to the latest data at the end of 2022, the average useful floor space per person in Sweden was 42 m2 (Sweden, 2023). The floor space per capita from the final dataset is 41.8 m2, which indicates that the model performed well across the entire dataset. However, the minimum and maximum values show that the model's performance can be improved. For Sundbyberg municipality, the floor space per capita from the results is 11.6 m2, whereas the SCB floor space per capita is 33 m2. Similarly, for Sorsele, the floor space per capita from the results is 229.5 m2, whereas the SCB floor space per capita is 48 m2. These discrepancies are most likely due to errors in the apartment floor space prediction, as discussed above.
As expected, Figure 6c shows that the municipalities around the three major cities (Stockholm, Gothenburg and Malmö) have the highest/newest average construction year at around the year 1970. The municipalities with the lowest average construction years are all rural municipalities where the demand for new housing is relatively low. The average population of the 10 municipalities with the lowest average construction year is only 8,554 people.
4.3 Feature importance
Plotting the feature importance of the ML model can help elucidate which features the model relies on the most when making predictions. Interpretability and explainability are important to foster trust in sustainability insights and are especially important when ML is employed to enhance decision-making (Donati et al., 2022). In this section, we present feature-importance plots for both the classification and regression models to improve the transparency of the results. We adopt the definition proposed by Doshi-Velez and Kim (2017) for interpretability to discuss feature importance, where interpretability is defined as the ability to explain an ML system in terms understandable to a human. Therefore, this section focuses on whether the feature importance can be corroborated with intuitive interpretations based on domain knowledge.
The top 20 most important features from the normalized feature importance plots are presented in Figure 7, for Classification (Figure 7a) and Regression (Figure 7b). A complete list of the feature importance values can be found in the Supplementary Information. The two features with the highest predictive power within the classification model (Figure 7a) are the X and Y coordinates of each building. This is somewhat expected as buildings that are located in close proximity to one another tend to have similar construction years (Nachtigall et al., 2023). Furthermore, the tessellation and block footprint area have higher predictive power than the building's footprint area. The most surprising result is that half of the most-predictive features are block features. We hypothesize that the model is predicting that the buildings in each block have similar construction years, which is a logical assumption, and since the model predicts the construction year in decades, this means that it is more likely that a building within a block will belong to the same decade of construction.
The two-panel horizontal bar chart shows feature importance rankings for classification in panel “(a)” and regression in panel “(b)”. In panel “(a)”, titled “Top 20 Feature Importance Classification”, the horizontal axis is labeled “Feature Importance” and ranges from 0.00 to 0.08 in increments of 0.02. The vertical axis lists features from top to bottom as follows: “X coordinate”, “Y coordinate”, “TessFootprintArea”, “BlockFootprintArea”, “Building type number”, “Municipality”, “BlockOrientation”, “TessPerimeter”, “BlockCorners”, “FootprintArea”, “BlockRectangularity”, “Block underscore ccd underscore means”, “BlockConvexity”, “BlockElongation”, “c c d underscore means”, “Block underscore c c d underscore stdev”, “BlockEquivalentRectangularIndex”, “BlockCircularCompactness”, “LongestAxisLength”, and “County”. The corresponding feature importance values from top to bottom are as follows: “X coordinate”: 0.078, “Y coordinate”: 0.075, “TessFootprintArea”: 0.071, “BlockFootprintArea”: 0.038, “Building type number”: 0.037, “Municipality”: 0.036, “BlockOrientation”: 0.035, “TessPerimeter”: 0.034, “BlockCorners”: 0.034, “FootprintArea”: 0.032, “BlockRectangularity”: 0.031, “Block underscore c c d underscore means”: 0.030, “BlockConvexity”: 0.029, “BlockElongation”: 0.028, “c c d underscore means”: 0.028, “Block underscore c c d underscore stdev”: 0.027, “BlockEquivalentRectangularIndex”: 0.026, “BlockCircularCompactness”: 0.026, “LongestAxisLength”: 0.025, and “County”: 0.024. In panel “(b)”, titled “Top 20 Feature Importance Regression”, the horizontal axis is labeled “Feature Importance” and ranges from 0.00 to 0.25 in increments of 0.05 units. The vertical axis lists features from top to bottom as follows: “FootprintArea”, “Building type number”, “Construction year”, “LongestAxisLength”, “Y coordinate”, “County”, “Perimeter”, “c c d underscore means”, “X coordinate”, “CircularCompactness”, “Elongation”, “Municipality”, “FractalDimension”, “Convexity”, “Orientation”, “TessFootprintArea”, “Rectangularity”, “BlockRectangularity”, “BlockFootprintArea”, and “BlockOrientation”. The corresponding feature importance values from top to bottom are as follows: “FootprintArea”: 0.28, “Building type number”: 0.26, “Construction year”: 0.05, “LongestAxisLength”: 0.04, “Y coordinate”: 0.02, “County”: 0.02, “Perimeter”: 0.02, “c c d underscore means”: 0.02, “X coordinate”: 0.02, “CircularCompactness”: 0.01, “Elongation”: 0.01, “Municipality”: 0.01, “FractalDimension”: 0.01, “Convexity”: 0.01, “Orientation”: 0.01, “TessFootprintArea”: 0.01, “Rectangularity”: 0.01, “BlockRectangularity”: 0.01, “BlockFootprintArea”: 0.01, “BlockOrientation”: 0.01. Note: All numerical data values are approximated.Top 20 normalized feature importance plots from XGBoost: (a) classification model feature importance and (b) regression model feature importance. Note that the values on the x-axes are different
The two-panel horizontal bar chart shows feature importance rankings for classification in panel “(a)” and regression in panel “(b)”. In panel “(a)”, titled “Top 20 Feature Importance Classification”, the horizontal axis is labeled “Feature Importance” and ranges from 0.00 to 0.08 in increments of 0.02. The vertical axis lists features from top to bottom as follows: “X coordinate”, “Y coordinate”, “TessFootprintArea”, “BlockFootprintArea”, “Building type number”, “Municipality”, “BlockOrientation”, “TessPerimeter”, “BlockCorners”, “FootprintArea”, “BlockRectangularity”, “Block underscore ccd underscore means”, “BlockConvexity”, “BlockElongation”, “c c d underscore means”, “Block underscore c c d underscore stdev”, “BlockEquivalentRectangularIndex”, “BlockCircularCompactness”, “LongestAxisLength”, and “County”. The corresponding feature importance values from top to bottom are as follows: “X coordinate”: 0.078, “Y coordinate”: 0.075, “TessFootprintArea”: 0.071, “BlockFootprintArea”: 0.038, “Building type number”: 0.037, “Municipality”: 0.036, “BlockOrientation”: 0.035, “TessPerimeter”: 0.034, “BlockCorners”: 0.034, “FootprintArea”: 0.032, “BlockRectangularity”: 0.031, “Block underscore c c d underscore means”: 0.030, “BlockConvexity”: 0.029, “BlockElongation”: 0.028, “c c d underscore means”: 0.028, “Block underscore c c d underscore stdev”: 0.027, “BlockEquivalentRectangularIndex”: 0.026, “BlockCircularCompactness”: 0.026, “LongestAxisLength”: 0.025, and “County”: 0.024. In panel “(b)”, titled “Top 20 Feature Importance Regression”, the horizontal axis is labeled “Feature Importance” and ranges from 0.00 to 0.25 in increments of 0.05 units. The vertical axis lists features from top to bottom as follows: “FootprintArea”, “Building type number”, “Construction year”, “LongestAxisLength”, “Y coordinate”, “County”, “Perimeter”, “c c d underscore means”, “X coordinate”, “CircularCompactness”, “Elongation”, “Municipality”, “FractalDimension”, “Convexity”, “Orientation”, “TessFootprintArea”, “Rectangularity”, “BlockRectangularity”, “BlockFootprintArea”, and “BlockOrientation”. The corresponding feature importance values from top to bottom are as follows: “FootprintArea”: 0.28, “Building type number”: 0.26, “Construction year”: 0.05, “LongestAxisLength”: 0.04, “Y coordinate”: 0.02, “County”: 0.02, “Perimeter”: 0.02, “c c d underscore means”: 0.02, “X coordinate”: 0.02, “CircularCompactness”: 0.01, “Elongation”: 0.01, “Municipality”: 0.01, “FractalDimension”: 0.01, “Convexity”: 0.01, “Orientation”: 0.01, “TessFootprintArea”: 0.01, “Rectangularity”: 0.01, “BlockRectangularity”: 0.01, “BlockFootprintArea”: 0.01, “BlockOrientation”: 0.01. Note: All numerical data values are approximated.Top 20 normalized feature importance plots from XGBoost: (a) classification model feature importance and (b) regression model feature importance. Note that the values on the x-axes are different
The feature importance plot for the regression model to predict floor space is presented in Figure 7b. For the regression model, the top two features of footprint area and building type contribute more than 50% of the total predictive power of the model. The usable floor space of a building should show a strong correlation to its footprint area, especially for detached houses. Surprisingly, the specific type of building has a very strong impact on the model predictions. We hypothesize that this is mainly due to the unique nature of the data for apartment buildings existing as whole buildings, such that the floor space values for apartments are magnitudes higher than those for other types of buildings.
The result of the feature importance analysis shows that the urban morphology features are predictive for the classifications of building construction year and usable floor space. The predictive power of urban morphology features should be further tested on a dataset that does not represent the apartment floor space in an aggregate form.
5. Discussion
5.1 Result validation and comparison
This subsection discusses the validation of results compared to similar studies in the context of difference in training features. At the time writing, this study is the only work that does not use 3D (height) data for building floor area or age prediction according to the author's knowledge. Therefore, the results cannot be directly compared to existing studies. We instead compare the results in the context of discussing the strength and weakness of using height data as training feature. As previously discussed, Dabrock et al. (2025) is the only similar study using morphological features to predict building construction years at the national level, but they use height as input data (Dabrock et al., 2025). In addition, the height data used in Dabrock et al. (2025) was imputed using machine learning in a previous study (Dabrock et al., 2024). The similarities between the evaluation metrics results (F1 score of 73.93% in (Dabrock et al., 2025) vs our AUPRC of 0.823) show that models that only use 2D data are sufficient to predict building attributes to the same degree of accuracy as models that use 3D data.
The usable floor space prediction errors are previously compared to the results from Milojevic-Dupont et al. (2020b) and the R2 values are also similar (0.72 vs 0.789). Research in more geographical areas is needed to confirm this finding, and ideally in large geographical areas with official height data. An alternative approach is to use ML-predicted and imputed building height datasets, such as the 3D-GloBFP dataset (Che et al., 2024) as training features to predict construction year and usable floor space and compare the results. However, using an imputed dataset carries the risk of propagating the uncertainties forward and increasing the errors in the final predictions. Based on the comparison of results with existing literature, we theorize that 2D data and urban form features are sufficient for predicting construction year and usable floor space, but 3D height data will be complementary to model by adding more predictive power. Therefore, the proposed approach should be used and tested in geographical areas where accurate and reliable height data are not available.
5.2 Generalizability and limitations of the machine learning method
A key rationale and advantage for developing the machine learning method that does not require 3D building data is that this approach could be easily tested in other geographical areas. However, as Nachtigall et al. (2023) highlight, urban form feature distributions can vary significantly between countries, which makes it difficult to generalize across regions, but having access to more data for the new region that is being generalized to can improve prediction accuracy (Nachtigall et al., 2023). By not requiring building height data that is scarcely available, our proposed approach offers higher potential for generalizability compared to approaches by previous studies (Nachtigall et al., 2023; Dabrock et al., 2025).
The lack of a generalizability test is a main limitation of this work. Unlike other ML workflows that predict building attributes at the urban or city level that test generalizability within the same region or country (Rosser et al., 2019), testing the generalizability would require obtaining a building registry dataset that contains both construction year and usable floor space for a different country. A potential option is to use OpenStreetMap (OSM) data, but the data quality and completeness of OSM vary greatly across countries and thus might not be reliable as input data (Biljecki et al., 2023). We suggest that the proposed ML workflow be tested in Norway or Finland, where the geography and social-demographic factors are similar to those in Sweden.
6. Conclusion
The decarbonization of the building sector requires the availability of more detailed building stock data to support Science-based decision-making. Machine learning is emerging as a tool to fill in the gaps in incomplete datasets for building stock modeling. Two relevant data variables are the construction year and the usable floor space. The construction year is important for understanding the material composition and energy efficiency of buildings, while usable floor space is commonly used as the basis for calculating both the energy demand and material stock of buildings. The application of ML methods to predict these variables has often been limited to the city level due to the prevalence of 3D data, such as height, which is not widely available as a predictive variable.
In this study, we present a workflow to predict both the building construction year and usable floor space without using 3D data, using residential buildings in Sweden as a case study. The construction year prediction is treated as a classification issue, while the usable floor space prediction is treated as a regression issue. The results suggest that urban form attributes are good predictors of building characteristics and are suitable as an alternative to using height data as predictors, especially at the national level. Based on the findings of this article, we can conclude that 2D data is sufficient for predicting important residential building attributes, such as construction year and usable floor area, which are crucial for many fields of study. The implication and relevance of this conclusion is that the applicability of ML models to impute missing building data can be extended to geographical areas where height data are unavailable or inaccurate. The proposed approach should also be tested and compared with ML methods that use 3D data to compare the accuracy of each approach. A key area of application for this approach and dataset is for building material stock and flow modeling. Existing bottom-up building material stock and flow rely on material intensity datasets that are aggregated based on age-cohorts that are typically in 10-year intervals and are thus more tolerant of errors in construction year predictions.
In future work, the completed building stock dataset will be used as the basis for a material stock and flow study for Sweden. The proposed approach could also be generalized and tested for other geographic areas where data is available. Furthermore, other urban form attributes and ML algorithms should be tested depending on data availability. The resulting dataset will be used in future work to quantify the material stock and flow of residential buildings in Sweden to analyze the potential to reduce embodied emissions. The dataset could also be used to analyze questions such as the benefits and trade-offs of renovating residential buildings (Chen and Lai, 2025).
The supplementary material for this article can be found online.

