Exposure to extreme heat poses significant health risks for construction workers due to multiple reasons. However, research on heat-related illnesses (HRIs) in this sector has often relied on statistical methods, limiting insights into causal mechanisms and feature interrelations. This study addresses this gap by employing machine learning (ML) techniques to predict the severity of HRIs (fatal vs non-fatal) and identify critical contributing factors.
A total of 370 Occupational Safety and Health (OSHA) HRI reports are analyzed using natural language processing (NLP) to extract key features. These features are combined with climatic and non-climatic variables. Four ML algorithms – random forest, extreme gradient boosting, k-nearest neighbors and logistic regression – are applied to classify HRIs as fatal and non-fatal. Feature importance and dependencies are explored using SHAP (SHapley Additive exPlanations).
The random forest model achieves the highest accuracy at 70%. Employers' negligence and workers aged over 35 years are key human-centric factors in fatal HRIs. Humid subtropical and hot-summer Mediterranean climate zones are the most hazardous environmental factors. The results suggest that small-scale projects with limited budgets, renovation projects and building projects (residential/commercial) are the key contributing project-related features in the severity of HRIs.
This study is among the first attempts to investigate HRI reports specifically in the construction industry by using artificial intelligence techniques. The findings emphasize the need for tailored heat stress prevention programs in construction projects based on their characteristics.
1. Introduction
Global average temperatures are rising due to climate change, resulting in more frequent and severe extreme heat events. These events are expected to increase the incidence of occupational incidents (OIs) (Gibb et al., 2024). A positive correlation between OIs and extreme heat exposure has been documented, primarily attributed to biological mechanisms such as fatigue and reduced concentration (Varghese et al., 2018). The construction industry is particularly susceptible to heat-related risks since it is characterized by physically demanding tasks that are carried out under extreme climatic conditions (Shakerian et al., 2021). Studies have reported a rise in safety incidents in this sector during periods of extreme heat (Rameezdeen and Elmualim, 2017). Given the adverse impacts of climate change, there is an urgent need to develop predictive and preventive strategies to address heat-related illnesses (HRIs) within the construction industry.
Analyzing OI reports has long been an established approach for predicting future workplace accidents (Goh and Ubeynarayana, 2017). Although considerable efforts have been made to associate construction-related accidents – including falls, trips and struck by objects – with extreme heat exposure, relatively few studies have focused specifically on HRIs, such as heat stroke and heat exhaustion. Furthermore, most existing research in this domain relies on descriptive statistical analysis (Kim et al., 2024), which may have limited capacity to uncover latent relationships between contributing factors. For instance, some studies have employed time-series methods to estimate the relative risk of occupational injuries associated with extreme outdoor temperatures (Marinaccio et al., 2019) and heat waves (Gariazzo et al., 2023). However, these analyses rarely explore the interactions among contributing features in the occurrence of such incidents. In contrast, recent advances in AI have demonstrated their effectiveness in mining data, identifying accident precursors and analyzing their complex interactions that ultimately contributing more effective safety management strategies (Tselentis et al., 2023).
Accordingly, the aim of this study is to identify key contributing factors to the severity of HRIs in the construction industry. We address this aim through the integration of structured and unstructured data sources, enabling a comprehensive analysis of contributing factors. To support this, machine learning (ML) techniques are applied as analytical tools for modeling the underlying relationships. The potential precursors are identified by analyzing both non-climatic and climatic factors. The non-climatic factors were extracted from historical HRI reports published by the Occupational Safety and Health Administration (OSHA) (OSHAGroup, 2024a) alongside climatic factors including climate zones (PlantmapsGroup, 2024) and high ambient temperature on the incident dates. The temperature data were retrieved from the official tool provided by the US National Weather Services (NWSGroup, 2024).
The remainder of the paper is organized into several sections. Section 2 presents a review of HRIs and the analysis of construction incident reports. Section 3 outlines the research methodology, including the development of a natural language processing (NLP) approach for extracting non-climatic features from OSHA reports as a foundational step in data preprocessing. These NLP-extracted features are then integrated with climatic data to develop ML models for predicting the severity of the HRIs. Section 4 evaluates the contribution of individual features using the SHAP (SHapley Additive exPlanations) framework and presents the results of the data analysis. Section 5 offers practical implications and recommendations for future research. Finally, Section 6 presents the study's conclusions and outlines its limitations.
2. Literature review
This section provides an overview of key concepts related to the assessment of HRIs. It explores the identification of extreme heat exposure, followed by a discussion of occupational HRIs and the application of AI techniques for predicting and preventing HIRs in the construction industry. This information establishes the conceptual framework for understanding the objectives of this study.
2.1 Heat-related illnesses and extreme heat exposure
Heat stress refers to the cumulative heat exposure an individual experiences from both external and internal sources, including ambient temperature, humidity, direct sunlight and physical activity (Coca et al., 2016). The human body responds to heat stress through physiological adaptations like sweating and increasing heartbeat rate – responses collectively referred to as heat strain (OSHAGroup, 2023b). However, when these mechanisms fail to maintain a stable core body temperature, HRIs occur. HRIs can arise at a wide range of ambient temperatures and vary in severity, from mild heat cramps to fatal heat strokes.
Although the physiological interaction between heat stress and heat strain is well established, translating the effects of heat stress into the understanding of heat strain remains complex (Epstein and Moran, 2006). Multiple factors contribute to the onset of extreme heat stress. In occupational safety research, efforts to formulate heat stress have been ongoing since 1905, initially focusing on climatic variables such as ambient temperature, humidity and air velocity (Jacklitsch et al., 2016). Over time, researchers have acknowledged that additional work-related and worker-specific factors also play a critical role in determining heat susceptibility. These factors include alcohol consumption, age, work experience (Asghari et al., 2017; Chan et al., 2013), body posture, the physical demands of tasks and preexisting medical conditions (Borg et al., 2023; OSHAGroup, 2023a). Despite significant progress, developing a universal heat index remains a challenge due to the complex interrelation of human and environmental variables (Epstein and Moran, 2006). As a result, each occupational sector requires a tailored formula for assessing heat stress and its potential health impacts.
High ambient temperature, commonly used as an index for extreme heat exposure, is a key climatic variable OI analysis. This approach typically involves establishing a threshold temperature and examining the relationship between OI rates and increases in temperature beyond that threshold. Notably, these thresholds vary by geographic region. For instance, in Adelaide, Australia, Xiang et al. (2014) realized that within a temperature range of 14.2 °C–37 °C, each 1°C increase in the daily maximum temperature was associated with a heightened risk of OIs. In contrast, in Quebec, Canada, a similar correlation between rising ambient temperatures and OI risk was observed within a higher threshold range of 33°C–38°C (Adam-Poupart et al., 2015). Fatima et al. (2021) further emphasized that extreme heat stress is region-specific and varies according to climate zones. Using the Köppen–Geiger Climate Classification (KGCC), Beck et al. (2018) identified that extreme heat exposure poses a significant risk in temperate climates. Specifically, humid subtropical and oceanic climate zones exhibit a stronger association between rising temperatures and the risk of OIs compared to other climatic regions (Fatima et al., 2021).
2.2 Incorporating heat stress to occupational incidents
Heat stress can lead to both direct and indirect illnesses and injuries. Direct effects include conditions such as heat cramps and fatal heat stroke cases (Dong et al., 2019), while indirect consequences may involve incidents such as loss of control, slips, collapses or falls from height (Gariazzo et al., 2023). The availability of OI data, including records of HRIs, has enabled researchers to investigate the relationship between OIs and various climatic variables. For example, Rameezdeen and Elmualim (2017) examined the severity of OIs during heatwaves in Adelaide, Australia, by analyzing 1,418 compensation claims in the construction industry from 2002 to 2013 across 29 heatwave events. They found that although the number of claims filed by experienced workers increased by 1.6%, the overall severity of these incidents declined by 2.8%. In a separate study, Lee et al. (2021) analyzed the seasonal distribution of OIs in South Korea and reported that extreme climatic conditions in summer and winter increased the likelihood of OIs. However, both fatal and non-fatal accidents were more frequently reported in autumn, a season marked by milder temperatures and relative humidity. This trend was attributed to mandatory construction shutdowns during summer when temperatures exceeded 35°C and in winter under strong wind conditions (above 10 m/s) (Lee et al., 2021).
Although numerous studies have explored the relationship between OIs and heat stress, relatively few have focused specifically on HRI reports in the construction industry. Tymvios et al. (2016) analyzed 58 HRI cases from OSHA's “Investigation Summaries” dataset from 2009 to 2013. They found that the majority of the incidents (53%) occurred on small construction projects with budgets under US$50,000, and most cases took place during summer (84%). In a separate study, Dong et al. (2019) identified 285 heat-related fatalities in the construction sector using data from the Census of Fatal Occupational Injuries between 1992 and 2016. Among the most at-risk occupations were cement and brick masons, as well as roofers – groups found to experience the highest rates of heat-related deaths (Dong et al., 2019; Tripathi and Mittal, 2024).
One potential reason for the limited focus on HRI reports in the construction sector is the relatively small number of documented HRIs in occupational health and safety databases. For instance, Safe Work Australia's “Worker's Compensation” dataset includes over 188,000 claims filed between 2008 and 2022. However, only 1,814 of these are attributed to the mechanism of “heat, electricity and other environmental factors” with no clear distinction for claims specifically linked to climatic heat exposure (SWAGroup, 2024). Similarly, OSHA's “Investigation Summaries” for the construction industry contain more than 27,300 reports from 1984 to 2023, of which only 371 involve HRIs (OSHAGroup, 2024b). This scarcity of documented cases has contributed to the underrepresentation of HRIs in construction safety research and analyses.
2.3 Application of artificial intelligence in extreme heat exposure studies
AI is revolutionizing various industries by addressing complex challenges, including those related to health and safety (Abioye et al., 2021). Although the construction industry remains one of the least digitized sectors, there has been a growing body of research employing AI to mitigate the risks associated with extreme heat exposure among workers (Zhou et al., 2023). A prominent application of AI in this field is the development of early-warning systems through experimental approaches. For instance, Yi et al. (2016) integrated environmental sensors with wearable biosensors to monitor workers' heat strain in hot and humid conditions. They developed an artificial neural network model capable of alerting safety managers through a monitoring platform whenever a worker's heart rate exceeded a predefined threshold. Similarly, Shakerian et al. (2021) utilized biosensors to track workers' physical and physiological parameters, applying supervised learning algorithms to predict heat strain states. Expanding on this concept, Seo et al. (2024) conducted experiments in controlled environments using multiple ML algorithms to predict individual heat strain levels based on personal characteristics such as age, height and weight.
Despite their promising outcomes, the reviewed studies did not validate their predictions against actual HRIs. In a recent study, Lyu and Song (2024) analyzed historical occupational HRIs from OSHA by applying AI techniques, including text mining and four model-free ML algorithms. While their findings are commendable, the prediction model was developed using data aggregated across multiple industries. Given the substantial differences in work environments and tasks between the construction sector and other industries – such as agriculture or fire protection – developing prediction models focused on a single industry can yield more precise and effective insights for identifying risk factors and preventing HRIs.
Despite growing research on HRIs, a critical gap remains in understanding how climatic and non-climatic factors interact and contribute to HRI severity in the construction sector. The current study addresses this gap by integrating factors from diverse structured and unstructured data sources and applying ML techniques as analytical tools to explore their complex interdependencies.
3. Data collection and research method
This section outlines the methodology adopted to address the research objectives. It begins with a detailed description of the dataset development process, followed by the presentation of the NLP-based approach used for feature extraction. Subsequently, the ML prediction models are introduced, and finally, a comparative analysis of their performance results is discussed.
3.1 Dataset development
The first step in this research involved constructing a dataset of HRIs that includes both climatic and non-climatic features. The publicly available OSHA “Investigation Summaries” database (OSHAGroup, 2024a) was utilized to extract HRI cases specific to the construction industry, which were identified using the corresponding industry classification code (NAICSAssociation, 2024). To distinguish HRIs from other types of injuries, a set of targeted keywords was applied from the database's toolkit. These keywords were “Heat”, “Heat Exhaustion”, “Heat Injury Prevention Program”, “Heat Stroke”, “Heat Index” and “Heat-related illness”. This filtering process reduced the initial pool of over 26,000 construction incident reports to 429 relevant cases. No date restrictions were imposed during data collection to ensure comprehensive coverage. After carefully reviewing these reports, 370 reports were selected for inclusion in the final dataset. Excluded reports were either misclassified or unrelated to heat exposure – such as those involving burns or gas tank explosions – which fall outside the scope of this study. The finalized reports span from 2001 to 2023, with one exception, from 1997, deliberately removed due to the unavailability of climatic data for that specific date and location. Among the selected cases, 197 involved fatal incidents, while 173 were non-fatal. A detailed list of the extracted features is provided in Table 1.
Contributing input and target features
| HRI precursor levels | Precursor types | Precursor categories | Features | Feature type |
|---|---|---|---|---|
| Non-climatic | Construction project | Project end-use | Tower, tank, storage elevator; sewer/water treatment plant; refinery; other buildings; single family or duplex dwelling; multi-family dwelling; commercial building; powerline, transmission line; pipeline; powerplant; manufacturing plant; highway, road, street; bridge; other heavy construction | Categorical |
| Project cost | Under $50,000; $50,000 to $250,000; $250,000 to $500,000; $500,000 to $1M; $1M to $5M; $5M to $20M; over $20M | Categorical | ||
| Project type | New project or new addition; alteration or rehabilitation (renovation); maintenance or repair; demolition; other | Categorical | ||
| Individual | Age | Younger than 35; older than 35 | Binary | |
| NLP-extracted phrases | Activity-related | A1-A12 | Categorical | |
| Worker-related | W1-W10 | Categorical | ||
| Employer-related | E1-E8 | Categorical | ||
| Climatic | Daily temperature | Temperature anomaly | Positive temperature departure; negative temperature departure | Binary |
| Climate zone | KGCC | Af; Am; Aw; BSh; BSk; BWh; BWk; Cfa; Csa; Csb; Dfa; Dfb | Categorical | |
| Target feature | Incident result | Level of HRI | Fatal; non-fatal | Binary |
| HRI precursor levels | Precursor types | Precursor categories | Features | Feature type |
|---|---|---|---|---|
| Non-climatic | Construction project | Project end-use | Tower, tank, storage elevator; sewer/water treatment plant; refinery; other buildings; single family or duplex dwelling; multi-family dwelling; commercial building; powerline, transmission line; pipeline; powerplant; manufacturing plant; highway, road, street; bridge; other heavy construction | Categorical |
| Project cost | Under $50,000; $50,000 to $250,000; $250,000 to $500,000; $500,000 to $1M; $1M to $5M; $5M to $20M; over $20M | Categorical | ||
| Project type | New project or new addition; alteration or rehabilitation (renovation); maintenance or repair; demolition; other | Categorical | ||
| Individual | Age | Younger than 35; older than 35 | Binary | |
| NLP-extracted phrases | Activity-related | A1-A12 | Categorical | |
| Worker-related | W1-W10 | Categorical | ||
| Employer-related | E1-E8 | Categorical | ||
| Climatic | Daily temperature | Temperature anomaly | Positive temperature departure; negative temperature departure | Binary |
| Climate zone | KGCC | Af; Am; Aw; BSh; BSk; BWh; BWk; Cfa; Csa; Csb; Dfa; Dfb | Categorical | |
| Target feature | Incident result | Level of HRI | Fatal; non-fatal | Binary |
The OSHA database classifies certain variables, including project end-use, project type and project cost. However, it does not provide predefined categories for the ages of the victims. Given the absence of a universally accepted standard for age categorization, researchers often adopt tailored approaches based on the objectives of their studies. Physiological research indicates that susceptibility to heat stress begins to rise around the fourth decade of life (Flouris et al., 2018). Older individuals tend to retain more heat than younger counterparts in dry and humid environments, thereby experiencing greater heat strain (Stapleton et al., 2014). In their analysis of the catastrophic 2003 heatwave in France, Fouillet et al. (2006) observed a significant increase in fatal heat-related incidents notably among individuals aged 35 and older. Based on this evidence, this study adopts 35 years of age as the threshold for categorizing victims’ ages.
As outlined in the literature review (Section 2.1), temperature anomaly data were collected for each incident report. To achieve this, daily temperature departure values were retrieved from the National Weather Service's official toolkit (NWSGroup, 2024). This metric is statistically more applicable than considering daily maximum or daily average temperatures for the incidents of this dataset (Mirhosseini et al., 2024). In this database, the departure temperature represents the difference between the observed average daily temperature and the long-term historical average for the period 1981–2010, measured in degrees Fahrenheit (NOAAGroup, 2024). The departure index serves as an indicator of temperature anomalies: positive values denote warmer-than-average days, while negative values indicate cooler-than-average days. In this study, incident days with positive departure temperatures were classified as “hot days”.
Additionally, climate zone information was incorporated by extracting KGCC codes based on each incident's location using an interactive online map (PlantmapsGroup, 2024), as shown in Table 1. For further details on the KGCC system, refer to Beck et al. (2018, 2020). The abbreviations for KGCC classifications are listed in Table A1.
3.2 Data preprocessing
The data collected from the three sources, as described in the previous section, were integrated for analysis. The next step involved preprocessing the raw data to prepare it for use in ML models, thereby ensuring the dataset's suitability for accurate and effective analysis. The overall workflow for data preprocessing and analysis followed in this study is illustrated in Figure 1.
The flowchart presents a multi-stage pipeline for predicting degree of injury, organized into four sections labeled along the left: “Phrase Extracted by N L P”, “Data Fusion from Multiple Sources”, “Predicting Degree of Injury”, and “Model Outcome”. At the top, under “Phrase Extracted by N L P”, a box labeled “Accident Report Unstructured Text” feeds into a process labeled “R o B E R T a”. This produces a box labeled “Contextualized Key Phrases”, which includes three highlighted categories: “Activity-related”, “Worker-related”, and “Employer-related”. A side arrow labeled “Validating Phrases” loops downward to connect with the dataset. In the next section, “Data Fusion from Multiple Sources”, a large rounded box labeled “Developed Dataset (72 features)” contains multiple contributing components: “Weather: KGCC (12), Temperature anomaly”; “Victim: Injury Age”; “Project: Type (5), End-use (14), Cost (7)”; and “Extracted Phrases (30)”. These components are combined using plus symbols, indicating feature aggregation. Below this, a downward arrow labeled “Data Imputing and One-Hot Encoding” leads below to the M L Models into the third section. In the “Predicting Degree of Injury” section, a box labeled “M L Models” lists four algorithms: “Random Forest”, “X G Boost”, “K N N”, and “Logistic Regression”. An arrow labeled “Hyperparameter optimization” connects this to a box labeled “Model Performance”, which includes evaluation metrics: “Accuracy”, “Precision”, “Recall”, “F 1 Score”, and “Confusion Matrix”. Finally, in the “Model Outcome” section, results flow downward to a box labeled “S H A P summary plot, S H A P dependency plot”, which then connects to “Practical Implication”, indicating interpretability and real-world application of the model outputs.Research flowchart. Source: Authors’ own work
The flowchart presents a multi-stage pipeline for predicting degree of injury, organized into four sections labeled along the left: “Phrase Extracted by N L P”, “Data Fusion from Multiple Sources”, “Predicting Degree of Injury”, and “Model Outcome”. At the top, under “Phrase Extracted by N L P”, a box labeled “Accident Report Unstructured Text” feeds into a process labeled “R o B E R T a”. This produces a box labeled “Contextualized Key Phrases”, which includes three highlighted categories: “Activity-related”, “Worker-related”, and “Employer-related”. A side arrow labeled “Validating Phrases” loops downward to connect with the dataset. In the next section, “Data Fusion from Multiple Sources”, a large rounded box labeled “Developed Dataset (72 features)” contains multiple contributing components: “Weather: KGCC (12), Temperature anomaly”; “Victim: Injury Age”; “Project: Type (5), End-use (14), Cost (7)”; and “Extracted Phrases (30)”. These components are combined using plus symbols, indicating feature aggregation. Below this, a downward arrow labeled “Data Imputing and One-Hot Encoding” leads below to the M L Models into the third section. In the “Predicting Degree of Injury” section, a box labeled “M L Models” lists four algorithms: “Random Forest”, “X G Boost”, “K N N”, and “Logistic Regression”. An arrow labeled “Hyperparameter optimization” connects this to a box labeled “Model Performance”, which includes evaluation metrics: “Accuracy”, “Precision”, “Recall”, “F 1 Score”, and “Confusion Matrix”. Finally, in the “Model Outcome” section, results flow downward to a box labeled “S H A P summary plot, S H A P dependency plot”, which then connects to “Practical Implication”, indicating interpretability and real-world application of the model outputs.Research flowchart. Source: Authors’ own work
Unlike the categorized features presented in Table 1, the main content of OSHA reports – the incident descriptions – is presented as unstructured text. To extract potential precursors to HRIs from this raw text data, this research employs an NLP approach. NLP encompasses a set of computational techniques that enable machines to analyze and interpret human language, with a particular focus on information retrieval and extraction (Chowdhary and Chowdhary, 2020). While information can also be extracted using alternative methods – such as formulating search functions in Microsoft Excel (Jung et al., 2022) or manually reviewing incident descriptions (Özkan and Ulaş, 2024) – these approaches are labor-intensive (Shamshiri et al., 2024) and may restrict analytical depth.
In contrast, applying text mining and NLP tools facilitates the systematic extraction of meaningful and valid information, reducing the risk of human error. This approach has been increasingly applied in safety research to identify incident precursors and analyze cause–effect relationships (Sahoo et al., 2024). Once relevant phrases are extracted, they are converted into categorized features and integrated with the structured dataset for further analysis. Combining unstructured textual data with structured variables represents a relatively novel approach in the analysis of occupational incident reports (Shamshiri et al., 2024).
One of the earliest efforts to automate content analysis involved manually coding keywords from OSHA reports and analyzing their frequency using NVIVO software to identify key phrases (Esmaeili et al., 2015). This supervised approach has since been enhanced through the application of NLP-based supervised ML algorithms (Tixier et al., 2016). More recently, advancements in NLP have shifted toward unsupervised learning techniques, which eliminate the need for extensive pre-labeled training data (Ma and Chen, 2024). These unsupervised algorithms transform textual data into numerical vectors that represent the semantic content of the original reports, enabling the contextual extraction of key phrases using pretrained transformer models (Grootendorst, 2024; Gupta et al., 2022). Such methods have been effectively used to categorize the causes of construction accidents and link them to high-risk construction activities (Khan et al., 2024).
To achieve the research objectives, NLP packages in Python were employed to extract key phrases from the unstructured incident descriptions. The phrase extraction process followed in this research is illustrated in Figure 2. The Python NLTK library was utilized for text preprocessing, including stopword removal and tokenization of the cleaned reports. Stopwords – commonly occurring non-informative words such as “and,” “the” and “is” – were excluded to improve the relevance of extracted phrases. Unlike fields such as healthcare and law, the construction domain lacks a fine-tuned, domain-specific corpus (Shamshiri et al., 2024). To address this limitation, additional non-informative nouns frequently appearing in HRI reports – such as “employee”, “employer”, “job”, “temperature” and “person” – were added to the stopword list since these terms do not contribute meaningful insights into the precursors of HRIs. The complete stopword list and the NLTK preprocessing code are provided in Figure A1 in Appendix.
The workflow diagram illustrates a step-by-step natural language processing pipeline that transforms an input accident report into contextualized key phrases. At the top left, a box labeled “Input Report” contains a paragraph describing an incident: an employee assisting with roofing activity during extreme heat, feeling unwell, and being hospitalized for heat exhaustion. An arrow labeled “Remove Stopwords” leads to a simplified “Cleaned Report” box on the right, which contains condensed text such as “assisting roofing activity performed complained feeling well hospitalized exhaustion”. From the cleaned report, a vertical arrow labeled “Tokenizing” leads below to a “Tokenized Words” box listing individual terms like “assisting”, “roofing”, “activity”, “performed”, “complained”, “feeling”, “well”, “hospitalized”, and “exhaustion”. A left pointing arrow from “Tokenized Words” leads to a large section titled “Embedded Words (with R o B E R T a)”, showing three groups of vector representations: “Unigram Vectors”, with ovals V 1 assisting, V 2 roofing, ellipsis, and V 9 exhaustion; “Bigram Vectors”, with ovals V 10 assisting roofing, V 11 roofing activity, ellipsis, and V 17 hospitalized exhaustion; and “Trigram Vectors”, with ovals V 18 assisting roofing activity, V 19 roofing activity performed, ellipsis, and V 24 well hospitalized exhaustion. Each vector is represented as an oval with labels. These vectors feed into a “Calculating Cosine Similarity” section at the bottom left, where arrows radiate from a central point labeled “V 0 (Cleaned Report)”. Each arrow corresponds to a vector (V 1, V 2, V 9, V 10, V 11, V 17, V 18, V 19, V 24), and the angle between vectors represents similarity using cosine theta. To the right, a “Ranking and Selecting Top Vectors” panel lists selected vectors such as V 24, V 10, V 11, and V 18 with a green check mark, while less relevant vectors (V 2, V 9, V 1) are crossed out in red. Finally, the selected vectors lead to a box labeled “Contextualized Key Phrases”, which outputs a summarized phrase such as “A 8: Tasks being conducted on roofs, an ellipsis”, representing the extracted contextual meaning from the original report.Phrase extraction process. Source: Authors’ own work
The workflow diagram illustrates a step-by-step natural language processing pipeline that transforms an input accident report into contextualized key phrases. At the top left, a box labeled “Input Report” contains a paragraph describing an incident: an employee assisting with roofing activity during extreme heat, feeling unwell, and being hospitalized for heat exhaustion. An arrow labeled “Remove Stopwords” leads to a simplified “Cleaned Report” box on the right, which contains condensed text such as “assisting roofing activity performed complained feeling well hospitalized exhaustion”. From the cleaned report, a vertical arrow labeled “Tokenizing” leads below to a “Tokenized Words” box listing individual terms like “assisting”, “roofing”, “activity”, “performed”, “complained”, “feeling”, “well”, “hospitalized”, and “exhaustion”. A left pointing arrow from “Tokenized Words” leads to a large section titled “Embedded Words (with R o B E R T a)”, showing three groups of vector representations: “Unigram Vectors”, with ovals V 1 assisting, V 2 roofing, ellipsis, and V 9 exhaustion; “Bigram Vectors”, with ovals V 10 assisting roofing, V 11 roofing activity, ellipsis, and V 17 hospitalized exhaustion; and “Trigram Vectors”, with ovals V 18 assisting roofing activity, V 19 roofing activity performed, ellipsis, and V 24 well hospitalized exhaustion. Each vector is represented as an oval with labels. These vectors feed into a “Calculating Cosine Similarity” section at the bottom left, where arrows radiate from a central point labeled “V 0 (Cleaned Report)”. Each arrow corresponds to a vector (V 1, V 2, V 9, V 10, V 11, V 17, V 18, V 19, V 24), and the angle between vectors represents similarity using cosine theta. To the right, a “Ranking and Selecting Top Vectors” panel lists selected vectors such as V 24, V 10, V 11, and V 18 with a green check mark, while less relevant vectors (V 2, V 9, V 1) are crossed out in red. Finally, the selected vectors lead to a box labeled “Contextualized Key Phrases”, which outputs a summarized phrase such as “A 8: Tasks being conducted on roofs, an ellipsis”, representing the extracted contextual meaning from the original report.Phrase extraction process. Source: Authors’ own work
The tokenized words were then embedded as unigram, bigram and trigram sets using the KeyBERT library to facilitate phrase extraction and further analysis. These n-gram embeddings represent the overall content of each incident report and are ranked using cosine similarity scores based on a pretrained transformer model. Embedded vectors with higher cosine similarity to the cleaned report were given higher rankings. The RoBERTa pretrained model was employed for this ranking process. This model has been extensively validated and demonstrated strong performance in multiple NLP task (Liu et al., 2019).
Unlike previous studies that sought to modify existing pretrained models (Fang et al., 2020) or develop new terminology classification frameworks (Goh and Ubeynarayana, 2017), this study employed RoBERTa purely as a phrase extraction tool. The objective was to summarize the reports and extract semantically meaningful phrases without further categorizing or classifying the document. This approach does not require additional validation procedures that are typically associated with topic modeling or classification and enables managers and engineers to gain insights from a large corpus of raw text without the need for fully reviewing the reports (Lin et al., 2020).
The selected word embeddings were subsequently contextualized into representative key phrases that describe features relevant to each incident report (Luo et al., 2023). To ensure content fidelity, the extracted phrases from the first 19 reports (representing 5% of the dataset) were manually cross-checked against their corresponding unstructured text. This step confirmed that the extracted phrases accurately reflected the key information in the original reports, with no significant omissions. A similar manual analysis was conducted by Liu et al. (2023), in which they reviewed extracted termed (of their pretrained model) with the explicit aim of identifying whether the domain-specific phrases composed of common words. In both cases, the manual contextualization served to ensure interpretability and domain relevance, without requiring additional validation steps such as intercoder agreement assessments. This should be noted that the manual validation in the current study was conducted by a single researcher and focused on content fidelity rather than annotation consistency. Moreover, formal inter-coder reliability metrics were not assessed, which is acknowledged as a limitation of validating the contextualization.
The contextualized extracted phrases are presented in Figure 3. To facilitate data handling and modelling, these phrases were arbitrarily encoded. This encoding scheme does not imply any inherent ranking, priority or hierarchical relationship among the extracted phrases.
The illustration titled “Conceptualized Key Phrases” presents three grouped sections of labeled precursor categories used for analysis: activity-related (A), worker-related (W), and employer-related (E). Each section contains multiple coded items with detailed descriptions. Activity-related precursors (A): It includes twelve items as follows: “A 1: carpentry tasks, including cutting woods for concrete frames, preparing lumbers to be cut, etcetera” “A 2: concreting activities, including pouring concrete, demolishing concrete, preparing the formworks, and finishing concrete” “A 3: equipment operating, including trucks, forklifts, graders, etcetera” “A 4: handling construction tools or materials, such as shovels, jackhammers, hammers, moving sandbags, cement pockets, etcetera” “A 5: Mechanical, Electrical and Plumbing tasks, including installing pipes in trenches, digging trenches, building maintenance activities, installing cables, etcetera” “A 6: masonry tasks, including bricklaying, making mortar, pouring mortar, etcetera” “A 7: tasks related to road construction, including pouring tar and asphalt, shoveling asphalt, finishing asphalt, preparing the road and sidewalks, etcetera” “A 8: tasks being conducted on roofs of structures and buildings, including general construction of roofs, installing gutters, cleaning the garbage on the roof, isolating the roof with tar, etcetera” “A 9: activities related to steel structures or reinforcing for concrete structures, including welding steel frames, joining, installing, and adjusting rebar cages, etcetera” “A 10: activities performed in outdoor areas with direct exposure to sun” “A 11: activities performed in confined spaces, like inside tanks” “A 12: tasks related to surveying”. Worker-related precursors (W): It includes ten items as follows: “W 1: worker denied their illness and or refused medical treatment” “W 2: drug or alcohol intoxication” “W 3: worker had no resting during their shift” “W 4: worker was not acclimatized to working in hot weather conditions; they were either new workers or returning to work after a long period” “W 5: worker did not drink enough water” “W 6: worker had preexisting medical conditions” “W 7: worker was temporarily hired” “W 8: worker was a trainee or apprentice” “W 9: worker was unlicensed for their assigned tasks” “W 10: worker was wearing multiple layers of clothes”. Employer-related precursors (E): It includes eight items as follows: “E 1: employer failed to provide protection from construction hazards (other than exposure to heat)” “E 2: employer did not have or developed a heat injury prevention program” “E 3: employer did not protect their workers from risks of inhaling hazardous chemicals” “E 4: employer did not provide any first aid supplies” “E 5: employer failed to provide sufficient air conditioning” “E 6: employer or their representatives did not inspect the workers frequently” “E 7: employer did not have any heat-related trained personnel” “E 8: employer asked their worker to perform a voluntary task with no written or formal agreement”.Conceptualized key phrases. Source: Authors’ own work
The illustration titled “Conceptualized Key Phrases” presents three grouped sections of labeled precursor categories used for analysis: activity-related (A), worker-related (W), and employer-related (E). Each section contains multiple coded items with detailed descriptions. Activity-related precursors (A): It includes twelve items as follows: “A 1: carpentry tasks, including cutting woods for concrete frames, preparing lumbers to be cut, etcetera” “A 2: concreting activities, including pouring concrete, demolishing concrete, preparing the formworks, and finishing concrete” “A 3: equipment operating, including trucks, forklifts, graders, etcetera” “A 4: handling construction tools or materials, such as shovels, jackhammers, hammers, moving sandbags, cement pockets, etcetera” “A 5: Mechanical, Electrical and Plumbing tasks, including installing pipes in trenches, digging trenches, building maintenance activities, installing cables, etcetera” “A 6: masonry tasks, including bricklaying, making mortar, pouring mortar, etcetera” “A 7: tasks related to road construction, including pouring tar and asphalt, shoveling asphalt, finishing asphalt, preparing the road and sidewalks, etcetera” “A 8: tasks being conducted on roofs of structures and buildings, including general construction of roofs, installing gutters, cleaning the garbage on the roof, isolating the roof with tar, etcetera” “A 9: activities related to steel structures or reinforcing for concrete structures, including welding steel frames, joining, installing, and adjusting rebar cages, etcetera” “A 10: activities performed in outdoor areas with direct exposure to sun” “A 11: activities performed in confined spaces, like inside tanks” “A 12: tasks related to surveying”. Worker-related precursors (W): It includes ten items as follows: “W 1: worker denied their illness and or refused medical treatment” “W 2: drug or alcohol intoxication” “W 3: worker had no resting during their shift” “W 4: worker was not acclimatized to working in hot weather conditions; they were either new workers or returning to work after a long period” “W 5: worker did not drink enough water” “W 6: worker had preexisting medical conditions” “W 7: worker was temporarily hired” “W 8: worker was a trainee or apprentice” “W 9: worker was unlicensed for their assigned tasks” “W 10: worker was wearing multiple layers of clothes”. Employer-related precursors (E): It includes eight items as follows: “E 1: employer failed to provide protection from construction hazards (other than exposure to heat)” “E 2: employer did not have or developed a heat injury prevention program” “E 3: employer did not protect their workers from risks of inhaling hazardous chemicals” “E 4: employer did not provide any first aid supplies” “E 5: employer failed to provide sufficient air conditioning” “E 6: employer or their representatives did not inspect the workers frequently” “E 7: employer did not have any heat-related trained personnel” “E 8: employer asked their worker to perform a voluntary task with no written or formal agreement”.Conceptualized key phrases. Source: Authors’ own work
Given that the reports span various geographical locations and time periods, some features in the dataset contain missing values. ML prediction models, however, require complete datasets with no missing entries. Eliminating incidents with missing values would significantly reduce the sample size, potentially compromising the robustness of the analysis. To address this issue, researchers commonly employ systematic data imputation techniques to preserve data integrity (Deb and Liew, 2016). In this study, the IterativeImputer from the Scikit-learn Python library was used to handle missing values. This imputation method is known for its low bias score (Klomp, 2022) and has been effectively applied in previous accident prediction research (Rabus et al., 2022).
The final step in data preprocessing involves encoding categorical features into binary format to make them suitable for use in the prediction models. This transformation allows the effect of each feature to be evaluated independently, rather than assessing the entire category as a single unit. One-hot encoding is a widely adopted method for this transformation (Alkaissy et al., 2023). In this research, the get-dummies function from the Pandas library in Python was used to perform the encoding. As a result, the final dataset includes 72 independent binary features which serve as input variables for predicting the severity level of HRIs, classified as either fatal or non-fatal.
Before proceeding with model development, it is essential to assess correlations among the independent features. High correlations can lead to redundancy and multicollinearity (Luo et al., 2023), which compromises the efficiency of prediction models. Strongly correlated features should ideally be combined or removed. In this study, the Spearman correlation coefficient was employed to evaluate pairwise relationships among features (Figure 4). The results indicated no significant correlations. The highest correlation coefficient was 0.44, between the “manufacturing plant” project end-use feature and “A11”, which remains below the critical threshold.
The triangular heatmap titled “Spearman Correlation Coefficient Matrix”, displays pairwise correlations between numerous variables. The horizontal and vertical axes list the same set of features, including “A 1 underscore 1”; “A 3 underscore 1”; “A 5 underscore 1”; “A 7 underscore 1”; “A 7 underscore 1”; “A 11 underscore 1”; “W 1 underscore 1”; “W 4 underscore 1”; “W 6 underscore 1”; “W 9 underscore 1”; “W 12 underscore 1”; “W 12 underscore 1”; “E 1 underscore 1”; “E 3 underscore 1”; “E 5 underscore 1”; “E 7 underscore 1”; “Older than 35 underscore 1”; “K G C C underscore A f”; “K G C C underscore A w”; “K G C C underscore B S k”; “K G C C underscore B W k”; “K G C C underscore C s a”; “K G C C underscore D f a”; “project end use underscore Bridge”; “project end use underscore Highway, road, street”; “project end use underscore Multi-family dwelling”; “project end use underscore Other heavy construction”; “project end use underscore Powerline, transmission line”; “project end use underscore Refinery”; “project end use underscore Single family or duplex dwelling”; “project cost underscore 1 million to 5 million”; “project cost underscore 250 thousand to 500 thousand”; “project cost underscore 50 thousand to 250 thousand”; “project cost underscore Under 50 thousand dollars”; “project type underscore Demolition”; and “project type underscore New project or new addition”. Only the lower left triangular portion of the matrix is shown, with the upper right triangle omitted to avoid redundancy. Each cell represents the Spearman correlation coefficient between a pair of variables, encoded by color. A vertical color bar on the right ranges from negative 1.00 to 1.00 with an interval of 0.25, where deep blue indicates strong negative correlation, white or light gray indicates near zero correlation, and deep red indicates strong positive correlation. Most cells appear in light gray or faint shades, suggesting weak or negligible correlations across the majority of feature pairs. Scattered light orange or light blue cells indicate mild positive or negative relationships, respectively, but there are no large contiguous regions of strong correlation. The diagonal is not emphasized, as it would represent self-correlation.Spearman correlation coefficient matrix. Source: Authors’ own work
The triangular heatmap titled “Spearman Correlation Coefficient Matrix”, displays pairwise correlations between numerous variables. The horizontal and vertical axes list the same set of features, including “A 1 underscore 1”; “A 3 underscore 1”; “A 5 underscore 1”; “A 7 underscore 1”; “A 7 underscore 1”; “A 11 underscore 1”; “W 1 underscore 1”; “W 4 underscore 1”; “W 6 underscore 1”; “W 9 underscore 1”; “W 12 underscore 1”; “W 12 underscore 1”; “E 1 underscore 1”; “E 3 underscore 1”; “E 5 underscore 1”; “E 7 underscore 1”; “Older than 35 underscore 1”; “K G C C underscore A f”; “K G C C underscore A w”; “K G C C underscore B S k”; “K G C C underscore B W k”; “K G C C underscore C s a”; “K G C C underscore D f a”; “project end use underscore Bridge”; “project end use underscore Highway, road, street”; “project end use underscore Multi-family dwelling”; “project end use underscore Other heavy construction”; “project end use underscore Powerline, transmission line”; “project end use underscore Refinery”; “project end use underscore Single family or duplex dwelling”; “project cost underscore 1 million to 5 million”; “project cost underscore 250 thousand to 500 thousand”; “project cost underscore 50 thousand to 250 thousand”; “project cost underscore Under 50 thousand dollars”; “project type underscore Demolition”; and “project type underscore New project or new addition”. Only the lower left triangular portion of the matrix is shown, with the upper right triangle omitted to avoid redundancy. Each cell represents the Spearman correlation coefficient between a pair of variables, encoded by color. A vertical color bar on the right ranges from negative 1.00 to 1.00 with an interval of 0.25, where deep blue indicates strong negative correlation, white or light gray indicates near zero correlation, and deep red indicates strong positive correlation. Most cells appear in light gray or faint shades, suggesting weak or negligible correlations across the majority of feature pairs. Scattered light orange or light blue cells indicate mild positive or negative relationships, respectively, but there are no large contiguous regions of strong correlation. The diagonal is not emphasized, as it would represent self-correlation.Spearman correlation coefficient matrix. Source: Authors’ own work
3.3 Predictive modeling of HRI severity
Following the completion of data preprocessing, ML prediction models were developed. This study employed four classification models: random forest (RF), extreme gradient boosting (XGB), K-nearest neighbor (KNN) and logistic regression (LR). These models are frequently applied in construction safety management research (Zhou et al., 2023). For model evaluation and performance assessment, the dataset was partitioned into 80% for training and 20% for testing and validation.
Each model includes multiple hyperparameters that influence predictive performance. Therefore, hyperparameter optimization is essential to ensure optimal model performance. In this study, the grid-search module was used to perform hyperparameter optimization. This approach has been shown to effectively enhance the performance of accident risk prediction models, particularly when applied to sparse and heterogeneous datasets (Moosavi et al., 2019).
Upon optimizing the hyperparameters of the classification models, further improvements in model performance can be achieved by addressing the class imbalance in the target variable. One common approach for mitigating class imbalance is generating synthetic samples to augment the minority class. However, this approach can potentially distort the original data distribution or lead to model overfitting, as it introduces artificial patterns into the training set (Alkhawaldeh et al., 2023). Therefore, such methods should be applied cautiously and are generally more appropriate when the dataset exhibits a substantial degree of class imbalance.
There is not a universal threshold for defining a dataset as imbalanced. Nevertheless, prior studies have suggested that oversampling is typically considered when the number of variables significantly exceeds the number of samples (Blagus and Lusa, 2013), or when the ratio of the majority to minority class instances exceeds 1.5 (Salarian et al., 2023). In the current study, the dataset consists of 370 samples and relatively lower number of variables (70 variables excluding the target variable). Moreover, the target variable has two classes: fatal and non-fatal HRIs with 197 and 173 instances, respectively. This gives the majority to minority class instances ratio of 1.14 which is less than 1.5. Thus, based on these characteristics, the dataset does not exhibit a level of class imbalance that would necessitate the application of synthetic oversampling techniques.
The following subsections provide a brief overview of the prediction models used in this research.
3.3.1 Random forest (RF)
RF combines multiple decision trees to generate predictions through a four-step process: splitting the data into subsets, constructing decision trees in parallel, generating predictions from each tree and calculating the mean of these predictions to produce the final output (Mesghali et al., 2024). One of the key strengths of RF is its capacity to handle multi-dimensional data effectively. Its hyperparameters include the sampling method (with or without replacement, defined by bootstrap), the maximum depth of each tree (max_depth), the minimum number of samples required at a leaf node (min_samples_leaf), the minimum number of samples required to split an internal node (min_samples_split) and the number of trees in the ensemble (n_estimators). When trained on high-quality datasets, RF has demonstrated strong performance, achieving accuracies of up to 97% in predicting fatal fall-from-height accidents (Zermane et al., 2023).
3.3.2 Extreme gradient boosting (XGB)
XGB is a tree-based algorithm known for its robustness against multicollinearity. Its hyperparameters are the maximum depth of each tree (max_depth), the subsample ratio of columns used per tree (colsample_bytree), the minimum loss reduction required to make a further partition on a leaf node (gamma), the learning rate or step size shrinkage to prevent data overfitting (learning_rate), the number of boosting rounds (n_estimators) and the subsample ratio of training instances (subsample). XGB has demonstrated exceptional prediction performance, achieving accuracy rates as high as 99% in road accidents detection and prediction tasks (Parsa et al., 2020).
3.3.3 K-nearest neighbor (KNN)
KNN is a simple yet effective classifier that assigns a class to a data point based on the majority class among its k nearest neighbors (Özkan and Ulaş, 2024). Despite its simplicity, the model's reliance on extensive memory and computational resources can limit its scalability and applicability to large datasets. Key hyperparameters include the number of nearest neighbors considered (n_neighbors), the metric used to calculate proximity between data points (metric) and the method for weighting neighbors based on their distance (weights) (Yang et al., 2022). This model has been applied in occupational safety research, such as identifying and predicting musculoskeletal disorders (Sánchez et al., 2016).
3.3.4 Logistic regression (LR)
LR is a linear classification algorithm for controlling confounding variables in predictive modeling (Kwak and Kho, 2016). It is widely used to establish multivariate regression relationships between a target variable and a set of independent features and is well-suited for binary classification tasks (Zhu et al., 2021). The hyperparameters of this model are the inverse regularization strength (C), the type of regularization applied (penalty), the method for handling multi-class classification (multi_class) and the optimization algorithm used during training (solver). In a relevant application, Gholizadeh and Esmaeili (2020) employed LR to analyze OSHA fatal accident data for electrical contractors and found that the highest fatality rates occurred in “alteration and rehabilitation” projects.
3.4 Evaluating model performance
Following the development of the ML models, the next step involves evaluating and comparing their performance using classification metrics. Model performance is assessed based on a four-class system compromising true positive (TP), true negative (TN), false positive (FP) and false negative (FN) (Alkaissy et al., 2023). A TP occurs when the model correctly predicts a positive actual class. A TN reflects a correct prediction of a negative actual class. An FP occurs when the model incorrectly predicts a positive class. And an FN occurs when the model incorrectly predicts a negative class. These classifications form the foundation of the confusion matrix, which provides a summary of the model's prediction outcomes by displaying the count for each category. An example of a confusion matrix is illustrated in Figure 5.
The diagram titled “Confusion Matrix” presents a two-by-two grid illustrating the outcomes of a binary classification model. The horizontal axis at the top is labeled “Predicted”, with two columns: “Positive” on the left and “Negative” on the right. The vertical axis on the left is labeled “Actual”, with two rows: “Positive” at the top and “Negative” at the bottom. The top-left cell represents “True Positive (T P)”, shown in green. The top-right cell represents “False Negative (F N)”, shown in red. The bottom-left cell represents “False Positive (F P)”, also shown in red. The bottom-right cell represents “True Negative (T N)”, shown in green.The basis of a confusion matrix with four defined classes. Source: Authors’ own work
The diagram titled “Confusion Matrix” presents a two-by-two grid illustrating the outcomes of a binary classification model. The horizontal axis at the top is labeled “Predicted”, with two columns: “Positive” on the left and “Negative” on the right. The vertical axis on the left is labeled “Actual”, with two rows: “Positive” at the top and “Negative” at the bottom. The top-left cell represents “True Positive (T P)”, shown in green. The top-right cell represents “False Negative (F N)”, shown in red. The bottom-left cell represents “False Positive (F P)”, also shown in red. The bottom-right cell represents “True Negative (T N)”, shown in green.The basis of a confusion matrix with four defined classes. Source: Authors’ own work
In the confusion matrix, each class is associated with its own TP, TN, FP and FN. This matrix serves as the foundation for calculating the evaluation metrics detailed below:
Accuracy: is the ratio of correct predictions to the total predictions.
Precision: this metric helps us to realize the actual correct proportion of positive identifications.
Recall (Sensitivity): is the metric for calculating correct actual positives. It is also commonly referred to as sensitivity (Naidu et al., 2023).
F1 score: this metric represents the harmonic mean of precision and recall that indicates the relative importance of recall compared to precision. It is most commonly used when the class distribution is not well-balanced.
Specificity: is the model's ability to predict a true negative of each category available.
Cohen's Kappa Statistic: it is used in classification models as a measure of agreement between observed and predicted classes in a testing dataset (Delgado and Tibau, 2019). It is defined as below.
In this equation Q is the observed outcome and P is the expected outcome.
Receiver operating characteristic curve (ROC) and the area under curve (AUC): ROC curve is a graphical indicator of a classification model's performance by considering the true positive rate (TPR) and the false positive rate (FPR) which are based on the equations below.
AUC is the area below the AUC curve which presents the performance of the model by summarizing the ROC curve. An AUC value of 1 means that the classification model is capable of distinguishing between all positive and negative classes (Naidu et al., 2023). The classification model, in the worst presentation, has an AUC of 0.5 which means that it is unable to ascertain between negative and positive classes.
The calculation of these metrics provides a baseline for comparing the performance of the ML models. The model demonstrating the highest overall performance will be selected for subsequent data analysis and interpretation.
4. Data analysis results and discussion
4.1 Selecting the best model
The developed dataset serves as the input to the ML predictive models. The performance metrics of these models are summarized in Table 2. Additionally, the optimized hyperparameters for each model are presented in Table A2. Furthermore, Figures A2 and A3 in Appendix display the confusion matrix and ROC curves, respectively, which were used to derive the results in Table 2.
Performances of prediction models
| ML model | Accuracy | Specificity | Kappa | ROC-AUC | HRI level | Precision | Recall (sensitivity) | F1 score |
|---|---|---|---|---|---|---|---|---|
| RF | 0.70 | 0.700 | 0.396 | 0.772 | Non-fatal | 0.62 | 0.70 | 0.66 |
| Fatal | 0.78 | 0.70 | 0.74 | |||||
| average | 0.70 | 0.70 | 0.70 | |||||
| XGB | 0.66 | 0.677 | 0.365 | 0.762 | Non-fatal | 0.57 | 0.70 | 0.63 |
| Fatal | 0.76 | 0.64 | 0.69 | |||||
| average | 0.66 | 0.67 | 0.66 | |||||
| KNN | 0.64 | 0.677 | 0.317 | 0.680 | Non-fatal | 0.54 | 0.63 | 0.58 |
| Fatal | 0.72 | 0.64 | 0.67 | |||||
| average | 0.63 | 0.63 | 0.63 | |||||
| LR | 0.69 | 0.677 | 0.365 | 0.764 | Non-fatal | 0.61 | 0.67 | 0.63 |
| Fatal | 0.76 | 0.70 | 0.73 | |||||
| average | 0.68 | 0.69 | 0.68 |
| ML model | Accuracy | Specificity | Kappa | ROC-AUC | HRI level | Precision | Recall (sensitivity) | F1 score |
|---|---|---|---|---|---|---|---|---|
| RF | 0.70 | 0.700 | 0.396 | 0.772 | Non-fatal | 0.62 | 0.70 | 0.66 |
| Fatal | 0.78 | 0.70 | 0.74 | |||||
| average | 0.70 | 0.70 | 0.70 | |||||
| XGB | 0.66 | 0.677 | 0.365 | 0.762 | Non-fatal | 0.57 | 0.70 | 0.63 |
| Fatal | 0.76 | 0.64 | 0.69 | |||||
| average | 0.66 | 0.67 | 0.66 | |||||
| KNN | 0.64 | 0.677 | 0.317 | 0.680 | Non-fatal | 0.54 | 0.63 | 0.58 |
| Fatal | 0.72 | 0.64 | 0.67 | |||||
| average | 0.63 | 0.63 | 0.63 | |||||
| LR | 0.69 | 0.677 | 0.365 | 0.764 | Non-fatal | 0.61 | 0.67 | 0.63 |
| Fatal | 0.76 | 0.70 | 0.73 | |||||
| average | 0.68 | 0.69 | 0.68 |
As shown in Table 2, the RF model achieved the highest accuracy of 0.70, outperforming the other models in terms of precision, recall and F1 scores for both fatal and non-fatal HRIs. The LR model followed closely with an accuracy of 0.69, while XGB and KNN achieved accuracy scores of 0.66 and 0.64, respectively. Based on these results, RF demonstrated the best overall performance among the evaluated models, making it the most suitable choice for analyzing the interrelationships among HRI precursors within the dataset. RF's capacity to model complex, nonlinear interactions between features further justifies its selection. This finding is consistent with prior research, where RF has frequently been identified as a top-performing algorithm in injury severity modeling (Santos et al., 2022).
The sensitivity (recall) score for “fatal” HRIs is 0.7 in both the RF and LR models – the highest among all prediction models. This means that 70% of fatal cases were correctly identified. This relatively high recall score is critical in safety contexts since false negatives (i.e. failing to detect fatal-risk cases) can lead to irreversible consequences. This recall score supports the use of either RF or LR models as risk-averse screening tools prioritize critical interventions, although there might be some false alarms. Moreover, specificity score, which reflects the ability to correctly identify non-fatal cases, is 0.7 for the RF model, higher than other models. High specificity ensures fewer false positives, which in turn reduces unnecessary interventions – a key consideration when safety resources and budget are limited. The RF model, by balancing both sensitivity and specificity scores, offers a suitable middle ground for real-world deployment.
While the achieved scores across all metrics are generally acceptable, the Kappa statistic, which adjusts for agreement occurring by chance, is relatively low across all four models. The highest Kappa value was observed in the RF model (0.396), which falls within the range typically interpreted as “fair agreement”. This outcome is not unexpected given the small sample size and the moderate class imbalance. Importantly, the Kappa score should not be interpreted in isolation. When considered alongside the RF model's high sensitivity, specificity, F1 score and ROC-AUC, the model demonstrates strong overall discriminatory power, despite the modest agreement beyond chance. The relatively low Kappa values across all models reflect the difficulty of achieving high agreement in real-world HRI data.
Although the RF model outperformed the other models across all evaluation metrics, the observed differences in performance were relatively small. Therefore, to assess whether these differences are statistically significant, Delong's test was applied to compare the ROC-AUC scores between each pair of models. Delong's test is a non-parametric method used to evaluate the significance of the difference based on the ROC-AUC scores. A significance level of 5% (p-value <0.05) was used to determine whether one model's ROC-AUC score is statistically superior to another's (Schütze, 2025).
The test was implemented using the MLstatkit library in Python. The results of this test are summarized in Table 3. According to the findings, the p-values for comparisons involving the KNN model were below the 0.05 threshold when compared with RF and XGB, and marginally above the threshold when compared with LR. It indicates that the KNN's predictive performance is significantly different and inferior relative to other models.
Delong's test result of the four models
| p-value | RF | XGB | LR |
|---|---|---|---|
| KNN | 0.0165 | 0.0451 | 0.0618 |
| LR | 0.8000 | 0.8596 | – |
| XGB | 0.5999 | – | – |
| p-value | RF | XGB | LR |
|---|---|---|---|
| KNN | 0.0165 | 0.0451 | 0.0618 |
| LR | 0.8000 | 0.8596 | – |
| XGB | 0.5999 | – | – |
The p-values of the remaining three models exceed the significance threshold, suggesting no statistically significant difference in their ROC-AUC scores. This implies that, from a statistical standpoint, the models perform similarly in predicting the severity of HRIs. Nonetheless, as the RF model attained the highest ROC-AUC score, it is selected for further data analysis.
To acknowledge the bias in the RF model's performance, a bootstrap resampling approach is proposed. Bootstrapping is a resampling simulation approach by repeatedly and randomly sampling subsets of data from the original dataset in order to estimate and quantify the uncertainty in the model's performance (Pei et al., 2016).
In this study, a bootstrapping procedure with 1,000 iterations was performed to estimate 95% confidence intervals for the RF model's predictive performance. The results of the bootstrapping are summarized in Table 4. According to this table, for most metrics, the distance between the main score and the lower or upper bound of the confidence interval is approximately ±15%, with the exception of slightly wider bounds for precision and F1 score. This range provides a practical estimate of the model's uncertainty and indicates that its predictive performance may vary by up to 15% depending on the sample distribution. Although the lower bounds of the results still reflect acceptable performance of the RF model, the observed fluctuation is relatively high when compared to clinical prediction models, where experts have reached model uncertainties of less than 5% (Noma et al., 2021). The level of uncertainty in this study is likely attributed to the limited sample size and the relatively large number of predictive features which can amplify variance in performance estimates.
Bootstrapped results of the RF model performance metrics
| Performance metric | Lower | Upper | Main score |
|---|---|---|---|
| Accuracy | 0.61 | 0.80 | 0.70 |
| Precision | 0.65 | 0.89 | 0.70 |
| Recall (Sensitivity) | 0.57 | 0.84 | 0.70 |
| F1 score | 0.63 | 0.83 | 0.70 |
| ROC-AUC | 0.67 | 0.87 | 0.77 |
| Performance metric | Lower | Upper | Main score |
|---|---|---|---|
| Accuracy | 0.61 | 0.80 | 0.70 |
| Precision | 0.65 | 0.89 | 0.70 |
| Recall (Sensitivity) | 0.57 | 0.84 | 0.70 |
| F1 score | 0.63 | 0.83 | 0.70 |
| ROC-AUC | 0.67 | 0.87 | 0.77 |
4.1.1 Justifying RF selection
It is always advised to perform prediction based on several models to enhance trust (Saarela and Jauhiainen, 2021). Predictive classification models have two primary objectives: performing well and being suitably interpretable for the practitioners. Although the performance of selected linear and nonlinear models (RF, LR and XGB) are almost the same, feature interrelation and contribution to the outcome is more explainable in non-linear models because logistic regression feature importances are less interpretable (Bondell and Reich, 2008). Moreover, logistic regression (and other linear classification) techniques calculate feature coefficients by considering all features as inputs, while random forest calculate the values for feature importances separately for each feature (Saarela and Jauhiainen, 2021). Therefore, due to the RF model's better interpretability, it is selected for further data analysis.
The next step in this research is to investigate the important features and their interrelations with SHAP analysis techniques. Because of Random Forest's scale-independency and nonlinear approach in capturing feature interactions, it outperforms in developing SHAP analysis models compared to linear models. Kim and Kim (2022) utilized the same approach in predicting heat-related mortalities in municipal regions. They applied SHAP models on their RF prediction model to investigate the contribution and interrelations of demographic and socioeconomic features on the predicted mortality rates. The following section presents the findings of SHAP analyses on the results of the RF model.
4.2 Analyzing features' contribution
As outlined in the previous section, the RF model was selected to analyze the contribution of features to the prediction of HRIs. To assess the relative importance of each feature, a SHAP summary plot was generated using the Python SHAP library. Figure 6 presents the top 20 features with the highest contributions to the model's performance, offering insight into their influence on predictive outcomes.
The horizontal axis is labeled “S H A P value (impact on model output)” and ranges from negative 0.2 to 0.2 with an interval of 0.2. A vertical reference line at zero marks neutral impact. The vertical axis lists features in descending order of importance, including “K G C C underscore C f a”, “E 2 underscore 1”, “K G C C underscore C s a”, “Older than 35 underscore 1”, “project cost underscore Under 50 thousand dollars”, “K G C C underscore B S k”, “project end use underscore Single family or duplex dwelling”, “project cost underscore 500 thousand to 1 million”, “K G C C underscore D f b”, “K G C C underscore D f a”, “project cost underscore 250 thousand to 500 thousand”, “K G C C underscore C s b”, “project end use underscore Multi-family dwelling”, “project end use underscore Powerline, transmission line”, “project type underscore Alteration or rehabilitation”, “A 2 underscore 1”, “project end use underscore Highway, road, street”, “project end use underscore Other building”, “A 4 underscore 1”, and “A 5 underscore 1”. Each feature row contains multiple dots representing individual observations. The horizontal position of each dot corresponds to its S H A P value, showing how much that feature contributed to increasing or decreasing the prediction for that instance. The color of the dots represents the feature value, with a color bar on the right labeled “Feature value” ranging from blue (low) to pink or red (high). For the most influential features at the top, such as “K G C C underscore C f a” and “E 2 underscore 1”, points are spread more widely across the horizontal axis, indicating a stronger impact on predictions. Lower-ranked features toward the bottom show tighter clustering around zero, indicating weaker influence. The distribution of colors across each row shows how high and low feature values correspond to positive or negative contributions. Note: All numerical data values are approximated.SHAP summary plot of the RF model. Source: Authors’ own work
The horizontal axis is labeled “S H A P value (impact on model output)” and ranges from negative 0.2 to 0.2 with an interval of 0.2. A vertical reference line at zero marks neutral impact. The vertical axis lists features in descending order of importance, including “K G C C underscore C f a”, “E 2 underscore 1”, “K G C C underscore C s a”, “Older than 35 underscore 1”, “project cost underscore Under 50 thousand dollars”, “K G C C underscore B S k”, “project end use underscore Single family or duplex dwelling”, “project cost underscore 500 thousand to 1 million”, “K G C C underscore D f b”, “K G C C underscore D f a”, “project cost underscore 250 thousand to 500 thousand”, “K G C C underscore C s b”, “project end use underscore Multi-family dwelling”, “project end use underscore Powerline, transmission line”, “project type underscore Alteration or rehabilitation”, “A 2 underscore 1”, “project end use underscore Highway, road, street”, “project end use underscore Other building”, “A 4 underscore 1”, and “A 5 underscore 1”. Each feature row contains multiple dots representing individual observations. The horizontal position of each dot corresponds to its S H A P value, showing how much that feature contributed to increasing or decreasing the prediction for that instance. The color of the dots represents the feature value, with a color bar on the right labeled “Feature value” ranging from blue (low) to pink or red (high). For the most influential features at the top, such as “K G C C underscore C f a” and “E 2 underscore 1”, points are spread more widely across the horizontal axis, indicating a stronger impact on predictions. Lower-ranked features toward the bottom show tighter clustering around zero, indicating weaker influence. The distribution of colors across each row shows how high and low feature values correspond to positive or negative contributions. Note: All numerical data values are approximated.SHAP summary plot of the RF model. Source: Authors’ own work
The SHAP summary plot ranks features according to their importance, with the most influential features displayed at the top. These features exert a greater impact on the model's predictions compared to those ranked lower. Based on the structure of the dataset, features positioned to the left of the vertical axis (negative SHAP values) are associated with the prediction of non-fatal HRIs, while those on the right side of the vertical axis (positive SHAP values) are linked to the prediction of fatal HRIs. Each point in the plot reflects a binary feature's contribution to a single prediction: red points represent the presence of a feature (value = 1), and blue points indicate its absence (value = 0).
Among the top 20 most important features identified by the model, six belong to different climate zones, underscoring the significant role that regional climate plays in predicting the severity of HRI. Notably, the Cfa (humid subtropical) and Csa (hot-summer Mediterranean) climate zones rank first and third, respectively. This finding aligns with the work of Fatima et al. (2021), who reported that temperature anomalies in humid subtropical regions are associated with an increased risk of occupational incidents. According to Figure 6, the Cfa, Dfb (mild-summer humid continental), Dfa (hot-summer humid continental) climate zones are primarily associated with fatal HRIs, whereas, Csa, BSk (cold semi-arid) and Csb (warm-summer Mediterranean) zones are linked to non-fatal HRIs. Among all features, the Dfb climate zone exhibits the highest positive SHAP values, suggesting that working in this zone presents the greatest risk for fatal HRIs. The presence of these features among the top predictors supports the need for climate zone-specific safety planning as a critical component of regionally adapted and environmentally sustainable construction.
The second most influential feature in this plot is E2, which represents cases where “the employer did not have or develop a heat prevention program”. The employer's responsibility in preventing occupational incidents has long been emphasized in safety regulations and standards (Gubernot et al., 2014). The prominent ranking of this feature underscores the critical role of employer-level interventions in mitigating HRIs – independent of temperature anomalies or potential worker errors during task execution. It suggests that organizational-level interventions are not merely safety measures, but a core component of sustainable workforce management. Embedding comprehensive heat prevention programs can effectively prevent or reduce the likelihood of fatal HRIs, a step toward achieving climate-adaptive construction principles.
The fourth-ranked feature is the age of the victims. The distribution of red points for this feature appears exclusively on the right side of the plot, indicating that being over the age of 35 is strongly associated with fatal HRIs. In contrast, younger workers are more likely to experience non-fatal HRIs when exposed to extreme heat. This finding highlights the need for age-aware work planning in extreme environments. Assigning less thermally intensive tasks to older workers supports inclusive workforce policies. Supporting this observation, Lyu and Song (2024) found that middle-aged and older workers – particularly in the construction, manufacturing and agriculture industries – are at a higher risk of fatal heat-related incidents.
The next top feature is related to project cost. “Projects with a budget lower than $50,000” ranks fifth in the SHAP summary plot and is positively associated with an increased risk of fatal HRIs. Additionally, two other project cost categories – representing budgets between $250,000 and $1 million – also appear among the top 20 features. These budget ranges reflect relatively small-scale projects in comparison to others in the dataset with budgets exceeding $1 million. Previous research has shown that, when construction companies are unable to secure adequate funding, safety management practices are often compromised to reduce costs (Rowlinson and Jia, 2015). This highlights the importance of maintaining safety procedures as a core component of construction projects, regardless of their financial constraints – particularly for companies operating with limited budgets. Roles and policies that incentivize or require minimum heat risk safeguards could help mainstream sustainability across all tiers of the construction sector.
Among the top 20 most important features, five are related to the project's end-use, highlighting the significant role this factor plays in predicting the severity of HRIs. The features “single family or duplex dwelling”, “multifamily dwellings” and “other buildings (not commercial or residential)” suggest that construction workers engaged in building projects are generally vulnerable to HRIs. Notably, they may experience extreme heat stress even while working indoors, depending on the nature of their tasks (Al-Bouwarthan et al., 2019). These findings highlight the need for project managers not to underestimate the risk of HRIs in building projects, as the data clearly demonstrate their continued prevalence. The remaining projects in this category – “powerline, transmission line” and “highway, road, street” – typically involve physically demanding labor conducted under harsh environmental conditions in remote areas. These infrastructure projects are also linear in nature, extending across long distances, which complicates safety management due to the challenge of monitoring workers across expansive and dispersed worksites.
Among the project types, “alteration or rehabilitation (renovation)” appears within the top 20 most important features. According to the SHAP summary plot, this feature does not exhibit a clear pattern in distinguishing between fatal and non-fatal HRIs, as the red and blue points are evenly distributed. Nevertheless, its overall importance – especially when compared to other project types such as new construction or maintenance – emphasizes the need for vigilant heat stress monitoring in renovation projects. It is important to note that retrofitting projects often lack a comprehensive safety management framework and present distinct occupational hazards compared to other project types (Doukari et al., 2024).
In terms of activity-related precursors, three features rank among the top 20: A2 “concreting activities”, A4 “handling heavy construction tools or materials” and A5 “mechanical, electrical, and plumbing tasks”. These tasks are physically demanding, typically require the use of personal protective equipment and are frequently performed in constrained postures, leading to elevated levels of metabolic heat generation (Cheung et al., 2016; Tripathi and Mittal, 2024). Accordingly, safety managers should exercise heightened caution during the execution of these activities to mitigate the risk of HRIs. They can establish risk-aware planning that considers both physical exertion and environmental exposure as part of sustainable workplace management.
4.3 Analyzing feature dependency
To explore the interactions between features, SHAP dependency plots were utilized. In these plots, the x-axis represents the value of a binary feature (0 or 1), and the y-axis reflects its SHAP value. Positive SHAP values indicate the feature's contribution to the prediction of fatal HRIs, whereas negative values reflect its association with non-fatal outcomes. The color of each point in the plot denotes the binary value of a secondary interacting feature: blue indicates a value of 0 and red indicates a value of 1. Selected SHAP dependency plots are illustrated in Figure 7.
The three-panel scatter plots are labeled (a), (b), and (c), each showing S H A P interaction effects between pairs of features. In all plots, the horizontal axis represents a binary feature (values 0 or 1), and the vertical axis represents the corresponding S H A P value. Each point represents an observation, and color indicates the interacting feature value, ranging from blue (low, 0.0) to red (high, 1.0), as shown by a vertical color bar. In plot (a), titled “Impact of E 2 and older than 35 on model output”, the horizontal axis is “E 2 underscore 1”, ranging from 0.0 to 1.0 with an interval of 0.02, and the vertical axis is “S H A P value for E 2 underscore 1”, ranging from negative 0.05 to 0.20 with an interval of 0.05. Points at E 2 underscore 1 equal to 0.0 cluster around negative S H A P values from negative 0.07 to negative 0.03, while points at E 2 underscore 1 equal to 1.0 cluster at higher positive S H A P values from 0.10 to 0.22. Color represents “Older than 35 underscore 1”, with both blue and red points present in each cluster. In plot (b), titled “Impact of alteration project type and project cost 250 thousand to 500 thousand on model output”, the horizontal axis is “project type underscore Alteration or rehabilitation”, ranging from 0.0 to 1.0 with an interval of 0.02, and the vertical axis is “S H A P value for project type underscore Alteration or rehabilitation”, ranging from negative 0.03 to 0.02 with an interval of 0.01. Points at project type underscore Alteration or rehabilitation equal to 0.0 cluster around values from negative 0.01 to 0.02, while points at value 1.0 extend more into negative S H A P values, ranging from negative 0.03 to 0.015. Color represents “project cost underscore 250 thousand to 500 thousand”, with both blue and red points distributed across the clusters. In plot (c), titled “Impact of project cost 250 thousand dollars to 500 thousand dollars and E 2 on model output”, the horizontal axis is “project cost underscore 250 thousand to 500 thousand”, ranging from 0.0 to 1.0 with an interval of 0.02, and the vertical axis is “S H A P value for project cost underscore 250 thousand to 500 thousand”, ranging from negative 0.03 to 0.02 with an interval of 0.01. Points at project cost underscore 250 thousand to 500 thousand equal to 0.0, cluster around small positive S H A P values from 0.00 to 0.02, while points at value 1.0 cluster around lower values ranging from negative 0.036 to 0.01. Color represents “E 2 underscore 1”, with blue and red points indicating variation across both clusters. Note: All numerical data values are approximated.SHAP feature dependency plots. Source: Authors’ own work
The three-panel scatter plots are labeled (a), (b), and (c), each showing S H A P interaction effects between pairs of features. In all plots, the horizontal axis represents a binary feature (values 0 or 1), and the vertical axis represents the corresponding S H A P value. Each point represents an observation, and color indicates the interacting feature value, ranging from blue (low, 0.0) to red (high, 1.0), as shown by a vertical color bar. In plot (a), titled “Impact of E 2 and older than 35 on model output”, the horizontal axis is “E 2 underscore 1”, ranging from 0.0 to 1.0 with an interval of 0.02, and the vertical axis is “S H A P value for E 2 underscore 1”, ranging from negative 0.05 to 0.20 with an interval of 0.05. Points at E 2 underscore 1 equal to 0.0 cluster around negative S H A P values from negative 0.07 to negative 0.03, while points at E 2 underscore 1 equal to 1.0 cluster at higher positive S H A P values from 0.10 to 0.22. Color represents “Older than 35 underscore 1”, with both blue and red points present in each cluster. In plot (b), titled “Impact of alteration project type and project cost 250 thousand to 500 thousand on model output”, the horizontal axis is “project type underscore Alteration or rehabilitation”, ranging from 0.0 to 1.0 with an interval of 0.02, and the vertical axis is “S H A P value for project type underscore Alteration or rehabilitation”, ranging from negative 0.03 to 0.02 with an interval of 0.01. Points at project type underscore Alteration or rehabilitation equal to 0.0 cluster around values from negative 0.01 to 0.02, while points at value 1.0 extend more into negative S H A P values, ranging from negative 0.03 to 0.015. Color represents “project cost underscore 250 thousand to 500 thousand”, with both blue and red points distributed across the clusters. In plot (c), titled “Impact of project cost 250 thousand dollars to 500 thousand dollars and E 2 on model output”, the horizontal axis is “project cost underscore 250 thousand to 500 thousand”, ranging from 0.0 to 1.0 with an interval of 0.02, and the vertical axis is “S H A P value for project cost underscore 250 thousand to 500 thousand”, ranging from negative 0.03 to 0.02 with an interval of 0.01. Points at project cost underscore 250 thousand to 500 thousand equal to 0.0, cluster around small positive S H A P values from 0.00 to 0.02, while points at value 1.0 cluster around lower values ranging from negative 0.036 to 0.01. Color represents “E 2 underscore 1”, with blue and red points indicating variation across both clusters. Note: All numerical data values are approximated.SHAP feature dependency plots. Source: Authors’ own work
Figure 7a presents the interaction between the E2 feature (absence of employer heat program) and workers older than 35. When the E2 feature is active (value = 1), the distribution of red and blue dots remains relatively balanced, suggesting that the absence of a heat prevention program increases the fatal HRI risk regardless of worker's age. This finding underscores the critical importance of implementing robust heat safety management plans as a foundational element of HRI prevention.
Figure 7b illustrates the relationship between “alteration or rehabilitation” projects and project budgets ranging from $250,000 to $500,000. A clear pattern emerges in this plot, with red dots (indicating the presence of the budget feature) predominantly concentrated on the right side and blue dots (absence of the feature) on the left. This distribution suggests that most alteration projects in the dataset fall within this budget range. Such projects are typically characterized by constrained budgets, shorter durations and smaller teams. Employers overseeing these projects may also be managing multiple worksites concurrently with fragmented supervision. In small- and medium-sized construction firms, prioritizing environmental and equipment-related safety protocols remains a persistent challenge (Kima et al., 2024). Combined with the inherent characteristics of alteration projects, such limitations often result in insufficient supervision and inadequate heat safety monitoring. As shown in the SHAP dependency plot, this combination contributes to a heightened risk of HRIs and signals the need for policy adjustments that scale safety requirements based on both budget and project complexity.
Figure 7c examines the characteristics of projects with budgets between $250,000 and $500,000. It demonstrates that small- to mid-scale projects are particularly susceptible to fatal outcomes when employer-level heat programs are absent. This finding suggests that current safety frameworks may be insufficiently adaptive to project scale. Developing tiered safety guidelines based on both budget and activity type could improve real-world implementation and contribute to more equitable, sustainable construction environments.
5. Future directions: research and practice implications
Sustainable construction requires proactive strategies that safeguard worker health, particularly under extreme climate conditions. The findings of this study provide insights for achieving heat-resilient construction environments. Several directions for future research and industry practice are discussed below.
Several directions for the future can be identified. One notable NLP-extracted precursors to HRIs are W1: “Worker denied their illness and/or refused medication”. Previous studies have highlighted that poor individual situational awareness and inadequate behavioral responses are contributing factors to HRIs (Jia et al., 2020; Rowlinson and Jia, 2015), yet these issues are often overlooked in existing occupational safety guidelines. To date, no studies have specifically investigated the behavioral and perceptual causes of such responses among construction workers, despite their implications for long-term workforce resilience. Future researchers could explore this issue to support behaviorally informed training, contributing to a safety culture embedded within sustainable construction frameworks. Enhancing workers' safety perception has the potential to improve their situational awareness and commitment to safety standards (Baah et al., 2023). Scholars could employ wearable biosensors, smart protective equipment (Rashidi et al., 2025) or artificial intelligence techniques to examine the drivers of unsafe behaviors (Purushothaman et al., 2025).
Moreover, while this study focused on extracting precursors to HRIs from unstructured text, the OSHA reports also contain valuable information on HRI symptoms, such as vomiting, shivering or fainting. Although symptom analysis fell outside the scope of this study, future research could extract symptom-related phrases to inform real-time health monitoring systems. Integrating this information into early-warning tools could support proactive heat illness intervention and enhance sustainable workforce management in the construction sector.
The findings of this study hold valuable implications for sustainable construction management. By incorporating NLP-extracted features and critical contributing precursors into risk assessment tools, safety managers can develop targeted, context-specific heat prevention programs that account for project type, budget and worker demographics. These targeted interventions go beyond compliance, enabling organizations to embed heat resilience into the planning and execution of construction projects – a localized sustainable development endeavor having a positive impact on workers' health and well-being (Salama et al., 2024).
Notably, the study highlighted that the absence of employer-initiated heat prevention programs is a critical factor contributing to the severity of HRIs. To support sustainable construction practices, regulatory bodies should consider enhancing existing codes and compliance requirements – particularly for small-scale projects, where the safety protocols are often underprioritized due to limited resources. Strengthened regulations would encourage employers to prioritize the prevention of excessive heat stress, ensuring long-term workforce resilience.
6. Conclusion
The construction industry encompasses a wide range of indoor and outdoor activities with varying project budgets and environmental conditions, making the prevention and management of HRIs increasingly complex. This study sought to uncover critical contributing features to HRI severity by integrating unstructured incident reports with other structured climatic and non-climatic data, offering a novel approach to analyzing risk across diverse project and workforce contexts. The study identified key risk factors, such as the absence of employers' heat prevention programs and the workers' age, as strong predictors of severe HRI outcomes in the developed prediction models. These findings underscore the need to embed heat resilience strategies into the planning and execution of construction projects.
This research had some limitations. First, the number of extracted incidents was relatively limited, with only 370 cases obtained from the OSHA database. A larger dataset would have offered a more robust foundation for analyzing feature dependencies and may have improved the predictive performance of the models. Moreover, the OSHA-documented incidents do not have any recorded biological features of the victims (heartbeat rate, blood pressure, etc.), their level of experience or their ethnicity. Integrating other HRI reports into this dataset can enhance the diversity and representativeness of the collected data. Some potential databases to diversify the collected data from HRI victims and different climatic regions are: the “Census of Fatal Occupational Injuries” database maintained by the US Bureau of Labor Statistics, and the “Workers' Compensation” database provided by Safe Work Australia, which documents occupational incidents across Australia.
Second, the incidents analyzed in this study were exclusively from the United States, where not all KGCC classes are represented. Expanding the dataset to include incidents from other regions of the world could increase the diversity of climate zones associated with HRIs, thereby offering a more comprehensive understanding of the influence of climate classification on HRI severity. Lastly, the study was limited by the unavailability of reliable sources for obtaining additional climate variables, such as humidity, wind speed and air pollution levels on the dates of the incidents. Incorporating these factors could have significantly enriched the data analysis by enabling a more detailed assessment of the climatic conditions contributing to HRIs.
Ethics approval
For this study, no human participants or animals were involved.
AI
During the presentation of this work, the authors used ChatGPT 4.0 as an AI-editing assistant tool for refining the manuscript, with caution. After using this tool, the authors reviewed and edited the content as needed and take full responsibility for the content of the publication.
Appendix
The snippet shows a Python script for extracting keywords from text data using the Key B E R T library. At the top, required libraries are imported, including: Line 1 : At indentation level 0, import n l t k. Line 2 : At indentation level 0, import pandas as p d. Line 3 : At indentation level 0, from n l t k dot corpus import stopwords. Line 4 : At indentation level 0, n l t k dot download open parenthesis single quote stopwords single quote close parenthesis. Line 5 : At indentation level 0, from keybert import Key B E R T. Line 6 : At indentation level 0, stop underscore words equals stopwords dot words open parenthesis single quote english single quote close parenthesis plus custom stopwords. Line 7 : At indentation level 0, k w model equals Key B E R T open parenthesis single quote roberta minus base single quote close parenthesis. Line 8 : At indentation level 0, texts equals d f open square bracket single quote raw text single quote close square bracket dot tolist open parenthesis close parenthesis. Line 9 : At indentation level 0, keywords equals k w underscore model dot extract underscore keywords open parenthesis texts comma keyphrase underscore ngram underscore range equals open parenthesis 1 comma 3 close parenthesis. Line 10 : At indentation level 0, comma stop underscore words equals stop underscore words comma top underscore n equals 10 close parenthesis. Line 11 : At indentation level 0, keywords underscore only equals open square bracket open square bracket k w for k w comma score in k w list close square bracket for k w underscore list in keywords close square bracket. Line 12 : At indentation level 0, scores underscore only equals open square bracket open square bracket score for k w comma score in k w list close square bracket for k w list in keywords close square bracket. Line 13 : At indentation level 0, keywords underscore only equals p d dot DataFrame open parenthesis keywords underscore only close parenthesis. Line 14 : At indentation level 0, scores underscore only equals p d dot DataFrame open parenthesis scores underscore only close parenthesis. Line 15 : At indentation level 0, keywords separated equals p d dot concat open parenthesis open square bracket keywords underscore only comma scores underscore only close square bracket comma axis equals 1 close parenthesis. Below the code, a long list of default N L T K English stopwords is displayed, including common words such as: N L T K English stopwords equals open square bracket single quote a single quote comma single quote about single quote comma single quote above single quote comma single quote after single quote comma single quote again single quote comma single quote against single quote comma single quote ain single quote comma. single quote all single quote comma single quote am single quote comma single quote an single quote comma single quote and single quote single quote any single quote comma single quote are single quote comma single quote aren single quote comma double quotes aren single quote t double quotes comma single quote as single quote comma single quote at single quote comma. single quote be single quote comma single quote because single quote comma single quote been single quote comma single quote before single quote comma single quote being single quote comma single quote below single quote comma single quote between single quote comma single quote both single quote comma single quote but single quote comma single quote by single quote comma. single quote can single quote comma single quote couldn single quote comma double quotes couldn single quote t double quotes comma single quote d single quote comma single quote did single quote comma single quote didn single quote comma double quotes didn single quote t double quotes comma single quote do single quote comma single quote does single quote comma single quote doesn single quote comma. double quotes doesn single quote t double quotes comma single quote doing single quote comma single quote don single quote comma double quotes don single quote t double quotes comma single quote down single quote comma single quote during single quote comma single quote each single quote comma single quote few single quote comma single quote for single quote comma single quote from single quote comma. single quote further single quote comma single quote had single quote comma single quote hadn single quote comma double quotes hadn single quote t double quotes comma single quote has single quote comma single quote hasn single quote comma double quotes hasn single quote t double quotes comma single quote have single quote comma single quote haven single quote comma. double quotes haven single quote t double quotes comma single quote having single quote comma single quote he single quote comma double quotes he single quote d double quotes comma double quotes he single quote ll double quotes comma single quote her single quote comma single quote here single quote comma single quote hers single quote comma single quote herself single quote comma double quotes he single quote s double quotes comma. single quote him single quote comma single quote himself single quote comma single quote his single quote comma single quote how single quote comma single quote i single quote comma double quotes i single quote d double quotes comma single quote if single quote comma double quotes i single quote ll double quotes comma double quotes i single quote m double quotes comma single quote in single quote comma single quote into single quote comma single quote is single quote comma. single quote isn single quote comma double quotes isn single quote t double quotes comma single quote it single quote comma double quotes it single quote d double quotes comma double quotes it single quote ll double quotes comma double quotes it single quote s double quotes comma single quote its single quote comma single quote itself single quote comma double quotes i single quote ve double quotes comma single quote just single quote comma single quote 1l single quote comma. single quote m single quote comma single quote ma single quote comma single quote me single quote comma single quote mightn single quote comma double quotes mightn single quote t double quotes comma single quote more single quote comma single quote most single quote comma single quote mustn single quote comma double quotes mustn single quote t double quotes comma single quote my single quote comma. single quote myself single quote comma single quote needn single quote comma double quotes needn single quote t double quotes comma single quote no single quote comma single quote nor single quote comma single quote not single quote comma single quote now single quote comma single quote o single quote comma single quote of single quote comma single quote off single quote comma single quote on single quote comma. single quote once single quote comma single quote only single quote comma single quote or single quote comma single quote other single quote comma single quote our single quote comma single quote ours single quote comma single quote ourselves single quote comma single quote out single quote comma single quote over single quote comma single quote own single quote comma. single quote re single quote comma single quote s single quote comma single quote same single quote comma single quote shan single quote comma double quotes shan single quote t double quotes comma single quote she single quote comma double quotes she single quote d double quotes comma double quotes she single quote ll double quotes comma double quotes she single quote s double quotes comma single quote should single quote comma. single quote shouldn single quote comma double quotes shouldn single quote t double quotes comma double quotes should single quote ve double quotes comma double quotes so single quote comma single quote some single quote comma single quote such single quote comma single quote t single quote comma single quote than single quote comma single quote that single quote comma. double quotes that single quote ll double quotes comma single quote the single quote comma single quote their single quote comma single quote theirs single quote comma single quote them single quote comma single quote themselves single quote comma single quote then single quote comma single quote there single quote comma single quote these single quote comma. single quote they single quote comma double quotes they single quote d double quotes comma double quotes they single quote ll double quotes comma double quotes they single quote re double quotes comma double quotes they single quote ve double quotes comma single quote this single quote comma single quote those single quote comma single quote through single quote comma single quote to single quote comma. single quote too single quote comma single quote under single quote comma single quote until single quote comma single quote up single quote comma single quote ve single quote comma single quote very single quote comma single quote was single quote comma single quote wasn single quote comma double quotes wasn single quote t double quotes comma single quote we single quote comma double quotes we single quote d double quotes comma. double quotes we single quote ll double quotes comma double quotes we single quote re double quotes comma single quote were single quote comma single quote weren single quote comma double quotes weren single quote t double quotes comma double quotes we single quote ve double quotes comma single quote what single quote comma single quote when single quote comma single quote where single quote comma. single quote which single quote comma single quote while single quote comma single quote who single quote single quote whom single quote comma single quote why single quote comma single quote will single quote comma single quote with single quote comma single quote won single quote comma double quotes won single quote t double quotes. single quote wouldn single quote comma double quotes wouldn single quote t double quotes comma single quote y single quote comma single quote you single quote comma double quotes you single quote d double quotes comma double quotes you single quote ll double quotes comma single quote your single quote comma double quotes you single quote re double quotes comma single quote yours single quote comma. single quote yourself single quote comma single quote yourselves single quote comma double quotes you single quote ve double quotes close square bracket. At the bottom, another list of custom stopwords is displayed, including words such as: custom underscore stopwords equals open square bracket single quote employee single quote comma single quote employer single quote comma single quote emergency single quote comma single quote employees single quote comma single quote heat single quote comma single quote environment single quote comma single quote job single quote comma single quote temperature single quote comma single quote temperatures single quote comma single quote degree single quote comma single quote degrees single quote comma single quote health single quote comma single quote index single quote comma single quote Fahrenheit single quote comma single quote fahrenheit single quote comma single quote injury single quote comma single quote died single quote comma single quote injuries single quote comma single quote illness single quote comma single quote illnesses single quote comma single quote disease single quote comma single quote diseases single quote comma single quote project single quote comma single quote building single quote comma single quote construction single quote comma single quote exposure single quote comma single quote exposures single quote comma single quote exposed single quote comma single quote exposing single quote comma single quote environmental single quote comma single quote day single quote comma single quote due single quote comma single quote work single quote comma single quote incident single quote comma single quote person single quote comma single quote injured single quote close square bracket.Full list of the stopwords and the NLTK preprocessing code. Source: Authors’ own work
The snippet shows a Python script for extracting keywords from text data using the Key B E R T library. At the top, required libraries are imported, including: Line 1 : At indentation level 0, import n l t k. Line 2 : At indentation level 0, import pandas as p d. Line 3 : At indentation level 0, from n l t k dot corpus import stopwords. Line 4 : At indentation level 0, n l t k dot download open parenthesis single quote stopwords single quote close parenthesis. Line 5 : At indentation level 0, from keybert import Key B E R T. Line 6 : At indentation level 0, stop underscore words equals stopwords dot words open parenthesis single quote english single quote close parenthesis plus custom stopwords. Line 7 : At indentation level 0, k w model equals Key B E R T open parenthesis single quote roberta minus base single quote close parenthesis. Line 8 : At indentation level 0, texts equals d f open square bracket single quote raw text single quote close square bracket dot tolist open parenthesis close parenthesis. Line 9 : At indentation level 0, keywords equals k w underscore model dot extract underscore keywords open parenthesis texts comma keyphrase underscore ngram underscore range equals open parenthesis 1 comma 3 close parenthesis. Line 10 : At indentation level 0, comma stop underscore words equals stop underscore words comma top underscore n equals 10 close parenthesis. Line 11 : At indentation level 0, keywords underscore only equals open square bracket open square bracket k w for k w comma score in k w list close square bracket for k w underscore list in keywords close square bracket. Line 12 : At indentation level 0, scores underscore only equals open square bracket open square bracket score for k w comma score in k w list close square bracket for k w list in keywords close square bracket. Line 13 : At indentation level 0, keywords underscore only equals p d dot DataFrame open parenthesis keywords underscore only close parenthesis. Line 14 : At indentation level 0, scores underscore only equals p d dot DataFrame open parenthesis scores underscore only close parenthesis. Line 15 : At indentation level 0, keywords separated equals p d dot concat open parenthesis open square bracket keywords underscore only comma scores underscore only close square bracket comma axis equals 1 close parenthesis. Below the code, a long list of default N L T K English stopwords is displayed, including common words such as: N L T K English stopwords equals open square bracket single quote a single quote comma single quote about single quote comma single quote above single quote comma single quote after single quote comma single quote again single quote comma single quote against single quote comma single quote ain single quote comma. single quote all single quote comma single quote am single quote comma single quote an single quote comma single quote and single quote single quote any single quote comma single quote are single quote comma single quote aren single quote comma double quotes aren single quote t double quotes comma single quote as single quote comma single quote at single quote comma. single quote be single quote comma single quote because single quote comma single quote been single quote comma single quote before single quote comma single quote being single quote comma single quote below single quote comma single quote between single quote comma single quote both single quote comma single quote but single quote comma single quote by single quote comma. single quote can single quote comma single quote couldn single quote comma double quotes couldn single quote t double quotes comma single quote d single quote comma single quote did single quote comma single quote didn single quote comma double quotes didn single quote t double quotes comma single quote do single quote comma single quote does single quote comma single quote doesn single quote comma. double quotes doesn single quote t double quotes comma single quote doing single quote comma single quote don single quote comma double quotes don single quote t double quotes comma single quote down single quote comma single quote during single quote comma single quote each single quote comma single quote few single quote comma single quote for single quote comma single quote from single quote comma. single quote further single quote comma single quote had single quote comma single quote hadn single quote comma double quotes hadn single quote t double quotes comma single quote has single quote comma single quote hasn single quote comma double quotes hasn single quote t double quotes comma single quote have single quote comma single quote haven single quote comma. double quotes haven single quote t double quotes comma single quote having single quote comma single quote he single quote comma double quotes he single quote d double quotes comma double quotes he single quote ll double quotes comma single quote her single quote comma single quote here single quote comma single quote hers single quote comma single quote herself single quote comma double quotes he single quote s double quotes comma. single quote him single quote comma single quote himself single quote comma single quote his single quote comma single quote how single quote comma single quote i single quote comma double quotes i single quote d double quotes comma single quote if single quote comma double quotes i single quote ll double quotes comma double quotes i single quote m double quotes comma single quote in single quote comma single quote into single quote comma single quote is single quote comma. single quote isn single quote comma double quotes isn single quote t double quotes comma single quote it single quote comma double quotes it single quote d double quotes comma double quotes it single quote ll double quotes comma double quotes it single quote s double quotes comma single quote its single quote comma single quote itself single quote comma double quotes i single quote ve double quotes comma single quote just single quote comma single quote 1l single quote comma. single quote m single quote comma single quote ma single quote comma single quote me single quote comma single quote mightn single quote comma double quotes mightn single quote t double quotes comma single quote more single quote comma single quote most single quote comma single quote mustn single quote comma double quotes mustn single quote t double quotes comma single quote my single quote comma. single quote myself single quote comma single quote needn single quote comma double quotes needn single quote t double quotes comma single quote no single quote comma single quote nor single quote comma single quote not single quote comma single quote now single quote comma single quote o single quote comma single quote of single quote comma single quote off single quote comma single quote on single quote comma. single quote once single quote comma single quote only single quote comma single quote or single quote comma single quote other single quote comma single quote our single quote comma single quote ours single quote comma single quote ourselves single quote comma single quote out single quote comma single quote over single quote comma single quote own single quote comma. single quote re single quote comma single quote s single quote comma single quote same single quote comma single quote shan single quote comma double quotes shan single quote t double quotes comma single quote she single quote comma double quotes she single quote d double quotes comma double quotes she single quote ll double quotes comma double quotes she single quote s double quotes comma single quote should single quote comma. single quote shouldn single quote comma double quotes shouldn single quote t double quotes comma double quotes should single quote ve double quotes comma double quotes so single quote comma single quote some single quote comma single quote such single quote comma single quote t single quote comma single quote than single quote comma single quote that single quote comma. double quotes that single quote ll double quotes comma single quote the single quote comma single quote their single quote comma single quote theirs single quote comma single quote them single quote comma single quote themselves single quote comma single quote then single quote comma single quote there single quote comma single quote these single quote comma. single quote they single quote comma double quotes they single quote d double quotes comma double quotes they single quote ll double quotes comma double quotes they single quote re double quotes comma double quotes they single quote ve double quotes comma single quote this single quote comma single quote those single quote comma single quote through single quote comma single quote to single quote comma. single quote too single quote comma single quote under single quote comma single quote until single quote comma single quote up single quote comma single quote ve single quote comma single quote very single quote comma single quote was single quote comma single quote wasn single quote comma double quotes wasn single quote t double quotes comma single quote we single quote comma double quotes we single quote d double quotes comma. double quotes we single quote ll double quotes comma double quotes we single quote re double quotes comma single quote were single quote comma single quote weren single quote comma double quotes weren single quote t double quotes comma double quotes we single quote ve double quotes comma single quote what single quote comma single quote when single quote comma single quote where single quote comma. single quote which single quote comma single quote while single quote comma single quote who single quote single quote whom single quote comma single quote why single quote comma single quote will single quote comma single quote with single quote comma single quote won single quote comma double quotes won single quote t double quotes. single quote wouldn single quote comma double quotes wouldn single quote t double quotes comma single quote y single quote comma single quote you single quote comma double quotes you single quote d double quotes comma double quotes you single quote ll double quotes comma single quote your single quote comma double quotes you single quote re double quotes comma single quote yours single quote comma. single quote yourself single quote comma single quote yourselves single quote comma double quotes you single quote ve double quotes close square bracket. At the bottom, another list of custom stopwords is displayed, including words such as: custom underscore stopwords equals open square bracket single quote employee single quote comma single quote employer single quote comma single quote emergency single quote comma single quote employees single quote comma single quote heat single quote comma single quote environment single quote comma single quote job single quote comma single quote temperature single quote comma single quote temperatures single quote comma single quote degree single quote comma single quote degrees single quote comma single quote health single quote comma single quote index single quote comma single quote Fahrenheit single quote comma single quote fahrenheit single quote comma single quote injury single quote comma single quote died single quote comma single quote injuries single quote comma single quote illness single quote comma single quote illnesses single quote comma single quote disease single quote comma single quote diseases single quote comma single quote project single quote comma single quote building single quote comma single quote construction single quote comma single quote exposure single quote comma single quote exposures single quote comma single quote exposed single quote comma single quote exposing single quote comma single quote environmental single quote comma single quote day single quote comma single quote due single quote comma single quote work single quote comma single quote incident single quote comma single quote person single quote comma single quote injured single quote close square bracket.Full list of the stopwords and the NLTK preprocessing code. Source: Authors’ own work
The grouped confusion matrix compares four models: “R F”, “X G B”, “K N N”, and “L R”, with values color-coded according to the legend (black for R F, green for X G B, blue for K N N, and red for L R). The matrix is divided into four sections based on actual and predicted classes. The top-left section, labeled “Real Non-fatal predicted as Non-fatal”, represents true negatives. The values are R F 21, X G B 21, K N N 19, and L R 20. The top-right section, labeled “Real Non-fatal predicted as Fatal”, represents false positives. The values are R F 9, X G B 9, K N N 11, and L R 10. The bottom-left section, labeled “Real Fatal predicted as Non-fatal”, represents false negatives. The values are R F 13, X G B 16, K N N 16, and L R 13. The bottom-right section, labeled “Real Fatal predicted as Fatal”, represents true positives. The values are R F 31, X G B 28, K N N 28, and L R 31. The background shading differentiates the four quadrants, with neutral tones for correct classifications and red-tinted areas for misclassifications.Integrated confusion matrix of the prediction models. Source: Authors’ own work
The grouped confusion matrix compares four models: “R F”, “X G B”, “K N N”, and “L R”, with values color-coded according to the legend (black for R F, green for X G B, blue for K N N, and red for L R). The matrix is divided into four sections based on actual and predicted classes. The top-left section, labeled “Real Non-fatal predicted as Non-fatal”, represents true negatives. The values are R F 21, X G B 21, K N N 19, and L R 20. The top-right section, labeled “Real Non-fatal predicted as Fatal”, represents false positives. The values are R F 9, X G B 9, K N N 11, and L R 10. The bottom-left section, labeled “Real Fatal predicted as Non-fatal”, represents false negatives. The values are R F 13, X G B 16, K N N 16, and L R 13. The bottom-right section, labeled “Real Fatal predicted as Fatal”, represents true positives. The values are R F 31, X G B 28, K N N 28, and L R 31. The background shading differentiates the four quadrants, with neutral tones for correct classifications and red-tinted areas for misclassifications.Integrated confusion matrix of the prediction models. Source: Authors’ own work
The chart is titled “ROC Curves – All Models (Fatal H R I s equals Positive Class)”. The horizontal axis is labeled “False Positive Rate”, ranging from 0.0 to 1.0 with an interval of 0.2, and the vertical axis is labeled “True Positive Rate”, also ranging from 0.0 to 1.0 with an interval of 0.2. A diagonal dashed line from the bottom left at (0.0, 0.0) to the top right at (1.0, 1.0) represents a random classifier with an A U C of 0.5. Four colored curves represent different machine learning models. The orange curve for “Random Forest” has an A U C of 0.772 and generally lies above the others across most of the range. The green curve for “X G Boost” has an A U C of 0.762 and follows closely behind Random Forest. The purple curve for “Logistic Regression” has an A U C of 0.766 and performs similarly to the top models, overlapping in several regions. The blue curve for “K N N” has a lower A U C of 0.680 and remains below the other model curves for most values of false positive rate. All curves start near the origin and rise toward the top right corner, with improved true positive rates as false positive rates increase. The Random Forest, Logistic Regression, and X G Boost models show better separation from the diagonal baseline, indicating stronger predictive performance compared to K N N. Note: All numerical data values are approximated.The ROC curves of the classification models. Source: Authors’ own work
The chart is titled “ROC Curves – All Models (Fatal H R I s equals Positive Class)”. The horizontal axis is labeled “False Positive Rate”, ranging from 0.0 to 1.0 with an interval of 0.2, and the vertical axis is labeled “True Positive Rate”, also ranging from 0.0 to 1.0 with an interval of 0.2. A diagonal dashed line from the bottom left at (0.0, 0.0) to the top right at (1.0, 1.0) represents a random classifier with an A U C of 0.5. Four colored curves represent different machine learning models. The orange curve for “Random Forest” has an A U C of 0.772 and generally lies above the others across most of the range. The green curve for “X G Boost” has an A U C of 0.762 and follows closely behind Random Forest. The purple curve for “Logistic Regression” has an A U C of 0.766 and performs similarly to the top models, overlapping in several regions. The blue curve for “K N N” has a lower A U C of 0.680 and remains below the other model curves for most values of false positive rate. All curves start near the origin and rise toward the top right corner, with improved true positive rates as false positive rates increase. The Random Forest, Logistic Regression, and X G Boost models show better separation from the diagonal baseline, indicating stronger predictive performance compared to K N N. Note: All numerical data values are approximated.The ROC curves of the classification models. Source: Authors’ own work
KGCC code description
| Climate zone | KGCC code |
|---|---|
| Tropical Rainforest | Af |
| Tropical Monsoon | Am |
| Tropical Savanna (Wet and Dry Climate) | Aw |
| Hot Semi-Arid Climate | BSh |
| Cold Semi-Arid Climate | BSk |
| Hot Desert Climate | BWh |
| Cold Desert Climate | BWk |
| Humid Subtropical Climate | Cfa |
| Hot-Summer Mediterranean Climate | Csa |
| Warm-Summer Mediterranean Climate | Csb |
| Humid Continental Hot Summers | Dfa |
| Humid Continental Mild Summer, Wet All Year | Dfb |
| Climate zone | KGCC code |
|---|---|
| Tropical Rainforest | Af |
| Tropical Monsoon | Am |
| Tropical Savanna (Wet and Dry Climate) | Aw |
| Hot Semi-Arid Climate | BSh |
| Cold Semi-Arid Climate | BSk |
| Hot Desert Climate | BWh |
| Cold Desert Climate | BWk |
| Humid Subtropical Climate | Cfa |
| Hot-Summer Mediterranean Climate | Csa |
| Warm-Summer Mediterranean Climate | Csb |
| Humid Continental Hot Summers | Dfa |
| Humid Continental Mild Summer, Wet All Year | Dfb |
Performances of the prediction models
| ML model | Hyperparameter | Value |
|---|---|---|
| RF | “bootstrap” | False |
| “max_depth” | None | |
| “min_samples_leaf” | 2 | |
| “min_samples_split” | 2 | |
| “n_estimators” | 100 | |
| XGB | “colsample_bytree” | 0.8 |
| “gamma” | 0 | |
| “learning_rate” | 0.01 | |
| “max_depth” | 10 | |
| “n_estimators” | 300 | |
| “subsample” | 1.0 | |
| KNN | “metric” | “manhattan” |
| “n_neighbors” | 11 | |
| “weights” | “distance” | |
| LR | “C” | 1 |
| “multi_class” | “multinominal” | |
| “penalty” | “l1” | |
| “solver” | “saga” |
| ML model | Hyperparameter | Value |
|---|---|---|
| RF | “bootstrap” | False |
| “max_depth” | None | |
| “min_samples_leaf” | 2 | |
| “min_samples_split” | 2 | |
| “n_estimators” | 100 | |
| XGB | “colsample_bytree” | 0.8 |
| “gamma” | 0 | |
| “learning_rate” | 0.01 | |
| “max_depth” | 10 | |
| “n_estimators” | 300 | |
| “subsample” | 1.0 | |
| KNN | “metric” | “manhattan” |
| “n_neighbors” | 11 | |
| “weights” | “distance” | |
| LR | “C” | 1 |
| “multi_class” | “multinominal” | |
| “penalty” | “l1” | |
| “solver” | “saga” |

