The flowchart presents a multi-stage pipeline for predicting degree of injury, organized into four sections labeled along the left: “Phrase Extracted by N L P”, “Data Fusion from Multiple Sources”, “Predicting Degree of Injury”, and “Model Outcome”. At the top, under “Phrase Extracted by N L P”, a box labeled “Accident Report Unstructured Text” feeds into a process labeled “R o B E R T a”. This produces a box labeled “Contextualized Key Phrases”, which includes three highlighted categories: “Activity-related”, “Worker-related”, and “Employer-related”. A side arrow labeled “Validating Phrases” loops downward to connect with the dataset. In the next section, “Data Fusion from Multiple Sources”, a large rounded box labeled “Developed Dataset (72 features)” contains multiple contributing components: “Weather: KGCC (12), Temperature anomaly”; “Victim: Injury Age”; “Project: Type (5), End-use (14), Cost (7)”; and “Extracted Phrases (30)”. These components are combined using plus symbols, indicating feature aggregation. Below this, a downward arrow labeled “Data Imputing and One-Hot Encoding” leads below to the M L Models into the third section. In the “Predicting Degree of Injury” section, a box labeled “M L Models” lists four algorithms: “Random Forest”, “X G Boost”, “K N N”, and “Logistic Regression”. An arrow labeled “Hyperparameter optimization” connects this to a box labeled “Model Performance”, which includes evaluation metrics: “Accuracy”, “Precision”, “Recall”, “F 1 Score”, and “Confusion Matrix”. Finally, in the “Model Outcome” section, results flow downward to a box labeled “S H A P summary plot, S H A P dependency plot”, which then connects to “Practical Implication”, indicating interpretability and real-world application of the model outputs.Research flowchart. Source: Authors’ own work
Sharing content requires targeting cookies to be enabled. Please update your cookie preferences to use this feature.