Despite the growing integration of machine learning (ML) in humanitarian supply chain (HSC), herein HSCML, no prior review has established a systematic application domain taxonomy or examined model quality assurance, limiting sustainable adoption. This study aims to address three gaps: the absence of a systematic taxonomy for HSCML application domains, the lack of frameworks for HSCML data and model quality and the disconnect between technical development and operational deployability.
A sequential mixed-methods systematic literature review was used. In Phase 1, latent Dirichlet allocation topic modelling was applied to 302 HSCML journal papers to identify thematic clusters and derive a context–intervention–mechanism–outcome (CIMO) classification system. In Phase 2, a CIMO-guided systematic review analysed 93 HSCML model development studies across contextual, methodological and quality dimensions.
ML contributed to HSC through six domains: disaster consequence prediction, relief operations optimisation, real-time crisis intelligence, infrastructure monitoring and recovery, population mobility and evacuation and strategic HSC network planning. Research efforts are unevenly distributed across disaster types, management phases and regions, with fairness, explainability and uncertainty quantification systematically neglected across all domains. In response, three frameworks are established covering data quality dimensions, model quality dimensions and factors influencing model quality.
This is the first systematic examination of data and model quality in HSCML. It delivers a deployability-focused quality assurance framework and a data-centric research agenda addressing 14 knowledge gaps, providing practical references for rigorous HSCML development in line with humanitarian obligations.
1. Introduction
Humanitarian supply chain (HSC) focuses on the preparation for, response to and recovery from humanitarian crises, to save lives and alleviate the suffering of affected populations (Kembro et al., 2024). The timely and effective assistance of HSC has been demonstrated in the preparation for and response to many disasters, benefiting millions of beneficiaries. Recovery and mitigation efforts, such as event and consequence forecasting (Kondraganti et al., 2022), preparedness (Kumar et al., 2022) and post-disaster recovery responses (Yağcıoğlu, 2025) were also a part of humanitarian operations. HSCs are guided primarily by non-profit objectives and must function under extreme uncertainty, resource constraints and time-critical conditions (de Camargo Fiorini et al., 2022; Zhao et al., 2025), requiring collaborative decision-making among stakeholders under different mandates and with diverse logistics capabilities (Altay et al., 2023; Sentia et al., 2023).
In such managerial and operational environments, artificial intelligence (AI) has emerged as a technology of significant potential, capable of delivering insights and recommendations tailored to operational situations. Among AI approaches, machine learning (ML), in which models are trained on historical data to make predictions, classifications or optimisation decisions, can handle complex relationships to generate accurate outputs. As ML enables complex and high-value predictive and prescriptive analytics for decision-making, and given that the LSCM literature is overwhelmingly ML-based (Vlachos and Reddy, 2025), this review focuses specifically on ML in HSC (HSCML).
HSCML has evolved from early applications in disaster and risk management, driven primarily by computer science communities, towards a rapidly expanding and multi-disciplinary field encompassing operational logistics, real-time intelligence and network resilience design (e.g. damage assessment (Chen and Cho, 2022), situational awareness (Ullah et al., 2021; Zhang et al., 2024b), UAV for disaster reconnaissance and communication restoration (Zhu and Wang, 2023; Yağcıoğlu, 2025). Despite its emergence and advocates for adoption, the HSCML’s foundational taxonomy and quality assurance have not been established. As evidenced in Table 1, existing HSC reviews address operational, organisational and sustainability dimensions without ML as a primary focus, and Vlachos and Reddy (2025) covered general ML-focused LSCM rather than humanitarian contexts. Where ML appears in humanitarian reviews, it is either incidental or prescribed as a future research direction, with no prior work establishing an application domain taxonomy or examining data and model quality dimensions.
Meanwhile, HSC stakeholders are struggling with adoption, with major concerns centred on data and model quality. Organisations like the UN Refugee Agency, the Office for the Coordination of Humanitarian Affairs (OCHA) and the International Committee of the Red Cross raised concerns about AI-based solutions producing problematic outputs stemming from data quality issues and algorithmic bias (OCHA, 2024). Empirical evidence confirms these concerns, documenting how opaque AI systems cause distrust and operational harm (Behl et al., 2026), while an insufficient understanding of model quality dimensions and their trade-offs leaves operators ill-equipped to govern and maintain AI accountability in line with humanitarian principles and legal obligations (Dorsey, 2026). Addressing these challenges requires establishing quality assurance frameworks and a structured awareness of the factors that influence data and model quality. Through evidence-based insights into HSCML contexts, applications and outcomes, this study responds to these gaps through three key research questions (RQs):
How was HSCML proposed and implemented across different contexts and application domains?
How have data quality dimensions been recognised and attended to across HSCML application domains?
What dimensions and factors characterise HSCML model quality?
RQ2 and RQ3 target data and model quality as the mechanisms through which the gap between HSCML technical development and operational deployability manifests, directly identifying dimensions and factors whose systematic oversight obstructs adoption. Using latent Dirichlet allocation (LDA) topic modelling on 302 ML-relevant HSC papers and a context–intervention–mechanism–outcome (CIMO)-structured SLR of 93 HSCML model development studies, this review characterises HSCML across six application domains, detailing ML applications, implementations and HSC contributions in each context. This review systematically examines how data quality and model quality dimensions have been addressed and overlooked across domains. In response, three frameworks are established for data quality dimensions, model quality dimensions and factors influencing model quality, along with a deployability-focused, data-centric research agenda that addresses 14 identified knowledge gaps. Together, these contribute to principled, deployable HSCML development that aligns technical rigour with humanitarian obligations.
The rest of the paper is organised as follows. Section 2 details the sequential mixed-methods SLR, including LDA topic modelling results that support the CIMO structure. Section 3 presents the results of the systematic literature review. Section 4 discusses key findings, research gaps and future directions, before the conclusion in Section 5.
2. Review methodology
This study uses a sequential mixed-methods SLR design, integrating quantitative and qualitative strands in sequence (Grant and Booth, 2009; Harden and Thomas, 2010). As no established HSCML application domain taxonomy existed before this study, Phase 1 deploys LDA topic modelling as the quantitative-dominant strand to identify the thematic structure of the HSCML corpus, particularly the application domains and expected outcomes of HSCML, from which an evidence-grounded CIMO classification framework is derived to structure Phase 2’s CIMO-guided systematic review of model development studies, curated through uniformly applied quality and relevance criteria.
2.1 Data identification and collection
Based on the three targeted core elements: humanitarian, supply chain management and ML, a terminological expansion was carried out to develop a comprehensive three-field keyword structure detailed in Figure 1. The ML field keyword use generalised parent-category and technologically related terms, e.g. “deep learning”, “neural network”, “predictive analytics” and “prescriptive analytics” that encompass all specific technique variants and methodological overlaps, avoiding the risk of omitting novel or hybrid methods that technique-specific search terms might overlook. This structure was used to search titles, abstracts and keywords for English-language journal papers in Scopus and Web of Science (updated to 31 December 2025), with no publication date restrictions. This review excludes grey literature (e.g. technical reports, white papers) to ensure peer-review processes, reproducibility and technical specifications.
The papers’ metadata were imported into EndNote. An aggregation with duplication removal yielded 501 papers. Abstracts were examined by two independent reviewers for relevance and originality (e.g. excluding papers addressing ML but not to support HSC, retraction and literature reviews). The data set was finalised at 302 journal articles. The inclusion and exclusion criteria, along with the number of papers after each screening step, are illustrated in Figure 1. To avoid missing relevant papers in emerging or specialised journals, the filtering based on journal ranking is placed only before the SLR phase.
2.2 Topic modelling
LDA was used to identify underlying thematic structures within the 302-paper corpus at the semantic level. LDA represents topics as probability distributions over words, while papers have probability distributions over topics (Zhao et al., 2024b), thus enabling the direct examination of semantic content (i.e. title, keywords and abstract) when forming research clusters. The LDA models are governed by Dirichlet priors and built through an iterative training process. Inference is then conducted for each reference to determine its topic-probability mixture (Figure 1), with a membership threshold of ≥ 0.3 considered as significant topic membership. Each topic is uniquely described in two aspects. Saliency prioritises frequently occurring terms (Chuang et al., 2012), whereas relevance ranks terms by their distinctiveness, emphasising word exclusivity (Sievert and Shirley, 2014).
This study used natural language processing (NLP) packages in Python: NLTK for text preprocessing, including tokenisation, stop-word removal and lemmatisation; gensim for LDA model implementation, topic extraction and n-gram detection; and pyLDAvis for term list extraction. Domain terminology standardisation and abbreviation handling were implemented through a custom thesaurus. In Step 2, different model configurations were tried and their performance compared. Title and keyword weights (1–5) relative to abstracts were applied. Weight combinations and topic numbers ranging from 5 to 20 were grid-searched, with the formation achieving the highest topic coherence score ( measures semantic coherence based on the co-occurrence patterns of top topic terms) selected for in-depth analysis and model parameters fixed (end of Step 3).
Following topic membership inference for each paper (Step 4), reviewers engaged qualitatively with the member papers and the top-ranked saliency and relevance terms for each topic in Step 5 (Figure 1), prioritising studies with higher topic membership probabilities as more exclusive and representative of the topic’s thematic core. For each topic, reviewers identified the distinguishing ML application domain, characterised by its HSC purpose and operational focus, which constitutes the intervention (I) class in the CIMO framework. Concurrently, the improvements and claimed impacts reported in those papers, including the rationale for the ML application, constitute the outcome (O) classes. Where a first-level topic contained a disproportionately large number of studies and exhibited internal thematic heterogeneity upon qualitative review, another iteration of Step 2 – Step 4 was applied to that topic’s corpus alone to identify finer-grained subtopics (i.e. Topics 1, 2, 3, 7 and 8). In total, 14 topics were characterised. Topics whose member studies exhibited a higher probability of belonging to other topics (Topics 6 and 8.2) were not characterised. Across the topics and subtopics, six I classes and five O classes were identified (Figures 2 and 3).
2.3 Systematic review
Metadata and full texts of 302 papers were further filtered based on the journals, following a multi-year combined list of ABS, ABDC and SCImago Q1 and the research focus on ML model development. Studies that solely focused on socio-technical aspects or lacked sufficient development details (e.g. data, model specifications) were excluded (Figure 1). Both criteria are methodologically required to ensure that RQs are answered by analyses on good (i.e. journal filtering) and complete and relevant data (i.e. model development filter). The final set of 93 papers was analysed in the systematic review.
Applied flexibly as Denyer and Tranfield (2009) intended, an adapted CIMO classification framework was used to characterise each HSCML study, i.e. in which context an ML application was deployed (C), what intervention was proposed (I), how it was technically developed through data and algorithms (M) and what outcomes were expected (O) (Figure 3). The C classes are adapted from previous HSC literature reviews (Chiappetta Jabbour et al., 2017; de Camargo Fiorini et al., 2022). The I and O classes are derived from LDA topic characterisation (see Section 2.2). For the M classes, while standard CIMO mechanisms describe human reasoning, motivation or social processes, ML intervention mechanisms are inherently technical. Data determines what patterns can be learnt, while algorithms determine how learning occurs (training) and how inputs are transformed into outputs (inference). Defining M as data and algorithms thus preserves the analytical separation between what is done (I) and how it is done (M), directly enabling the answers to RQ2 and RQ3. Multiple class tags can be used for each paper.
Table 2 summarises the quality assurance practices applied throughout this review. It is worth noting that the analytical scope does not extend to specific developing pipeline practices (e.g. data processing, feature engineering and hyperparameter tuning). These context-, data- and modelling-purpose-specific investigations call for sub-field SLRs that could be facilitated using frameworks aggregated in this paper.
3. Systematic literature review results
Three reviews were conducted, covering multiple aspects of the selected 93 HSCML papers, following the C, I, M and O classes. Figure 4 illustrates the complete analytical flow, from the primary analytical focus of each subsection.
3.1 Context-focused analysis
The heatmaps in Figure 5 examine Context (C) × Intervention (I) × Outcome (O) intersections to address RQ1. The mechanism – M dimensions (data and algorithms) will be analysed in Sections 3.2 and 3.3, as problem requirements primarily drive technical choices.
Regarding country type and urbanisation, HSCML applications in developed countries and LMICs primarily focus on real-time situational awareness and logistics optimisation, based on predictive analytics (e.g. tree-based and deep learning) and prescriptive analytics (e.g. RL) [Figure 5(b)]. RL was used to ensure computational efficiency while maintaining optimality close to that of exhaustive optimisation methods (Pan et al., 2025). HSCML in developed-country contexts tends to lean more towards critical infrastructure management, whereas in LMICs, it is more often focused on consequence prediction. LDCs are underrepresented in HSCML studies (5.7%), which are mostly set in rural contexts. HSCML in the context of LDCs uniquely prioritised disaster impact forecasting [e.g. mapping flood emergency aid requirements (Rahman et al., 2021)], while continuous monitoring capabilities in these countries focused on pressing issues, such as water quality monitoring, war and conflict (Heuschmid et al., 2025) [Figure 5(d)].
Natural disasters with sudden onset is the dominant focus, primarily addressing real-time situational awareness in response [e.g. UAV-based detections (Hazarika et al., 2025)], while natural disasters with slow onset, such as droughts (Araya et al., 2023) or pandemics (Bhullar et al., 2022), demand resource optimisation with a focus on preparation [Figures 5(a) and (e)]. Earthquakes and seismic events are the most common, focusing on building damage assessment and monitoring (Chen and Cho, 2022) and response coordination (Amin Amani et al., 2025), with tsunamis addressed as cascading events requiring early warning, vulnerability assessment and evacuation (Song et al., 2017; van Steenbergen et al., 2023). Flood and storm, as well as wildfire, studies emphasise infrastructure monitoring (e.g. communication and power systems) and social media analytics for targeted humanitarian logistics. By contrast, storm, flood, drought and landslide studies feature risk management with occurrence and consequence prediction (e.g. groundwater salinity prediction).
Human-made disasters with slow onset exclusively target resource optimisation, such as refugee resettlement and long-term aid planning (Ahani et al., 2021). Meanwhile, human-made disasters with sudden onset prioritise situational awareness in war and military contexts, such as satellite imagery in monitoring humanitarian crises (Zhao et al., 2024a), search and rescue in collapsed structures (Hu et al., 2022) and anti-personnel demining (Heuschmid et al., 2025). This is also the context for compound disasters, which necessitate adaptive policies, e.g. armed conflict, infrastructural collapse and service disruptions (Gong et al., 2024; Zhao et al., 2024a). The deployment of AI in conflict settings introduces distinct legal and ethical requirements, including compliance with proportionality obligations under International Humanitarian Law that current HSCML technical literature has yet to address (Dorsey, 2026).
Preparation phase prioritises strategic HSC design on pre-event allocation optimisation, such as inventory management and procurement planning (Quiliche et al., 2023). The response phase focuses on real-time intelligence for rapid damage assessment and search and rescue operation (Hu et al., 2022) and immediate resource distribution and procurement [e.g. multi-agent emergency response (Nadi and Edrisi, 2017)]. Recovery phase demonstrates higher emphasis on resilience and sustainability for long-term reconstruction [e.g. communication and transportation infrastructure monitoring and restoration (Ghannad et al., 2021)] [Figure 5(c)]. The mitigation phase prioritises risk assessment and prevention through predictive likelihood and consequences insights [Figure 5(b)] [e.g. risk mapping (Rahman et al., 2021), continuous infrastructure monitoring and contingency planning (Jana et al., 2023) and essential supply resilience and sustainability (Zhao et al., 2025)] [Figure 5(c)].
Apart from social media analytics, the demand forecasting stage used remote sensing to anticipate event consequences (e.g. building damage assessment for the prioritisation of evacuation (Xia et al., 2023) or relief demand prediction based on dynamic population distribution (Lin et al., 2020)) [Figure 5(a)]. ML in distribution emphasises real-time intelligence on humanitarian demands [e.g. NLP (Kiavash et al., 2026)] to optimise precise and robust last-mile delivery [e.g. RL (van Steenbergen et al., 2023; Peng et al., 2025)]. HSCML of the logistics stage predominantly addresses transportation limitations in disrupted disaster zones [e.g. truck-UAV fleet deployment (van Steenbergen et al., 2023)], rapid resource optimisation decision-making [e.g. RL for optimisation (Yu et al., 2021)] and transportation networks monitoring and recovery (Jana et al., 2023). Procurement remains underexplored, focusing on HSC design [e.g. optimal vendor selection under uncertainties (Hu et al., 2024)]. Due to the lack of actual disaster data and the requirement for data completeness, synthetic data were widely used in these contexts. Apart from ML-integrated information systems for extracting HSC insights (e.g. NLP for social media pipelines), the information flow saw an emerging use case of re-establishing communication infrastructure, such as 5G networks (Yağcıoğlu, 2025). Financial aspects remain significantly underexplored, with a focus on socioeconomic-based resource allocation [e.g. the distribution of wealth and poverty (Chi et al., 2022), humanitarian relief and stabilising food prices through national fiscal policies (Xiong et al., 2024)].
3.2 Domain-focused analysis: applications and data
This section details the six application domains. Together with Section 3.1, this section directly addresses RQ1 by characterising how HSCML has been operationalised and RQ2 by reviewing the data quality dimensions relevant to each domain. The findings from this section also support the model quality dimensions and the influencing factors examined in Section 3.3 in addressing RQ3.
3.2.1 Disaster consequence prediction
This domain focuses on the predictive analytics of potential disasters’ outcomes, including casualties, infrastructure damage, population displacement and resource requirements, focusing on earthquakes, landslides and floods (Terti et al., 2019; Rahman et al., 2021), providing crucial inputs for the preparation and mitigation phases. The targets of prediction could be modelled as either categorical (e.g. damage states, injury severity) or numerical (e.g. numbers of deaths and injuries) (Galera-Zarco and Floros, 2023; Biswas et al., 2024; Zhang et al., 2025).
Recorded data of the affected were critical in gaining insights into likelihood and severity, e.g. environmental and geophysical data to predict earthquake consequences, duration of precipitation to predict vehicle-related fatality and patients’ records in predicting injury classes (Khosravi and Rabbani, 2026). Socioeconomic indicators are also involved both as the potential scale of vulnerability (e.g. house size, number of commuters) and the population’s preparedness (e.g. number of emergency centres) (Terti et al., 2019; Gong et al., 2024; Zhang et al., 2024a).
Regarding ML algorithms, tree-based ensemble methods are the most prevalent, with random forest and other gradient boosting methods achieving high accuracy (Terti et al., 2019; Araya et al., 2023; Gong et al., 2024). Naturally, this domain is evolving towards multidimensional impact assessments, e.g. combining death and injury counts, infrastructure damage levels (Biswas et al., 2024), cascading effects and the temporal evolution of impacts. These demands encourage the use of deep learning architectures, such as the sequence-to-sequence long short-term memory (LSTM) model for disaster needs prediction in different phases (Nguyen et al., 2022). This domain also advanced from deterministic to probabilistic approaches, enabled by tree-based ensemble methods (e.g. random forest) to estimate the probability of groundwater salinity severity (Araya et al., 2023).
Strong attention to data accuracy and credibility was observed through authoritative sources and cross-validation (Rahman et al., 2021; Xiong et al., 2024), or expert verification (Quiliche et al., 2023). As the scopes of cause-effect are broad, data completeness and relevance were addressed through feature-engineering and data fusion, e.g. multi-year, multi-aspect, multi-source coverage (Rahman et al., 2021; Biswas et al., 2024). The class imbalance resulting from extreme events’ rareness (specific to supervised classification ML) was addressed through sampling and stratification techniques (Araya et al., 2023). Less explored are the domain-specific impacts of data temporal-spatial resolutions, data harmonisation standardisation (e.g. attribute definitions, fusion techniques) and factors that constitute data credibility/trustworthiness.
3.2.2 Relief operations optimisation
This domain focuses on resource allocation at the operational level during or immediately post-event, such as specific relief items based on population dynamics, seasonal and regional characteristics (Lin et al., 2020). Apart from traditional multilayer perceptron (MLP) networks, more sophisticated approaches emerged to handle temporal dependencies and complex patterns, such as LSTM [e.g. fuel demand prediction (Fuqua and Hespeler, 2022)]. Allocation plans can also be optimised while balancing cost efficiency and service-level effectiveness across multiple affected areas and time periods [e.g. RL for relief items (Ahmad et al., 2025) and medicines (Vanvuchelen et al., 2024) distributions, dynamic vehicle routeing and scheduling (Hssini, 2025; Peng et al., 2025) and team-based relief effort optimisation (Wu et al., 2021)]. An advantage of RL is in overcoming model scaling issues to provide fast decision-making in catastrophic and complicated events (van Steenbergen et al., 2023; Hssini, 2025).
The primary inputs vary, including disaster parameters, recorded impacts and consequences, population demographics, resource availability and allocation/transportation costs and typically require data fusion. At this operational level, studies often use high-dimensional, high-resolution temporal and spatial data (Polushko et al., 2024; Zhao et al., 2024a), even real-time or near-real-time information, ranging from coordinates, velocity, direction, to agents’ specifications and population density map (Lin et al., 2020). Concurrently, the reliance on RL techniques necessitates synthetic data for a comprehensive, plausible and relevant representation of potential scenarios (Peng et al., 2025), e.g. sequential action-reward interactions of disaster scenarios. Therefore, the realism of the scenarios used for RL learning was frequently mentioned. Domain expertise was used through either human-in-the-loop or ML-based verification [e.g. deep learning segmentation (Polushko et al., 2024)], ensuring alignment with relevant theories (Lin et al., 2020; Yu et al., 2021; van Steenbergen et al., 2023). Still, no framework was established to ensure data comprehensiveness and reliability, given the dubious robustness of deep learning techniques in extrapolation. Additionally, with the short timeframe for decision-making, potential biases were acknowledged in RL-recommended actions, especially with unreliable data as model inference inputs (Vanvuchelen et al., 2024; Hssini, 2025).
3.2.3 Real-time crisis intelligence
This domain focuses on gaining insights and perceptions about the current emergency. Satellite images were used to assess the level of impact, especially in natural disasters, such as building damage (Hu et al., 2023; Xia et al., 2023; Zhang et al., 2024b), while ground-penetrating radar can be used to map the situation of buildings (Hu et al., 2022). On-site, convolutional neural network (CNN)-based computer vision studies concentrated on person detection in disaster scenarios, using data from cameras, thermal signals or LiDAR inputs from UAVs and CCTV to detect, navigate and locate people in need (Speth et al., 2022; Wen and Chen, 2023; Diez-Tomillo et al., 2024). Additionally, transformer- and LSTM-based NLP techniques can classify social media data into actionable categories, such as sub-events like emotional outbursts or local cases of disease transmission for relief and prevention measures [e.g. COVID-19 (Khatoon et al., 2021; Gour et al., 2022], or casualties, infrastructure damage and rescue progress in other disaster contexts (Alam et al., 2019; Kumar and Singh, 2019; Fan et al., 2020). A natural extension has been to develop multimodal deep learning models that combine CNN-based visual encoders and NLP-based encoders for text to rapidly assess situations with reduced uncertainty through cross-modal corroboration (Fan et al., 2020; Garcia et al., 2023; Zhang et al., 2024b) (Figure 6).
Various techniques are used to ensure text data’s completeness, including stemming, lemmatisation, word segmentation and translation expansion (Khatoon et al., 2021; Ullah et al., 2021; Bhullar et al., 2022; Gour et al., 2022). Apart from corroboration for credibility and filtering for relevance (Ullah et al., 2021; Feng et al., 2025), data accuracy was improved by preprocessing steps, such as spelling correction, phonetic substitution handling and denoising, super-resolution for image data (Fan et al., 2020; Bhullar et al., 2022; Gour et al., 2022).
However, the complications of multi-source data processing pipelines could introduce additional complexity and, in turn, uncertainties. These pipelines are rarely transparent, follow established frameworks or are well-reasoned based on the data sets at hand, obscuring the impacts on data quality dimensions. For example, inconsistent damage assessment scales and ambiguous definitions of “disaster-related” content affect data representativeness and completeness. Additionally, data are costly to label and validate (Fuqua and Hespeler, 2022). The scarcity of training data for deep learning models may also cause latent issues, including feature distortion, shortcut and non-coverage biases and domain shift, which are acute for CNN and transformer architectures (Zhu and Wang, 2023). Finally, despite the significance of low-latency high-resolution data in this domain, the deliberate trade-off between timeliness and accuracy in modelling (Speth et al., 2022) has not been systematically addressed. These data quality concerns are likely the culprit for the significant mismatches observed by Kiavash et al. (2026) between social media-extracted demand signals and other reports.
3.2.4 Infrastructure monitoring, assessment and recovery
This domain supports physical damage and resilience assessment and the restoration of critical infrastructure, ensuring the continuity of essential services and responses. At the tactical level, deep learning techniques were used extensively to investigate the functionality and reliability of the logistics network, i.e. road and bridge systems and prioritising critical components to recover [e.g. radar-based structural assessment (Hu et al., 2022), graph neural networks for network topology (Jana et al., 2023)]. Damages to buildings’ structural integrity and safety were also detected and assessed (Galera-Zarco and Floros, 2023; Zhao et al., 2024a). Multi-agent RL models were proposed to coordinate UAV swarms for communication and computing infrastructure recovery (Faraci et al., 2023).
HSC prioritises functionality in the objective function’s components [e.g. call completion and service availability over communication quality and energy efficiency (Almalki and Angelides, 2019)] while keeping the feasibility of the solutions through the authenticity of scenarios [e.g., GenAI (Hazarika et al., 2025)] and data realism [e.g. visual Turing test (Hu et al., 2022), data abstraction quality (Jana et al., 2023)]. This domain shows attention to data completeness, particularly in infrastructure data, e.g. heterogeneity through merging video footage and physics-based simulation (Galera-Zarco and Floros, 2023; Hu et al., 2023; Zhao et al., 2024a). A typical synthetic data pipeline includes foundation data acquisition, parameter space definition, synthetic data generation and realism enhancement (e.g. generative AI) (Hu et al., 2022; Galera-Zarco and Floros, 2023).
However, presumptions regarding critical data configuration were not always strongly addressed, e.g. anomaly detection assuming cracked points are minor, with arbitrary thresholds (Chen and Cho, 2022). Frequently, data processing decisions were reasoned concerning individual quality regardless of how they might negatively affect other dimensions [e.g. deliberate annotation bias inconsistency (Zhao et al., 2024a) and realistic scenarios by Generative AI may not be comprehensive (Hazarika et al., 2025)].
3.2.5 Population mobility and evacuation
This domain focuses on the movement pattern of population during and after events, which is valuable for HSC operation optimisation at the tactical and operational levels [e.g. predictions of travel and service times for evacuation using XGBoost (Nabavi et al., 2022) and emergency behaviour and emergency mobility using deep RL (Song et al., 2017)]. Due to the domain’s dynamic nature and its relationships with different social and economic factors, data processing heavily involved advanced location discovery and behavioural analytics, e.g. location’s functions and familiarity and social relationships (Song et al., 2017). The challenges were acknowledged in fusing positioning, event and transportation network data, news reports and feature engineering to ensure data relevance and completeness [e.g. transportation mode with destination labelling for millions of anonymised users for several years with billions of GPS records (Song et al., 2017)]. Validating data quality is often overlooked, especially in complex data fusion processes, despite its systematic implications for model quality. Another issue is the practice of assessing data quality using model accuracy indicators, an indirect approach that may create false expectations about the model’s performance in real-world operations.
3.2.6 Strategic humanitarian supply chain networks planning
This domain optimises HSC facility decisions and network planning to enhance resilience. For example, network optimality can be estimated using topological measurements (e.g. centrality indices). A classification model trained on that pattern can quickly and efficiently optimise distribution centres’ locations [e.g. deep neural network (Taouktsis and Zikopoulos, 2024)]. Similarly, the importance of facilities can be used to strategise resource deployment to precisely improve system resilience (Jana et al., 2023). Models can predict operational performance across different scenarios and management strategies [e.g. cumulative demand, initial inventory, management strategies and cumulative delivered supply to predict the average cumulative delivery percentage and the daily cumulative percentage of fulfilled demand (Zhang et al., 2024a)].
These studies require both synthetic and authentic data to ensure completeness, a manageable computational workload and model generalisability. For example, clustering was used to effectively and efficiently stratify disaster scenarios (Hu et al., 2024). ML-based surrogate modelling was used to efficiently approximate the results of computationally heavy and repetitive programming tasks in optimisation [e.g. two-stage optimisation of pre-event location and resource allocation and post-event procurement, distribution and inventory (Pan et al., 2025)]. Data credibility was classified into three levels, with historical data ranked the highest – used to validate the simulation models, which are then used to validate ML model results (Zhang et al., 2024a).
While data generation parameters are mostly transparent, their rationale and verification are not always clear, affecting the basis for validity and realism (e.g. synthetic coordinates, distances between distribution centres and impact/consequence levels for each synthetic scenario). Some studies relied on experts for plausibility and representativeness verification (Zhang et al., 2024a). However, data representativeness from an expert’s perspective can be ambiguous regarding statistical distributions, completeness, or “operational possibility”. Additionally, while data imbalance was frequently mentioned as a quality-related issue [e.g. nodes being the minority among many other nodes (Taouktsis and Zikopoulos, 2024)], the imbalance as a nature of the phenomenon is not inherently a data quality issue and should be considered under the appropriateness of modelling techniques.
The key HSCML data and ten data quality dimensions derived from analyses across different application domains are aggregated in Figure 7. It is noteworthy that operationalising these dimensions requires domain-specific indicators and thresholds (see Section 4).
3.3 Humanitarian supply chain machine learning model quality
Model quality considerations can be categorised into a priori rationales (i.e. reasoning for algorithms and pre-trained models) and post-training testing and evaluation. This separation not only provides insights into the dimensions of model quality and the technical considerations in the modelling process but also helps identify quality gaps in model evaluation.
3.3.1 A priori rationales for machine learning techniques
Data-related rationales address only partially the data-related challenges in the humanitarian context (see Section 3.2). Apart from big data analytics, handling noisy data became critical for NLP and transformer-based models working with unofficial, unvalidated sources, such as real-time handling of colloquial expressions, abbreviations and multilingual social media posts (Kumar and Singh, 2019; Saleem et al., 2024) (Figure 8). To gain preprocessing insights into data sets, data exploration is usually deployed, in which unsupervised ML can be valuable [e.g. LDA (Zhao et al., 2024b)].
Some rationales are specific and decisive towards methodological choice. For example, the rarity of disaster events can be addressed through transfer learning (Zheng et al., 2022), or data imbalance and scarcity are addressed by stratification techniques (Hu et al., 2024). Additionally, purpose-built architectures and pre-trained models have proven critical in complex HSCML solutions, such as object recognition and image segmentation and reconstruction [e.g. YOLOv5 for building damage (Hu et al., 2023), segmentation models (Polushko et al., 2024)]. Handling complex multi-factor interactions and cascading effects is another popular rationale, such as to handle temporally dependent data [e.g. LSTM for demand forecasting (Nguyen et al., 2022)] and complex multi-objective optimisation of spatial-temporal decisions under uncertainty [e.g. multi-agent RL for UAV coordination (Hazarika et al., 2025)].
Modelling technique rationales focus on operational and deployment requirements. Modelling accuracy remains paramount in HSCML, e.g. resource misallocation can be catastrophic, motivating the choice based on known theoretical performance [e.g. deep learning (Song et al., 2017), tree-based and boosting algorithms (Nabavi et al., 2022), stacking ensembles (Zhang et al., 2025)]. Computational efficiency in real-time inference and training/retraining addresses the demands for rapid predictions and model updates in emergency responses [e.g. surrogate modelling for relief logistics (Pan et al., 2025), CNN for victim detection (Wen and Chen, 2023)], in which model size is a deliberate trade-off between accuracy and deployability (Hu et al., 2023; Heuschmid et al., 2025).
Another trade-off is between interpretability and other model quality dimensions, such as explainability, reliability [e.g. tree-based algorithms for transparent and deterministic feature importance (Araya et al., 2023; Quiliche et al., 2023)]. Apart from the advantage of ML in its non-assumption of predictor-target relationships, model robustness is cited as critical to ensure performance when dealing with outliers and data uncertainties, e.g. robustness to viewpoint variations and illumination conditions (Wen and Chen, 2023; Polushko et al., 2024; Zhao et al., 2024a) and seasonality (Vanvuchelen et al., 2024). Finally, the method/technique’s proven track record is deemed significant, citing the fit with data and prior studies (Garcia et al., 2023; Yağcıoğlu, 2025).
However, significant gaps exist. Firstly, there is a potential mismatch between the rationales’ generality and techniques and algorithms’ specificity, i.e. studies reasoned technical choices (e.g. neural networks, or pre-trained models) based on ML’s umbrella characteristics, such as nonlinear relationships and big data handling. Studies used vague, ambiguous and superficial reasons, such as “high-performance”, “high accuracy”, or even “widely used”. These weaknesses in the methodological basis and premature conclusions could lead to overlooking suitable methodologies, especially given the trade-offs among model quality dimensions and the strong individuality of implementations.
Secondly, despite the growing recognition of uncertainty quantification in ML (Abdar et al., 2021), HSCML studies provide limited rationale for handling and expressing uncertainty beyond data-related robustness (17/93 reviewed articles ∼ 18.3%). The lack of uncertainty-conscious methodological choices (e.g. confidence and prediction intervals, noise decomposition, quantile predictions, ensemble variance) might affect model applicability, as both aleatoric and epistemic uncertainties can propagate through HSC, from strategic to operational levels.
Thirdly, despite explainable AI has become recognised as essential for high-stakes decision-making (Barredo Arrieta et al., 2020) and even a structural requirement for humanitarian AI (Dorsey, 2026), there is still limited consideration of explainability beyond feature importance. In HSC’s collaborative environment, stakeholder-centred explainability is non-negotiable for execution. More critical decisions favour the use of tractable calculations and situation-solution relationships at the expense of model accuracy (e.g. tree-based models rather than neural networks).
Fourthly, despite the criticality of fair and equitable ML (Caton and Haas, 2024), efforts in HSCML to incorporate and consider these qualities in solution design are still very limited (11/93 reviewed articles ∼ 11.8%), e.g. resource allocation with fairness-integrated objective functions (Vanvuchelen et al., 2024; Ahmad et al., 2025). This negligence indicates a disconnection from the technical choices and the desired model quality.
Overall, in comparison with well-known, established general ML system design frameworks (Sculley et al., 2015), HSCML is currently behind in some key areas, including overall ML-based solution design, interpretability and conscious and predetermined trade-offs between aspects of model performance (e.g. errors’ costs) and equity and fairness. These shortcomings highlight the urgency of a framework for model quality dimensions.
3.3.2 Model quality dimensions and influencing factors
The identified HSCML model quality dimensions can be categorised into three groups (Figure 9). Technical performance dimensions were examined beyond mere accuracy. Computational efficiency and responsiveness address time-critical humanitarian operations through inference latency measurements or initialisation time requirements [e.g. CNN-based models (Diez-Tomillo et al., 2024)] and training duration constraints [e.g. transformer architectures (Saleem et al., 2024)]. Robustness and generalisability ensure reliable performance through training monitoring with [e.g. early stopping to control overfitting and model bias (Hu et al., 2023), accuracy under imperfect data (Gong et al., 2024) and cross-validation across (Jana et al., 2023)].
Regarding trustworthiness and transparency, a similar result to Section 3.3.1 was derived, as explainability and interpretability are conflated with feature importance quantification [e.g. Shapley-based importance (Feng et al., 2025), permutation importance (Terti et al., 2019)] and calculation traceability [e.g. to explain HSC’s predicted performance aspects (Zhang et al., 2024a)]. Although not always included, uncertainty was quantified through different techniques, such as the variance of results from ensembles and bootstrap samples (Ahani et al., 2021; De Santi et al., 2021). Regarding operational performance, solution optimality evaluates the objective achievement of HSCML-optimised outputs through indicators [e.g. survivability metrics (Lee and Lee, 2021)], demand fulfilment metrics (Vanvuchelen et al., 2024), multi-criteria performance trade-offs, such as efficiency and equity (Ahmad et al., 2025) and flow-specific metrics, such as energy consumption and delay in network recovery (Yağcıoğlu, 2025).
Regarding fairness and equity, the disaster consequence prediction domain demonstrates emerging but limited considerations, primarily through socioeconomic input variables (e.g. cost of living, per capita income) and a vulnerability-centred model design (Biswas et al., 2024). Cost of errors (e.g. deprivation cost of misclassifying at-risk households) can be used to deliberately adjust the balance between accuracy and humanitarian optimality (Quiliche et al., 2023). Infrastructure monitoring, assessment and recovery rarely engage with fairness and equity. In RL-based optimisation, considerations can be injected into the reward function [e.g. social vulnerability based on income, education and housing conditions (Ghannad et al., 2021) and fairness index for equitable resource allocation (Yağcıoğlu, 2025)]. Similarly, some relief operation optimisation studies balanced efficiency, effectiveness and equity through deprivation costs [e.g. penalty costs to ensure no survivors are severely disadvantaged) (Yu et al., 2021; Ahmad et al., 2025]. Other studies addressed service equity for disadvantaged communities, such as unserved area penalties (Peng et al., 2025) and geographic fairness in medicine access by minimising variance across facilities (Vanvuchelen et al., 2024). Real-time crisis intelligence acknowledged geographic biases, such as social media attention is not uniformly distributed across disaster-affected areas (Fan et al., 2020). Studies in strategic HSC network planning and population mobility and evacuation domains demonstrated a lack of consideration for fairness and equity, prioritising only technical efficiency (e.g. network efficiency, traditional supply chain performance metrics, prediction accuracy, evacuation time and costs).
Across domains, fairness and equity remain critically underexamined and are often considered as post hoc assessments rather than predetermined objectives. Both real-time crisis intelligence and disaster consequence prediction recognise that data availability and attention are geographically uneven, yet no corrective frameworks or systematic fairness integration have been proposed, thereby perpetuating spatial inequities and marginalisation (e.g. digitally disconnected populations). The balance among efficiency, effectiveness and equity is adjustable, but the quantification of weights, influenced by political, societal and cultural factors, was overlooked. In HSC networks, planning and infrastructure management, facility location and network design decisions can compound social disparities (e.g. transportation access, socioeconomic constraints and vulnerable populations’ evacuation capacity), calling for more dedicated efforts in demographic disaggregation to ensure tailor-made yet equitable assistance.
Additionally, the existing dimensions focus primarily on static, lab-environment metrics rather than situational adaptability and application-oriented conditions, such as reliability in handling data imperfections and concept drifts (e.g. population dynamics, evolving conflicts, climate change) (Nguyen et al., 2023). While transfer learning can enable generalisability, HSCML’s ability to correctly render fundamental understandings was not evaluated, undermining adoption confidence. Moreover, there is a lack of discussion of the importance of model quality dimensions and how their trade-offs should be considered. Finally, uncertainty quantification is not widely and systematically applied (e.g. aleatoric and epistemic), which could lead to ineffective and inefficient decision-making.
Figure 10 presents a comprehensive taxonomy of factors influencing HSCML model quality, thematically aggregated from the reviewed papers. Model quality is directly shaped by data-centric, computational and infrastructure, algorithmic and methodological and complexity factors. The outermost layer captures environmental and domain-specific influences, where decisions should be grounded in legal, ethical and situational awareness (Dorsey, 2026). This framework guides a systematic consideration of the factors that influence ML model quality during planning, development and improvement. Considering the recent development of ML in LSCM (Vlachos and Reddy, 2025; Nguyen et al., 2026), there is a gap in the sustainability of HSCML models in deployment, i.e. functional prototypes are evaluated and tested locally, while regimes for HSCML application-oriented testing are nonexistent, overlooking critical elements (e.g. performance monitoring and retraining). In addition, HSCML relies on software, hardware and infrastructure that may be subject to bottlenecks and dependencies. Examples in the software realm include data supply chain and software dependencies (e.g. data quality dimensions along the pipeline, computational overheads and biases within base and foundational models). The complexity of solutions is even higher for multi-model chaining, coupling and stacking, but these were not discussed.
4. Discussions and future research directions
HSCML efforts were unevenly distributed across disaster types, management phases and geographic regions. There is limited attention to slow-onset events (e.g. droughts, pandemics) and human-made crises (e.g. conflicts and industrial accidents), despite their complex and emerging impacts. This review confirms the role of digital innovations, here HSCML, in the response phase (Maric et al., 2022; Altay et al., 2023) and indicates a lack of attention to recovery and mitigation, as well as strategic-level decisions on HSC sustainability and resilience. Compound and cascading disasters pose modelling challenges that have not been adequately addressed, representing opportunities for HSCML applications. HSCML research also neglected LDCs despite their humanitarian challenges. While human factors such as know-how and competency and technological constraints contribute to the barriers (de Camargo Fiorini et al., 2022; Heuschmid et al., 2025; Behl et al., 2026), data scarcity is the most self-reinforcing obstruction across the ecosystem. The current sparse and fragmented historical disaster records preclude the applicability of ML to more complex, gradual and latent phenomena, further limiting our understanding of long-term sustainability and performance of HSCML systems and calling for cross-country data sharing, funded regional databanks and local expertise development to enable and generalise HSCML research.
Across the six HSCML application domains, two model architectures dominate: tree-based ensemble methods (e.g. random forest, XGBoost) and neural-network-based methods (e.g. CNN, LSTM, transformer). In the disaster consequence prediction domain, where tabular environmental and socioeconomic data predominate, tree-based methods prevail, offering native probabilistic outputs, interpretable feature importance and reliable performance. However, they are critically limited in extrapolation, i.e. unable to generalise beyond the training distributions, thereby constraining their applicability to significant and rare events. Deep learning architectures dominate high-dimensional, multimodal domains, including CNNs and graph neural networks in crisis intelligence and infrastructure monitoring for visual and spatial assessment, LSTMs for temporal modelling in demand forecasting and transformer-based encoders for NLP and multimodal pipelines. Despite their versatility, they are vulnerable to shortcut learning and feature distortion under data scarcity, require explicit uncertainty quantification methods and depend on post hoc approaches for explainability and fairness. RL, most frequently implemented through neural network architectures, has been mainly applied in distribution optimisation and infrastructure recovery contexts but carries its own quality considerations, such as reward-signal sensitivity, out-of-distribution state failures and limited explainability across decision trajectories.
In practice, deployable HSCML systems will increasingly involve multi-model pipelines where architectures are chained sequentially, e.g. transformer-based NLP, such as BERTopic, maps the demand feeding RL-based distribution optimisation and GNN-based network criticality coordinates with multi-agent RL for infrastructure recovery. Methodological decisions must therefore be considered at the pipeline (system) level rather than for each sub-problem in isolation, where architectural compatibility, propagation of failure modes and system-wide computational constraints collectively determine operational feasibility. The data quality framework (Figure 7), model quality dimensions (Figure 9) and influencing factors taxonomy (Figure 10) provide the reference for this system-level design, ensuring quality assurance is built into the pipeline architecture rather than assessed post hoc. This review also expands Kondraganti et al. (2022) by specifying six additional data categories beyond social media and clarifying their roles across application domains.
In response to the calls of Anjomshoae et al. (2022) and de Camargo Fiorini et al. (2022) for data-driven knowledge for humanitarian operations, four operational necessities for sustainable HSCML are recommended. Firstly, regarding the lack of good-quality data (Anjomshoae et al., 2022; Kondraganti et al., 2022), HSCML studies address data quality concerns piecemeal, focusing on data set-specific challenges and conforming to common practices without a cross-domain standard. This narrowness and individuality are critical, as they cause field-wide inconsistencies that go unaddressed and unresolvable. The aggregated data quality dimensions framework (Figure 7) addresses this by offering a comprehensive, domain-transferable reference for data quality.
Secondly, while ML-based systems need retraining (Vlachos and Reddy, 2025), data processing pipelines in HSCML are rarely transparent or well-grounded and regimes for continuous data quality assurance (e.g. pipeline maintenance) are nonexistent. As found in this review, HSCML studies are predominantly proof-of-concept evaluated in controlled settings. While retraining, monitoring and pipeline maintenance fall outside the typical design scope, they are critical gaps for long-term system sustainability in resource-constrained environments (Ülkü et al., 2024). This confirms Maric et al.’s (2022) assessment that HSCML remains primarily technology-centric rather than humanitarian- and supply-chain-integrated, for which tensions with operational readiness have been observed (Dorsey, 2026).
This review indicates critical prerequisites for deployable HSCML, including data source integration and compatibility, pipeline-level quality management and broader data and model quality dimensions (Figures 7 and 9), calling for a comprehensive framework of HSC performance indicators, which could be built on the established LSCM foundation (Anjomshoae et al., 2022). Data processing pipelines and modelling technique selection should be designed with purpose and tailored to contexts, rather than replicated from prior studies, given the individuality of data-model configurations and evolving data quality landscapes, e.g. improving sensor quality and deteriorating social media data caused by chatbots and generative AI content.
Thirdly, despite HSCML’s reliance on multi-source data fusion, stakeholder participation and data governance factors (Figure 10) remain largely overlooked. HSCML data is inherently cross-organisational, requiring separate access and sharing agreements, yet the reviewed literature predominantly treats data as pre-given, bypassing the governance required to secure, integrate and maintain it. Given the identified barriers of resistance due to resource competition and accountability concerns (Anjomshoae et al., 2022) and the prerequisite of data privacy, transparency and data management for stakeholder trust (Vlachos and Reddy, 2025), future research on governance, such as data valuation based on contributions to model performance, provenance tracking, cross-organisational quality assurance will be critical.
Fourthly, domain expertise through human-in-the-loop or expert verification is proven critical (de Camargo Fiorini et al., 2022). As model complexity scales, the expertise required to validate model behaviours and outputs grows accordingly. Yet, this capability has not been nurtured and scaled proportionately, preventing the realisation of HSCML’s true potential. It is known that automation bias and cognitive offloading can progressively erode judgement operators’ independent judgement and oversight capacity through over-reliance on AI outputs, i.e. “cognitive shifting” (Behl et al., 2026; Dorsey, 2026). Explainability is therefore the functional prerequisite for an operationally ready HSCML, enabling operators to interrogate model decisions, identify failures and maintain systems throughout their operational life cycle.
The review results indicate dominance of technical-induced model quality dimensions, both in a priori rationales and in the evaluation of model quality. At the same time, HSC characteristics, such as fairness and equity, are underexamined. Where present, the considerations are implemented narrowly through socioeconomic input variables and RL reward functions and treated as post hoc assessments. Incorporating fairness adds complexity to model design and may require trade-offs, as the framework in Figure 10 indicates, without improving the more obvious, easy-to-defend accuracy and computational metrics. Therefore, it is systematically deprioritised. This finding extends the concern by Nawazish et al. (2023) that the social pillar of sustainability in HSC remains nascent. It underscores the difficulty of translating humanitarian principles, such as impartiality, into measurable performance indicators. Future research should integrate equity throughout the HSCML development lifecycle, e.g. incorporating related indices as input features and model constraints, implementing demographic disaggregation in model outputs, developing tuneable multi-objective efficiency-equity trade-off frameworks, correcting geographic and demographic biases and establishing participatory design protocols that involve beneficiaries in the co-design of equity metrics. A view beyond social equity towards a full quadruple bottom-line sustainability framework, including economic, environmental, social and cultural elements (Ülkü et al., 2024), will also benefit adoption, as environmental and cultural sustainability remain largely underexplored in the current HSCML literature.
Lastly, the review reveals persistent weaknesses in methodological rigour and operational readiness, i.e. a tendency to justify methodological choices by appealing to ML’s umbrella characteristics rather than to contextual suitability and the specificity of algorithms/techniques. While concurring with Vlachos and Reddy (2025) on the versatility of neural networks, this review identified a multidimensional range of factors influencing model quality (Figure 10) that do not always favour deep learning. These loose reasonings and the limited consideration of trade-offs between model quality dimensions indicate again that HSCML publications’ focus remains on proof-of-concept rather than method optimality (Sculley et al., 2015; Maric et al., 2022; Altay et al., 2023). HSCML spans across HSC and computer science communities with different evaluation standards and has inherited the latter’s benchmark-oriented confidence without the former’s insistence on quality control. Integration and system-scale testing and accountability frameworks with regimes for performance monitoring and retraining (e.g. Nguyen et al. (2026)) should be developed to facilitate HSCML deployment.
The discussions above highlighted the gaps that hinder HSCML from maturing from technical prototypes to impact-oriented, operationally ready applications. These discussions are not only enabled by but also operationalise the three synthesised frameworks, providing a principled reference for development that prioritises deployability. The research directions in Figure 11 translate the identified gaps into a structured agenda for the proper adoption and assimilation of HSCML.
5. Conclusion
This study addresses the critical gap in understanding how ML contributes to HSC through an LDA-guided, CIMO-structured, mixed-methods SLR. As the first systematic examination of data and model quality in HSCML, this review produced three novel deliverables:
a CIMO-structured characterisation of HSCML across contexts and six application domains;
aggregated frameworks for data quality dimensions, model quality dimensions and factors influencing model quality; and
a deployability-focused, data-centric research agenda comprising 14 knowledge gaps and recommended directions across context, data and model dimensions.
Extending previous reviews, this study reveals the methodological impact mechanisms driving the disconnect between technical development and operational requirements, which are gaps that were overlooked by prior work.
This review offers structured guidance for multiple stakeholders. For researchers, the quality frameworks and research agenda provide systematic entry points for future work, particularly in the areas of uncertainty quantification, fairness-integrated and multi-model HSCML systems and deployment-oriented testing regimes identified across all six application domains. The context-focused review also identified underserved areas, such as LDCs, neglected disaster types and management phases and critical HSC stages and flows, such as procurement and finance.
For practitioners and decision-makers,HSCML operates in high-stakes environments where errors and biases translate directly into humanitarian harm. The systematic underinvestment in data quality, fairness, uncertainty quantification and explainability documented is not merely a technical deficiency but a governance risk that calls for institutional action, such as HSCML accountability frameworks, bias monitoring and demonstrable quality assurance as prerequisites for deployment, in line with obligations under international humanitarian law. The three synthesised frameworks from this literature review provide practical guidance for developing, evaluating and adopting HSCML with appropriate considerations.
Three limitations of this review should be acknowledged. Firstly, given the breadth of HSC domains identified, certain aspects, including ethical considerations, socio-political dimensions and organisational implementation factors, warrant dedicated systematic reviews beyond our scope. Secondly, this synthesis of published research would benefit from triangulation with primary data from HSCML practitioners, providing deeper insights into implementation challenges and operational barriers that the academic literature may not fully capture. Thirdly, given the field-level positioning of this review, specific pipeline design discussions and best practice recommendations for individual HSCML application domains were not covered. A dedicated sub-field SLR that examines pipeline-level practices in depth will be valuable, for which the established frameworks in this paper are intended to inform and scaffold.
The authors would like to thank Ms Nguyet Huynh from RMIT Vietnam Library for her valuable support in the curation of literature for this review and the editor-in-chief and two anonymous reviewers for their constructive comments and helpful suggestions on a previous draft of this paper.












