Purpose

This study examines whether a multidimensional supply-chain readiness system predicts export performance across African economies and whether the relationships persist when export performance is measured through aggregate value, extensive product breadth and intensive product depth rather than exports relative to GDP.

Design/methodology/approach

The analysis begins with a 54-country, 1985–2024 country-year grid while retaining original missing observations. Export intensity is retained as a benchmark. Additional outcomes comprise log merchandise export value, the number of exported HS-6 products and log export value per HS-6 product. Ridge, Random Forest, Extra Trees and Gradient Boosting are assessed with temporal, random, country-group and rolling validation.

Findings

Extra Trees remains the strongest benchmark model (R-squared = 0.651). Predictive performance is also high for log merchandise exports (R-squared = 0.926), exported HS-6 product breadth (R-squared = 0.891) and log exports per HS-6 product (R-squared = 0.824). The complete-case benchmark, estimated without numerical imputation, records R-squared = 0.737. Country demeaning lowers R-squared to 0.242 for export intensity and 0.283 for log merchandise exports, confirming that pooled machine learning captures substantial between-country structure. Within-country importance nevertheless retains industry employment for export intensity and manufacturing value added for merchandise export value as leading readiness signals.

Originality/value

The study combines interpretable machine learning with alternative export-performance margins, explicit missing-data sensitivity and a pooled-versus-within-country bridge. The results distinguish predictive cross-country structure from within-country change and provide a more cautious basis for connecting supply-chain readiness to export breadth and depth.

Export performance remains central to African development because it connects domestic production to foreign demand, scale economies, learning and structural transformation. Trade theory links export outcomes to technology, factor costs, firm productivity, market access and trade costs (Melitz and Redding, 2021; Larch and Yotov, 2024; Antràs and Chor, 2022). These channels matter in Africa, where uneven infrastructure, fragmented logistics systems, narrow production bases and uneven digital connectivity constrain cross-border participation (Kuteyi and Winkler, 2022).

A supply-chain readiness perspective explains why export outcomes differ across African subregions. Exporting requires comparative advantage, but it also requires supplier coordination, reliable electricity, information systems, finance, transport connectivity and dependable border processes. Supply-chain research links performance to coordination, resilience, information flows and logistics reliability (Govindan et al., 2022; Jha et al., 2022; Castillo, 2023), while African trade evidence identifies infrastructure gaps and border delays as persistent constraints (Tandrayen-Ragoobur et al., 2023; Olyanga et al., 2022; Wassie et al., 2025).

The determinant set is broadened beyond general macroeconomic proxies. Vogel (2025) shows that African export diversification reflects structural conditions, tariffs, regional integration, institutions, services, resource dependence, finance, foreign investment and exchange-rate stability. Consistent with that evidence, the present framework incorporates digital readiness, energy access and efficiency, manufacturing capacity, industrial employment, private credit, production-cost conditions, trade facilitation, transport connectivity, export capacity and direct logistics measures. GDP per capita, population, density, services, lagged FDI, resource rents and geographic characteristics are treated as controls rather than direct readiness measures.

The research question is therefore whether supply-chain readiness predicts export performance across African economies, whether those predictive relationships remain stable in later periods, and how strongly they persist under alternative aggregate, extensive-margin and intensive-margin outcomes. Exports of goods and services as a percentage of GDP is retained only as the export-intensity benchmark. Trade (% of GDP), imports (% of GDP) and current account balance remain excluded from the main specifications because they are mechanically or substantively close to that benchmark and may create accounting overlap or trade-openness endogeneity.

The contribution is in fivefold. Firstly, the study defines supply-chain readiness as an antecedent national capacity and distinguishes it from realised logistics performance, border-process facilitation, disruption resilience and foreign-market preparedness. Second, it separates direct readiness indicators from structural controls and outcome-related exclusions. Third, it uses temporal, country-group and rolling validation rather than relying on a random split alone. Fourth, it tests export performance through aggregate value, product breadth and product depth to reduce dependence on the exports-to-GDP ratio. Fifth, it combines SHAP, permutation importance, complete-case sensitivity, country-and-year fixed effects and country-demeaned machine learning so that pooled structural prediction is not mistaken for within-country progress or causal policy impact.

Supply-chain readiness is defined here as the ex-ante capacity of an economy to coordinate production inputs, infrastructure, information, finance, border processes and delivery systems before export transactions occur. It is broader than logistics performance, which reflects realised customs efficiency, infrastructure quality, shipment arrangement, logistics competence, tracking and timeliness (World Bank, 2023). It is also broader than trade facilitation, which concentrates on the time, cost, documentation and predictability of border procedures (Wassie et al., 2025).

Supply-chain resilience concerns the ability of production and logistics networks to absorb, adapt to and recover from disruption (Ivanov and Dolgui, 2020; Castillo, 2023), whereas export readiness commonly concerns preparedness to enter and serve foreign markets. The present construct treats these concepts as related but non-equivalent: readiness represents the underlying capacity set; logistics performance and border outcomes represent observable manifestations; resilience concerns response under disruption; and export readiness is closer to firm- or product-level market-entry preparedness.

Two complementary theoretical perspectives provide the foundation for the readiness construct. Heterogeneous-firm trade theory emphasises productivity and the fixed and variable costs of entering and serving foreign markets; at the national level, electricity, finance, productive capacity, digital connectivity and transport conditions shape the environment in which firms absorb those costs and scale export activity (Melitz and Redding, 2021; Larch and Yotov, 2024). A transaction-cost and supply-chain capability perspective adds the coordination mechanism: reliable information, supplier linkages, logistics competence and predictable border processes reduce search, contracting, delay and delivery frictions across fragmented production networks (Antràs and Chor, 2022; Govindan et al., 2022; Jha et al., 2022). These mechanisms imply effects on both margins of trade. Readiness can broaden the set of viable product-market combinations and can deepen established export relationships by lowering recurring coordination and delivery costs. The alternative outcome analysis therefore treats aggregate value, HS-6 product breadth and average value per HS-6 product as distinct empirical manifestations of export capacity.

Digital connectivity lowers search, communication and coordination costs. Internet access and mobile connectivity can support buyer discovery, transaction records, shipment coordination and regional market participation (Kere and Zongo, 2023; Xing et al., 2026). Energy access and lower power losses support production continuity, cold storage, digital systems and service delivery. Productive and financial readiness is reflected in manufacturing value added, industrial employment, capital formation and private-sector credit, consistent with evidence that productive capabilities shape African export performance (Avenyo et al., 2021).

Trade facilitation and transport connectivity are added as separate readiness dimensions. Border and documentary compliance time and cost, export lead time, tariffs, air freight and liner-shipping connectivity capture frictions that are not represented by broad income or sectoral shares. Direct logistics measures cover customs, infrastructure, international shipments, logistics competence, tracking and timeliness. Firm-level and Africa-focused evidence links these conditions to export outcomes and international competitiveness (Bugarčić et al., 2024; Wassie et al., 2025).

Controls are retained to represent structural heterogeneity rather than readiness interventions. GDP per capita captures development level; population and density capture market size and spatial concentration; services value added captures economic structure; lagged FDI captures external capital; resource rents capture commodity dependence; and subregion, landlocked and island status capture geography. Vogel (2025) reinforces the need to account for such contextual differences when analysing African export outcomes.

Conventional panel models are useful for estimating average conditional associations under a specified functional form. Machine learning serves a different purpose: it can accommodate nonlinear thresholds, interactions, missingness indicators and heterogeneous predictor combinations while evaluating out-of-sample accuracy (Masini et al., 2023; Goulet Coulombe, 2024). The approach is therefore selected for prediction and pattern discovery, not causal identification. A country-and-year fixed-effects model with lagged predictors is retained as a conventional benchmark so that predictive rankings can be compared with within-country associations.

Interpretability is treated as part of the design rather than an afterthought. Permutation importance evaluates loss of holdout accuracy after a feature is disrupted; SHAP values show the direction and size of local predictive contributions; dependence plots reveal nonlinear patterns and interaction colouring; subregional errors expose geographic heterogeneity; and rolling validation evaluates stability across successive later-period windows. (See Supplementary Table S5 and Supplementary Figures S3, S4, S5 and S6).

The data structure is a 54-country by 40-year grid covering 1985–2024, equivalent to 2,160 possible country-years. The grid is not treated as a complete balanced analytical sample. Original missing observations are retained and reported. The export-intensity benchmark is observed for 1,735 country-years. The core annual specification uses 1,711 observations from 51 economies; the extended specification uses 1,711 observations from 51 economies; and the direct-logistics specification uses 467 observations from 51 economies over 2007–2022. The alternative outcomes add 1,962 observations for log merchandise export value across 54 economies during 1988–2024 and 1,185 observations for HS-6 product breadth and product-depth measures across 49 economies during 1990–2023. (See Table 1). (See Supplementary Table S3 and Supplementary Figure S8).

Table 1

Analytical samples and missing-data treatment

Analytical scopeCountriesRowsPeriodTreatment
Target country-year grid542,1601985–2024Missing observations retained
Observed export-performance values541,7351985–2024Outcome available
Core annual specification511,7111985–2024Temporal cutoff: 2017
Extended readiness specification511,7111985–2024Temporal cutoff: 2017
Direct logistics specification514672007–2022Temporal cutoff: 2018
Aggregate export-value outcome541,9621988–2024Temporal cutoff: 2017
HS-6 product breadth/depth outcomes491,1851990–2023Temporal cutoff: 2017
Source(s): Computations from the study dataset (2026)

Export intensity, measured as exports of goods and services as a percentage of GDP, is retained as the benchmark rather than treated as a complete measure of export performance. To reduce structural scale bias, three estimable alternatives are analysed: the natural logarithm of total merchandise exports in current US dollars, the number of HS-6 products exported in a reporter-year, and the natural logarithm of merchandise export value per exported HS-6 product. The intended destination-partner breadth and per-partner depth outcomes did not yield a usable country-year series in the retrieved WITS indicators and are reported as unavailable rather than filled by imputation. The product-based extensive and intensive margins therefore provide the estimable decomposition of breadth and depth.

Table 2 separates direct readiness measures from controls and explicitly records the three outcome-related exclusions. This treatment prevents the high importance previously attached to imports from being interpreted as evidence of readiness when it may largely reflect general trade openness. (See Supplementary Table S1, Supplementary Table S4 and Supplementary Figure S2).

Table 2

Variable framework and modelling treatment

RoleDimensionIndicatorsTreatment
OutcomeExport performanceExports of goods and services (% of GDP)Dependent variable; described as export intensity
Core readinessDigital, energy, productive and financial capacityInternet use; mobile subscriptions; electricity access; manufacturing value added; industry employment; private-sector credit; inflationPrimary annual specification
Extended readinessTrade facilitation, transport, production cost and export capacityPower losses; tariffs; air freight; shipping connectivity; lead time; border and documentary time/cost; documents; REER; manufactured and high-technology exports; secure serversRobustness specification
Direct logisticsOperational logistics qualityCustoms; infrastructure; international shipments; logistics competence; tracking; timelinessSparse-year robustness specification
ControlsDevelopment, structure and geographyLog GDP per capita; log population; density; services; lagged FDI; resource rents; subregion; landlocked; islandNot interpreted as direct readiness levers
ExcludedOutcome-related aggregatesTrade (% of GDP); imports (% of GDP); current account balanceRemoved from the main specifications
Outcome robustnessAggregate value; extensive and intensive product marginsLog merchandise exports; exported HS-6 products; log exports per HS-6 productAlternative outcome analysis
Source(s): Computations from the study dataset (2026)

The benchmark pipelines retain numerical median imputation learned from each training sample, together with missingness indicators, while categorical imputation and one-hot encoding are also learned within training data. The sensitivity design directly tests whether this treatment drives the findings. For the alternative-outcome and pooled-versus-within-country analyses, predictors with more than 50% missingness in the relevant benchmark sample are excluded; the retained variables have missingness between 0% and 22.02%, while gross fixed capital formation is excluded because it has no usable observations. A separate complete-case Extra Trees model uses no numerical imputation. Sparse trade-facilitation and logistics variables remain confined to their period-specific robustness specifications. (See Supplementary Tables S2 and S10).

The supervised stage estimates a mean benchmark, ridge regression, Random Forest, Extra Trees and Gradient Boosting. The primary export-intensity test trains before 2017 and evaluates 2017 onward; the direct-logistics specification uses a 2018 cutoff because its measures are available only in selected years. The same 2017 temporal design is applied to log merchandise export value, exported HS-6 product breadth and log export value per HS-6 product. Random observation splitting remains a secondary comparison, country-group cross-validation keeps countries separated across folds, and four rolling three-year windows assess later-period stability. (See Supplementary Table S5 and Supplementary Figures S3 and S4).

The core specification includes annual digital, energy, productive and financial indicators. The extended specification adds energy efficiency, tariffs, air freight, liner-shipping connectivity, export lead time, border and documentary compliance, the number of export documents, the real effective exchange rate, manufactured and high-technology exports and secure internet servers. The direct-logistics specification uses the six Logistics Performance Index components together with available trade-facilitation and shipping measures.

Model performance is evaluated with mean absolute error, root mean squared error and R-squared. Permutation importance and SHAP are calculated for the strongest temporal tree model, and alternative-outcome permutation importance is used to compare structural controls with readiness signals. Two additional diagnostics bridge pooled prediction and panel inference. First, a complete-case temporal Extra Trees model removes numerical imputation. Second, country means are estimated from the training period and both outcomes and predictors are expressed as deviations from those training-period means before temporal testing. This country-demeaned machine-learning design measures within-country predictive variation without allowing time-invariant level differences to dominate. Country-level PCA and K-means clustering remain based on averages of readiness predictors only, while the fixed-effects benchmark includes one-year-lagged core predictors, country and year effects, and country-clustered standard errors. (See Supplementary Tables S6a, S6b, S7, S11, S12 and S13; Supplementary Figures S6, S7, S11, S12 and S13).

Figure 1 shows the distinction between readiness dimensions, controls, predictive assessment and the export-performance outcome. The excluded aggregates are stated below the framework to make the screening decision auditable.

Figure 1
A diagram illustrating the relationship between various readiness dimensions, controls, predictive assessment, and export performance.A diagram representing a framework for multidimensional supply-chain readiness and export performance. The diagram includes four main components: Digital readiness, Infrastructure readiness, Productive capacity, and External-sector conditions. Each of these components feeds into Interpretable machine learning, which includes methods such as Random validation, Temporal validation, Permutation importance, SHAP analysis, PCA, and K-means. The output of this process is Export performance, measured as exports of goods and services as a percentage of GDP. Digital readiness includes factors like Internet users and Mobile subscriptions. Infrastructure readiness includes factors like Electricity access and Population density. Productive capacity includes factors like Industry value added, Services value added, and GDP per capita. External-sector conditions include factors like Trade-facilitation measures, Transport-connectivity measures, Export-capacity measures, and Direct logistics measures.

Multidimensional supply-chain readiness and export-performance framework. Source: Authors' own construct (2026)

Figure 1
A diagram illustrating the relationship between various readiness dimensions, controls, predictive assessment, and export performance.A diagram representing a framework for multidimensional supply-chain readiness and export performance. The diagram includes four main components: Digital readiness, Infrastructure readiness, Productive capacity, and External-sector conditions. Each of these components feeds into Interpretable machine learning, which includes methods such as Random validation, Temporal validation, Permutation importance, SHAP analysis, PCA, and K-means. The output of this process is Export performance, measured as exports of goods and services as a percentage of GDP. Digital readiness includes factors like Internet users and Mobile subscriptions. Infrastructure readiness includes factors like Electricity access and Population density. Productive capacity includes factors like Industry value added, Services value added, and GDP per capita. External-sector conditions include factors like Trade-facilitation measures, Transport-connectivity measures, Export-capacity measures, and Direct logistics measures.

Multidimensional supply-chain readiness and export-performance framework. Source: Authors' own construct (2026)

Close Figure 1

All results are interpreted as predictive associations. Permutation and SHAP values indicate contributions to prediction, not the effect of a feasible policy intervention. Pooled machine-learning importance can reflect persistent differences across countries as well as changes over time, so the country-demeaned analysis is used to quantify the within-country component directly. The fixed-effects estimates and country-demeaned machine-learning results remain associational because neither lagging predictors nor removing country and year averages establishes exogenous variation.

Observed export performance varies across subregions and is not available for every possible country-year. Southern Africa has the highest mean export intensity in the observed sample, while West Africa has the largest number of observed outcomes. The unequal coverage reinforces the decision to report sample sizes and retain a missingness audit rather than constructing an artificial complete panel. (See Table 3).

Table 3

Subregional distribution of observed export performance

SubregionCountriesTarget grid rowsObserved outcomeMean export performance
Central Africa936029136.27
East Africa1872050128.54
North Africa624023528.47
Southern Africa520015040.88
West Africa1664055823.09
Source(s): Computations from the study dataset (2026)

Figure 2 shows persistent subregional differences and substantial year-to-year variation. The temporal structure supports the use of later-period validation rather than a single randomly mixed test sample. (See Supplementary Figure S1).

Figure 2
A line graph showing export-performance trends by African subregion over time.The line graph illustrates export-performance trends by African subregion from 1985 to 2025. The x-axis represents the years, ranging from 1985 to 2025, while the y-axis represents the percentage of exports of goods and services as a percentage of GDP, ranging from 20 to 50 percentage. The graph includes five data lines representing Central Africa, East Africa, North Africa, Southern Africa, and West Africa. Central Africa shows a significant decline from around 45 percentage in 1985 to around 30 percentage in 1990, followed by fluctuations and a peak around 2005. East Africa starts at around 20 percentage in 1985, with gradual increases and peaks around 2010. North Africa begins at around 20 percentage in 1985, with a steady increase and notable peaks around 2005 and 2015. Southern Africa starts at around 45 percentage in 1985, showing a decline to around 35 percentage by 1990, followed by fluctuations and peaks around 2005 and 2015. All values are approximated.

Export-performance trends by African subregion. Source: Computations from the study dataset (2026)

Figure 2
A line graph showing export-performance trends by African subregion over time.The line graph illustrates export-performance trends by African subregion from 1985 to 2025. The x-axis represents the years, ranging from 1985 to 2025, while the y-axis represents the percentage of exports of goods and services as a percentage of GDP, ranging from 20 to 50 percentage. The graph includes five data lines representing Central Africa, East Africa, North Africa, Southern Africa, and West Africa. Central Africa shows a significant decline from around 45 percentage in 1985 to around 30 percentage in 1990, followed by fluctuations and a peak around 2005. East Africa starts at around 20 percentage in 1985, with gradual increases and peaks around 2010. North Africa begins at around 20 percentage in 1985, with a steady increase and notable peaks around 2005 and 2015. Southern Africa starts at around 45 percentage in 1985, showing a decline to around 35 percentage by 1990, followed by fluctuations and peaks around 2005 and 2015. All values are approximated.

Export-performance trends by African subregion. Source: Computations from the study dataset (2026)

Close Figure 2

Extra Trees records the strongest core temporal performance, with MAE of 8.496, RMSE of 14.149 and R-squared of 0.651. Its random-split R-squared is 0.949, demonstrating that randomly mixing years and countries produce a much more optimistic assessment. The temporal result is therefore retained as the primary estimate of predictive performance. (See Table 4). (See Supplementary Figures S3 and S4).

Table 4

Core model performance under temporal and random validation

ValidationModelTraining NTest NMAERMSER-squared
Temporal holdoutExtra Trees1,3223898.49614.1490.651
Temporal holdoutRandom Forest1,32238910.17118.1100.429
Temporal holdoutGradient Boosting1,32238910.80918.9320.376
Temporal holdoutRidge1,32238914.21419.8680.312
Temporal holdoutMean benchmark1,32238915.73624.299−0.029
Random observation splitMean benchmark1,36834314.93921.859−0.001
Random observation splitRidge1,36834310.38516.6350.420
Random observation splitRandom Forest1,3683433.8745.7010.932
Random observation splitExtra Trees1,3683433.1974.9380.949
Random observation splitGradient Boosting1,3683438.23815.5690.492
Source(s): Computations from the study dataset (2026)

Figure 3 shows that the preferred model captures the central range of export intensity more accurately than extreme observations. This pattern explains why the temporal R-squared remains useful but materially lower than the random-split result.

Figure 3
A scatter plot showing the relationship between actual and predicted export performance.A scatter plot titled 'Actual versus predicted export performance Temporal holdout, Extra Trees' displays the relationship between actual export performance on the horizontal axis and predicted export performance on the vertical axis. The plot contains hundreds of data points, each representing a specific observation. A dashed line represents the ideal 1:1 relationship where actual and predicted values are equal. The data points are scattered around this line, with a noticeable concentration in the lower left quadrant, indicating a cluster of lower performance values. There is a general upward trend, suggesting a positive correlation between actual and predicted export performance. However, there are several outliers, particularly at higher performance values, where the predicted values deviate significantly from the actual values.

Actual versus predicted export performance under temporal validation. Source: Computations from the study dataset (2026)

Figure 3
A scatter plot showing the relationship between actual and predicted export performance.A scatter plot titled 'Actual versus predicted export performance Temporal holdout, Extra Trees' displays the relationship between actual export performance on the horizontal axis and predicted export performance on the vertical axis. The plot contains hundreds of data points, each representing a specific observation. A dashed line represents the ideal 1:1 relationship where actual and predicted values are equal. The data points are scattered around this line, with a noticeable concentration in the lower left quadrant, indicating a cluster of lower performance values. There is a general upward trend, suggesting a positive correlation between actual and predicted export performance. However, there are several outliers, particularly at higher performance values, where the predicted values deviate significantly from the actual values.

Actual versus predicted export performance under temporal validation. Source: Computations from the study dataset (2026)

Close Figure 3

Adding production-cost, trade-facilitation, transport and export-capacity indicators raises temporal R-squared from 0.651 to 0.667 and reduces RMSE from 14.149 to 13.845. The direct-logistics specification records R-squared of 0.809 and RMSE of 10.738, but it is based on 467 rows from intermittent measurement years. Its higher accuracy is therefore interpreted as evidence that direct logistics information is valuable, not as proof that the sparse specification is universally superior. (See Table 5).

Table 5

Temporal performance across core, extended and direct-logistics specifications

SpecificationRowsCountriesFirst yearLast yearCutoffBest modelR-squaredRMSE
Core annual readiness1,71151198520242017Extra Trees0.65114.149
Extended readiness1,71151198520242017Extra Trees0.66713.845
Direct logistics readiness46751200720222018Extra Trees0.80910.738
Source(s): Computations from the study dataset (2026)

Temporal errors are lowest in Southern Africa and West Africa, while East Africa records the highest RMSE. Bias also changes sign across subregions. These results show that pooled accuracy does not imply uniform performance and that regional structure remains important after conditioning on observed features. (See Table 6).

Table 6

Temporal prediction errors by African subregion

SubregionNMAEBiasRMSE
Central Africa7210.154−0.78514.972
East Africa12010.6302.46818.446
North Africa4610.1673.48016.407
Southern Africa404.906−1.6416.953
West Africa1115.7130.1737.703
Source(s): Computations from the study dataset (2026)

The overall permutation ranking is dominated by controls: log GDP per capita, log population, subregion and services value added. Among direct readiness indicators, industry employment ranks highest, followed by electricity access and private-sector credit. This separation matters because a variable can be highly predictive without representing a direct or feasible policy lever. (See Table 7).

Table 7

Top permutation contributions in the core temporal model

FeatureImportanceSD
GDP per capita, log0.36660.0395
Population, log0.19940.0335
African subregion0.13130.0136
Services value added0.11770.0121
Industry employment0.08990.0047
Natural-resource rents0.03910.0063
Island status0.02350.0076
Electricity access0.02180.0038
Private-sector credit0.01990.0035
Population density0.00960.0009
Source(s): Computations from the study dataset (2026)

The SHAP ranking similarly places log GDP per capita first, followed by log population, natural-resource rents, Southern Africa, services value added and electricity access. The dependence plots reveal nonlinear and interacting patterns. Higher income contributes increasingly positive predictions at the upper end of the distribution; larger population tends to contribute negatively after other features are considered; and resource rents are flat or negative through much of the range but become strongly positive among high-rent observations. The interaction colouring indicates that these patterns vary with income and population rather than operating independently. (See Figures 4–6). (See Supplementary Figures S9, S10a, S10b, S10c and S10d).

Figure 4
A bar graph showing SHAP feature importance for the preferred temporal model.The bar graph displays the SHAP feature importance for the preferred temporal model. The x-axis represents the mean absolute SHAP value, indicating the average impact on the model output magnitude. The y-axis lists the features in descending order of their importance. The features include log GDP per capita, log population, natural resource rents, subregion Southern Africa, services value added, electricity access, missing indicator population density, domestic credit private sector, subregion Central Africa, population density, subregion East Africa, industry employment, subregion North Africa, landlocked 0, and mobile subscriptions. Log GDP per capita has the highest importance, followed by log population and natural resource rents. The bars are horizontal and vary in length, reflecting the relative importance of each feature. The color scheme is uniform, with all bars in blue. All values are approximated.

SHAP feature importance for the preferred temporal model. Source: Computations from the study dataset (2026)

Figure 4
A bar graph showing SHAP feature importance for the preferred temporal model.The bar graph displays the SHAP feature importance for the preferred temporal model. The x-axis represents the mean absolute SHAP value, indicating the average impact on the model output magnitude. The y-axis lists the features in descending order of their importance. The features include log GDP per capita, log population, natural resource rents, subregion Southern Africa, services value added, electricity access, missing indicator population density, domestic credit private sector, subregion Central Africa, population density, subregion East Africa, industry employment, subregion North Africa, landlocked 0, and mobile subscriptions. Log GDP per capita has the highest importance, followed by log population and natural resource rents. The bars are horizontal and vary in length, reflecting the relative importance of each feature. The color scheme is uniform, with all bars in blue. All values are approximated.

SHAP feature importance for the preferred temporal model. Source: Computations from the study dataset (2026)

Close Figure 4
Figure 5
A scatter plot showing the relationship between log GDP per capita and SHAP values.A scatter plot titled 'SHAP dependence and interaction: log_gdp_per_capita' displays the relationship between log GDP per capita on the horizontal axis and SHAP values for log GDP per capita on the vertical axis. The plot includes dozens of data points, each colored according to the log population value, with a color scale ranging from blue to red. The data points show a positive correlation between log GDP per capita and SHAP values, with higher log GDP per capita values generally associated with higher SHAP values. There is a noticeable cluster of points around the lower end of the log GDP per capita range, and a few outliers at the higher end. The color gradient indicates that larger populations tend to have lower SHAP values after accounting for other features.

SHAP dependence and interaction pattern for GDP per capita. Source: Computations from the study dataset (2026)

Figure 5
A scatter plot showing the relationship between log GDP per capita and SHAP values.A scatter plot titled 'SHAP dependence and interaction: log_gdp_per_capita' displays the relationship between log GDP per capita on the horizontal axis and SHAP values for log GDP per capita on the vertical axis. The plot includes dozens of data points, each colored according to the log population value, with a color scale ranging from blue to red. The data points show a positive correlation between log GDP per capita and SHAP values, with higher log GDP per capita values generally associated with higher SHAP values. There is a noticeable cluster of points around the lower end of the log GDP per capita range, and a few outliers at the higher end. The color gradient indicates that larger populations tend to have lower SHAP values after accounting for other features.

SHAP dependence and interaction pattern for GDP per capita. Source: Computations from the study dataset (2026)

Close Figure 5
Figure 6
A scatter plot showing SHAP dependence and interaction for log population.A scatter plot titled SHAP dependence and interaction: log population. The x-axis represents log population values ranging from approximately -2 to 2. The y-axis represents SHAP values for log population, ranging from approximately -6 to 6. The data points are color-coded based on log GDP per capita, with a color scale ranging from blue to red. The plot shows several data points scattered across the graph, with a noticeable cluster of points around the log population value of 0 and SHAP value of 0. There are also several outliers, particularly at the higher and lower ends of the SHAP value range. The color gradient indicates that higher log GDP per capita values tend to be associated with higher SHAP values for log population, while lower log GDP per capita values tend to be associated with lower SHAP values for log population. The interaction coloring suggests that the patterns vary with income and population rather than operating independently. All values are approximated.

SHAP dependence and interaction pattern for population. Source: Computations from the study dataset (2026)

Figure 6
A scatter plot showing SHAP dependence and interaction for log population.A scatter plot titled SHAP dependence and interaction: log population. The x-axis represents log population values ranging from approximately -2 to 2. The y-axis represents SHAP values for log population, ranging from approximately -6 to 6. The data points are color-coded based on log GDP per capita, with a color scale ranging from blue to red. The plot shows several data points scattered across the graph, with a noticeable cluster of points around the log population value of 0 and SHAP value of 0. There are also several outliers, particularly at the higher and lower ends of the SHAP value range. The color gradient indicates that higher log GDP per capita values tend to be associated with higher SHAP values for log population, while lower log GDP per capita values tend to be associated with lower SHAP values for log population. The interaction coloring suggests that the patterns vary with income and population rather than operating independently. All values are approximated.

SHAP dependence and interaction pattern for population. Source: Computations from the study dataset (2026)

Close Figure 6

The country-and-year fixed-effects benchmark uses 798 complete observations from 46 economies. Lagged inflation and services value added are negatively associated with export intensity, while natural-resource rents are positively associated. Most direct readiness variables are statistically imprecise in this smaller complete-case sample. The difference from the machine-learning ranking illustrates that out-of-sample predictive contribution and within-country conditional association answer different questions. (See Table 8).

Table 8

Country and year fixed-effects benchmark with lagged core predictors

VariableCoefficientStandard errorp-value
Internet use, lagged−0.0950.0800.234
Mobile subscriptions, lagged0.0020.0310.937
Electricity access, lagged−0.0130.0810.876
Manufacturing value added, lagged−0.1100.1590.490
Industry employment, lagged0.3570.2180.101
Private-sector credit, lagged0.0460.0770.549
Inflation, lagged−0.0360.007<0.001
Log GDP per capita−0.1962.7130.942
Log population−18.69510.2100.067
Population density0.0300.0470.525
Services value added−0.4020.1340.003
FDI inflows, lagged0.0190.0540.727
Natural-resource rents0.5600.165<0.001

Note(s): Country and year effects are included; standard errors are clustered by country. The estimates are associational

A three-cluster solution has the highest silhouette score (0.421). The first three principal components explain 81.89% of the variance in the readiness indicators. Cluster 0 contains 15 higher-readiness economies; Cluster 1 contains 38 lower-readiness economies; and Cluster 2 contains one extreme case, the Democratic Republic of the Congo, whose profile is dominated by very high average inflation. The third cluster is therefore treated as an outlier profile rather than a general readiness regime. (See Table 9 and Figure 7). (See Supplementary Tables S6a, S6b and S7; Supplementary Figure S7).

Table 9

Mean country profile by readiness cluster

ClusterCountriesInternetElectricityManufacturing VAIndustry employmentPrivate creditInflationExport performance
01526.5677.2512.8722.1239.916.4638.00
1388.6731.049.5111.0612.6222.8827.97
214.7314.4612.448.393.551094.5429.00
Source(s): Computations from the study dataset (2026)
Figure 7
A scatter plot showing country-level supply-chain readiness clusters.A scatter plot titled 'Country-level supply-chain readiness clusters' displays the relationship between the first and second principal components of supply-chain readiness indicators. The horizontal axis represents the first principal component, and the vertical axis represents the second principal component. The plot includes dozens of data points, each representing a country. The data points are color-coded based on their readiness cluster, with a color bar on the right indicating the cluster values from 0 to 2. The plot shows several clusters, patterns, and an outlier. The outlier, labeled COD, is positioned far from the main clusters. The main clusters are grouped around the center of the plot, with some countries showing higher readiness scores and others lower. The plot does not include any regression or trend lines.

Country-level PCA map of supply-chain readiness clusters. Source: Computations from the study dataset (2026)

Figure 7
A scatter plot showing country-level supply-chain readiness clusters.A scatter plot titled 'Country-level supply-chain readiness clusters' displays the relationship between the first and second principal components of supply-chain readiness indicators. The horizontal axis represents the first principal component, and the vertical axis represents the second principal component. The plot includes dozens of data points, each representing a country. The data points are color-coded based on their readiness cluster, with a color bar on the right indicating the cluster values from 0 to 2. The plot shows several clusters, patterns, and an outlier. The outlier, labeled COD, is positioned far from the main clusters. The main clusters are grouped around the center of the plot, with some countries showing higher readiness scores and others lower. The plot does not include any regression or trend lines.

Country-level PCA map of supply-chain readiness clusters. Source: Computations from the study dataset (2026)

Close Figure 7

The alternative outcomes show that the predictive pattern is not confined to exports as a share of GDP. Extra Trees records R-squared of 0.926 for log merchandise export value, 0.891 for the number of exported HS-6 products and 0.824 for log export value per HS-6 product. These results cover aggregate export scale, product breadth and product depth, respectively. The destination-partner measures are not estimated because the retrieved partner series did not provide usable country-year coverage and were not reconstructed through imputation (See Table 10 and Figure 8). (See Supplementary Tables S8 and S9; Supplementary Figures S11, S12 and S13).

Table 10

Alternative export-performance outcomes under temporal validation

OutcomeObserved NCountriesPeriodBest modelR-squaredRMSE
Aggregate value — log merchandise exports1,962541988–2024Extra Trees0.9260.507
Extensive margin — export partners–––Not estimated––
Extensive margin — exported HS6 products1,185491990–2023Extra Trees0.891334.676
Intensive margin — log exports per partner–––Not estimated––
Intensive margin — log exports per HS6 product1,185491990–2023Extra Trees0.8240.629

Note(s): Partner-based breadth and depth outcomes did not meet usable coverage and were not filled by imputation

Source(s): Computations from the study dataset (2026)
Figure 8
A bar graph showing temporal predictive performance across export-performance margins.The bar graph compares the temporal predictive performance across three different export-performance margins. It features three vertical bars. The x-axis is labeled 'Outcome definition' and includes three categories: 'Aggregate value - log merchandise exports', 'Extensive margin - exported HS6 products', and 'Intensive margin - log exports per HS6 product'. The y-axis is labeled 'R-squared' and ranges from 0.0 to 1.0. The bars are colored blue. The first bar represents the R-squared value for log merchandise export value, the second bar represents the R-squared value for the number of exported HS6 products, and the third bar represents the R-squared value for log export value per HS6 product. The values are approximately 0.926, 0.891, and 0.824, respectively. The graph indicates that the predictive pattern is consistent across different export-performance margins. All values are approximated.

Temporal predictive performance across estimable export-performance margins. Source: Computations from the study dataset (2026)

Figure 8
A bar graph showing temporal predictive performance across export-performance margins.The bar graph compares the temporal predictive performance across three different export-performance margins. It features three vertical bars. The x-axis is labeled 'Outcome definition' and includes three categories: 'Aggregate value - log merchandise exports', 'Extensive margin - exported HS6 products', and 'Intensive margin - log exports per HS6 product'. The y-axis is labeled 'R-squared' and ranges from 0.0 to 1.0. The bars are colored blue. The first bar represents the R-squared value for log merchandise export value, the second bar represents the R-squared value for the number of exported HS6 products, and the third bar represents the R-squared value for log export value per HS6 product. The values are approximately 0.926, 0.891, and 0.824, respectively. The graph indicates that the predictive pattern is consistent across different export-performance margins. All values are approximated.

Temporal predictive performance across estimable export-performance margins. Source: Computations from the study dataset (2026)

Close Figure 8

The missing-data sensitivity test also changes the interpretation of the benchmark. The complete-case Extra Trees model, estimated on 630 training observations and 202 later-period observations with no numerical imputation, records MAE = 6.869, RMSE = 11.677 and R-squared = 0.737. This is higher than the pooled benchmark R-squared of 0.651, indicating that the benchmark signal is not created by median imputation. The result does not make sparse variables harmless; rather, it supports the 50% eligibility rule used for the bridge analyses and the separation of sparse logistics indicators into dedicated period-specific tests. (See Supplementary Tables S10 and S11).

Country demeaning provides a direct test of whether pooled machine learning is mainly learning persistent country differences. R-squared falls from 0.651 to 0.242 for export intensity and from 0.926 to 0.283 for log merchandise export value. The decline confirms that a substantial share of pooled predictive accuracy is between-country structural information. Importantly, the within-country rankings are not empty. Industry employment is the largest readiness contribution for within-country export intensity (permutation importance = 0.270). For within-country log merchandise exports, manufacturing value added is the leading readiness contribution (0.104), followed by mobile subscriptions (0.028), industry employment (0.027), private-sector credit (0.017) and internet use (0.015). (See Table 11). (See Supplementary Tables S11, S12 and S13).

Table 11

Pooled, complete-case and within-country predictive diagnostics

Outcome/diagnosticModel scopeR-squaredRMSE
Export intensity benchmarkPooled levels ML0.65114.149
Export intensity benchmarkComplete case; no numerical imputation0.73711.677
Export intensity benchmarkWithin-country ML0.24210.510
Log merchandise export valuePooled levels ML0.9260.507
Log merchandise export valueWithin-country ML0.2830.627
Source(s): Computations from the study dataset (2026)

The first finding is conceptual in the sense that export intensity can be predicted from a combination of readiness indicators and structural controls, but it should not be labelled export competitiveness. Renaming the outcome clarifies that the analysis concerns the share of domestic activity connected to exports rather than relative product advantage, sophistication or complexity.

The second finding concerns outcome and variable design. Removing imports, current account balance and the trade aggregate eliminates the principal source of accounting and trade-openness overlap, while the alternative outcomes show that the empirical story is not dependent on the exports-to-GDP ratio. Strong temporal performance for log merchandise export value, exported HS-6 product breadth and export value per HS-6 product indicates that readiness and structural conditions contain information about export scale, breadth and depth. The inability to estimate the partner-based margins is treated as a data-coverage limitation rather than resolved through imputation. The broader outcome design is consistent with evidence that African export performance reflects structural, trade and policy conditions beyond a single openness ratio (Vogel, 2025).

The third finding is that the largest predictive contributions come from controls. Income, population, subregion, services and resource dependence capture broad heterogeneity in productive structure and export composition. Their importance should not be translated directly into a policy ranking. Among direct readiness measures, industry employment, electricity access and private credit are more defensible intervention-related signals, although their permutation contributions are smaller.

The SHAP results add information beyond ranking. The income pattern is nonlinear, the population pattern is generally negative after other conditions are held constant, and high resource rents are associated with unusually positive export-intensity predictions. The resource-rent pattern is consistent with large commodity-export shares but does not imply diversified or sophisticated export structures. This distinction helps reconcile high export intensity with continuing concerns about concentration and structural transformation in Africa (Avenyo et al., 2021; Vogel, 2025).

The fixed-effects and country-demeaned results provide a stronger caution than a simple distinction between prediction and inference. The pronounced fall in R-squared after country demeaning confirms that pooled machine learning learns substantial between-country structure. That structural component is useful for cross-country prediction but cannot be interpreted as evidence that changing a high-ranked level variable within a country will reproduce the pooled ranking. At the same time, industry employment remains the strongest readiness signal for within-country export intensity, while manufacturing value added leads the readiness variables for within-country merchandise export value. The two approaches are therefore complementary only when the between-country component is made explicit.

Policy interpretation should begin with mechanisms rather than the pooled feature ranking. Electricity access affects production continuity, cold-chain reliability and the ability to meet delivery schedules. Industrial employment and manufacturing value added proxy the depth of productive capabilities and supplier networks that allow more products to reach exportable scale. Private-sector credit supports working capital, certification, inventory and shipment finance, while digital connectivity reduces buyer-search, coordination, documentation and tracking costs. Border time, shipping connectivity and logistics competence operate through lead-time reliability and the amount of capital tied up while goods move across borders. These pathways link readiness to both export breadth and export depth.

The within-country diagnostics narrow the actionable interpretation. Industry employment is the strongest readiness signal for within-country movements in export intensity, and manufacturing value added is the strongest readiness signal for within-country changes in log merchandise exports. These results suggest that governments should first diagnose whether the binding constraint is productive capacity, finance, energy, digital coordination or border/logistics reliability, and then compare interventions within that constraint. A country with repeated power interruptions may obtain little from advanced tracking systems until production reliability improves; a country with adequate production but high border delays may obtain more from customs process reform and corridor coordination.

The machine-learning results do not determine budget allocations on their own. A practical prioritisation process should combine the within-country readiness signals with intervention cost, implementation time, corridor exposure, distributional effects and credible causal evidence from pilots or quasi-experimental evaluation. Structural controls such as GDP per capita, population, subregion and island status remain useful for identifying context, but they are not investment levers. This converts explainable machine learning from a causal ranking into a screening tool for selecting bottlenecks that warrant economic appraisal.

This study examines supply-chain readiness and export performance across African economies using an interpretable predictive framework. Export intensity is retained as a benchmark, while log merchandise export value, exported HS-6 product breadth and log export value per HS-6 product provide alternative measures of aggregate scale, extensive product breadth and intensive product depth. Trade, imports and current account balance remain excluded from the main specifications, and direct readiness measures are separated from structural controls.

Extra Trees provides the strongest benchmark temporal result, with R-squared of 0.651. The alternative outcome models record R-squared of 0.926 for log merchandise exports, 0.891 for exported HS-6 product breadth and 0.824 for log exports per HS-6 product. A complete-case benchmark without numerical imputation records R-squared of 0.737. Country demeaning lowers predictive performance to 0.242 for export intensity and 0.283 for log merchandise exports, demonstrating that pooled results contain substantial between-country structure. Within-country importance nevertheless identifies industry employment and manufacturing value added as the leading readiness signals for the two principal outcomes.

These limitations provide a clear direction for future research. The destination-partner breadth and per-partner depth measures did not provide a usable country-year series in the retrieved WITS indicators and should be reconstructed directly from bilateral Comtrade records in future work. Product-level measures such as revealed comparative advantage, EXPY, economic complexity, export diversification and constant-market-share performance can further distinguish export quality from scale. Shipment-level logistics measures, structural-break tests, alternative missing-data strategies, firm and corridor evidence, and credible causal designs are also required before estimating the return to specific customs, infrastructure, energy or digital interventions.

The authors gratefully acknowledge the academic and institutional support provided by the University of Cape Coast and Yanshan University. The authors also appreciate the constructive comments and guidance received during the development and review of this study. Finally, appreciation is extended to the institutions and data providers whose publicly available datasets made the empirical analysis possible.

The supplementary material for this article can be found online.

Antràs
,
P.
and
Chor
,
D.
(
2022
), “Global value chains”, in
Handbook of International Economics
, Vol. 
5
, pp. 
297
-
376
, doi: .
Avenyo
,
E.K.
,
Tregenna
,
F.
and
Kraemer-Mbula
,
E.
(
2021
), “
Do productive capabilities affect export performance? Evidence from African firms
”,
European Journal of Development Research
, Vol. 
33
No. 
2
, pp. 
304
-
329
, doi: .
Bugarčić
,
F.Ž.
,
Stanišić
,
N.
and
Marinković
,
V.
(
2024
), “
Assessing the effects of logistics performance on export and competitiveness using SEM methodology: evidence from firm-level data
”,
International Journal of Logistics Management
, Vol. 
35
No. 
6
, pp. 
1847
-
1866
, doi: .
Castillo
,
C.
(
2023
), “
Is there a theory of supply chain resilience? A bibliometric analysis of the literature
”,
International Journal of Operations and Production Management
, Vol. 
43
No. 
1
, pp. 
22
-
47
, doi: .
Goulet Coulombe
,
P.
(
2024
), “
The macroeconomy as a random forest
”,
Journal of Applied Econometrics
, Vol. 
39
No. 
3
, pp. 
401
-
421
, doi: .
Govindan
,
K.
,
Kannan
,
D.
,
Jørgensen
,
T.B.
and
Nielsen
,
T.S.
(
2022
), “
Supply Chain 4.0 performance measurement: a systematic literature review, framework development, and empirical evidence
”,
Transportation Research Part E: Logistics and Transportation Review
, Vol. 
164
, 102725, doi: .
Ivanov
,
D.
and
Dolgui
,
A.
(
2020
), “
Viability of intertwined supply networks: extending the supply chain resilience angles towards survivability. A position paper motivated by COVID-19 outbreak
”,
International Journal of Production Research
, Vol. 
58
No. 
10
, pp. 
2904
-
2915
, doi: .
Jha
,
A.
,
Sharma
,
R.R.K.
,
Kumar
,
V.
and
Verma
,
P.
(
2022
), “
Designing supply chain performance system: a strategic study on Indian manufacturing sector
”,
Supply Chain Management: An International Journal
, Vol. 
27
No. 
1
, pp. 
66
-
88
, doi: .
Kere
,
S.
and
Zongo
,
A.
(
2023
), “
Digital technologies and intra-African trade
”,
International Economics
, Vol. 
173
, pp. 
359
-
383
, doi: .
Kuteyi
,
D.
and
Winkler
,
H.
(
2022
), “
Logistics challenges in Sub-Saharan Africa and opportunities for digitalization
”,
Sustainability
, Vol. 
14
No. 
4
, p.
2399
, doi: .
Larch
,
M.
and
Yotov
,
Y.V.
(
2024
), “
Estimating the effects of trade agreements: lessons from 60 years of methods and data
”,
The World Economy
, Vol. 
47
No. 
5
, pp. 
1771
-
1799
, doi: .
Masini
,
R.P.
,
Medeiros
,
M.C.
and
Mendes
,
E.F.
(
2023
), “
Machine learning advances for time series forecasting
”,
Journal of Economic Surveys
, Vol. 
37
No. 
1
, pp. 
76
-
111
, doi: .
Melitz
,
M.J.
and
Redding
,
S.J.
(
2021
),
Trade and Innovation (No. W28945)
,
National Bureau of Economic Research
, doi: .
Olyanga
,
A.M.
,
Shinyekwa
,
I.M.
,
Ngoma
,
M.
,
Nkote
,
I.N.
,
Esemu
,
T.
and
Kamya
,
M.
(
2022
), “
Export logistics infrastructure and export competitiveness in the East African Community
”,
Modern Supply Chain Research and Applications
, Vol. 
4
No. 
1
, pp. 
39
-
61
, doi: .
Tandrayen-Ragoobur
,
V.
,
Ongono
,
P.
and
Gong
,
J.
(
2023
), “
Infrastructure and intra-regional trade in Africa
”,
The World Economy
, Vol. 
46
No. 
2
, pp. 
453
-
471
, doi: .
Vogel
,
T.
(
2025
), “
Combining the pieces: identifying key determinants of export diversification in Africa amidst model uncertainty
”,
Review of World Economics
, Vol. 
161
No. 
1
, pp. 
257
-
307
, doi: .
Wassie
,
M.A.
,
Kornher
,
L.
and
Zaki
,
C.
(
2025
), “
Revisiting the impact of trade facilitation measures in Africa: a structural gravity approach
”,
Journal of International Trade and Economic Development
, Vol. 
34
No. 
7
, pp. 
1604
-
1634
, doi: .
World Bank
(
2023
),
Connecting to Compete 2023: Trade Logistics in an Uncertain Global Economy: The Logistics Performance Index and its Indicators
,
World Bank
.
Xing
,
L.
,
Cao
,
C.
and
Elahi
,
E.
(
2026
), “
Does digital technology development promote export efficiency?
”,
Review of Development Economics
, Vol. 
30
No. 
2
, pp. 
982
-
996
, doi: .
Antràs
,
P.
(
2020
), “
Conceptual aspects of global value chains
”,
The World Bank Economic Review
, Vol. 
34
No. 
3
, pp. 
551
-
574
, doi: .
Arkhangelsky
,
D.
,
Athey
,
S.
,
Hirshberg
,
D.A.
,
Imbens
,
G.W.
and
Wager
,
S.
(
2021
), “
Synthetic difference-in-differences
”,
The American Economic Review
, Vol. 
111
No. 
12
, pp. 
4088
-
4118
, doi: .
Dawson-Amoah
,
S.
and
Ekow de-Graft
,
O.
(
2026
), “
Replication data and code for ‘supply chain readiness and export performance in Africa: evidence from interpretable machine learning’
”,
Mendeley Data
, Vol. 
V2
, doi: .
Gollin
,
D.
and
Kaboski
,
J.P.
(
2023
), “
New views of structural transformation: insights from recent literature
”,
Oxford Development Studies
, Vol. 
51
No. 
4
, pp. 
339
-
361
, doi: .
Published in Journal of International Logistics and Trade. Published by Emerald Publishing Limited. This article is published under the Creative Commons Attribution (CC BY 4.0) licence. Anyone may reproduce, distribute, translate and create derivative works of this article (for both commercial and non-commercial purposes), subject to full attribution to the original publication and authors. The full terms of this licence may be seen at Link to the terms of the CC BY 4.0 licence.

Supplementary data

or Create an Account

Close subscription notice
Close access options