This study constructs a fully balanced panel dataset for 135 countries spanning 2013–2022 to explore the determinants of international trade. It employs classical econometric techniques – Robust Least Squares (RLS), Generalized Linear Model (GLM) and quantile regression – to capture linear effects, heterogeneity and distributional nuances. Complementing these, advanced Machine Learning (ML) methods – including Gradient Boosting Machine (GBM), bagging via Random Forest and an ensemble stacking model – uncover nonlinear relationships and complex interactions. All numeric variables are scaled, and a training/testing split is implemented, ensuring robust performance evaluation through metrics such as MAE, MSE, RMSE and R2.
Advanced ML techniques are utilized extensively for both regression and robustness checks. For regression, ML methods such as bagging via Random Forest, boosting and stacking with a meta-learner are employed.
Empirical evidence from both econometric and ML analyses reveals that a strong business environment (BE), high-tech exports (HTE), robust ICT services imports (ICTSI) and widespread ICT use (ICTU) significantly promote trade intensity across 135 countries from 2013 to 2022. Quantile regressions indicate that HTE’s positive impact intensifies at higher trade quantiles, whereas persistent underinvestment in R&D (RDC) consistently hampers trade performance. Advanced ML models, particularly GBM and ensemble stacking, further capture nonlinearities and interactions, reinforcing these findings and underscoring the critical role of digital infrastructure and innovation ecosystems in driving global trade competitiveness.
This study uniquely bridges classical econometrics with state-of-the-art ML to examine the trade–innovation nexus. It harnesses a fully balanced panel of 135 countries (2013–2022) and employs RLS, GLM, quantile regression, alongside advanced ML techniques like gradient boosting, bagging via Random Forest and stacking ensembles. This dual approach not only captures both linear and nonlinear dynamics but also enhances predictive accuracy and model interpretability. The integration of these methods sets a novel benchmark, offering robust, data-driven insights and context-specific policy recommendations that enrich the literature on global trade patterns amid rapid technological advancement.
1. Introduction
Understanding how international trade and innovation interact is central to addressing the dual imperatives of competitiveness and sustainability in the global economy. Empirical evidence shows that Foreign Direct Investment (FDI) and trade can generate innovation spillovers for domestic firms, but these effects are primarily concentrated at the firm level and do not extend broadly across entire industries (Gorodnichenko et al., 2020; Magazzino and Mele, 2022). The mechanisms linking trade and innovation are multifaceted and include the influence of trade on market structure, competition, and knowledge flows, each of which can amplify or constrain innovation incentives depending on context (Melitz and Redding, 2021). Given this complexity, conventional linear models may fail to capture the heterogeneity and interaction effects that characterize these relationships, calling for the integration of Machine Learning (ML) techniques capable of modeling nonlinearities, variable interactions, and higher-order dependencies that traditional econometric approaches may overlook.
This need is particularly salient in light of growing evidence that the impact of trade on innovation is highly uneven across sectors and countries. Trade openness alone does not guarantee the spread of innovation; rather, the diffusion of knowledge depends on sector-specific research and development (R&D) productivity, as well as cross-country and cross-sector heterogeneity in innovation efficiency (Cai et al., 2022). Innovation gains from Global Value Chain (GVC) participation are contingent upon foreign knowledge embedded in imported intermediate goods, and the strength of these effects is conditioned by intellectual property regimes, trade rules, and competition policy (Eissa and Zaki, 2023). Regional trade integration may reduce technological gaps – as observed in the European Union – but benefits accrue unevenly, favoring sectors and countries with stronger innovation capacities (Antimiani and Costantini, 2013).
In South Asia, green economic growth is positively associated with clean energy production, green innovation, and environmentally sustainable trade, indicating the relevance of coordinated policy strategies that link trade liberalization with energy and environmental transitions (Ahmed et al., 2022). These dynamics are not uniform across income groups: developing economies benefit more from innovation inputs and energy-driven exports, while advanced economies reap greater gains from innovation outputs and exchange rate competitiveness (Asghar et al., 2024). Causal links between trade and innovation also vary by sector, with imports typically stimulating innovation, which then drives exports, suggesting complex feedback loops rather than unidirectional effects (Johnson and Van Wagoner, 2021).
The expansion of digital infrastructure and e-commerce across Europe further reflects the tight coupling between innovation and trade performance. Higher innovation intensity is positively associated with better outcomes in business-to-consumer (B2C), business-to-business (B2B), and business-to-government (B2G) e-commerce sectors, which in turn support national economic development (Skare et al., 2023). However, these benefits are conditioned by the digital maturity of firms and countries. Similarly, the environmental costs of trade and Information and Communication Technology (ICT) expansion necessitate meticulous oversight, as trade openness and innovation can both reduce or exacerbate emissions depending on the energy mix, regulatory environment, and technology use (e.g. Haldar and Sethi, 2022; Zhangqi et al., 2022; Obobisa, 2023).
The strategic role of environmental R&D is also evident in efforts to mitigate trade-adjusted carbon emissions, particularly in developed economies. Increased investment in environmental innovation is associated with lower consumption-based carbon dioxide (CO2) emissions in G7 countries (Khan et al., 2020), while trade volumes and Gross Domestic Product (GDP) tend to raise emissions unless counterbalanced by regulatory action (Jiang et al., 2022). In the Organization for Economic Co-operation and Development (OECD) countries, green finance and green innovation are shown to reduce trade-adjusted emissions, particularly in high-emission contexts (Umar and Safi, 2023). However, even inclusive trade agendas risk reproducing structural inequalities unless redistribution, representation, and recognition are addressed in tandem (Goff, 2021).
In light of these challenges and following recent literature (see Soylu et al., 2023) on the ICT-trade nexus on competitiveness, this work explores how trade openness interacts with innovation-related drivers, such as high-tech transfers, ICT infrastructure, and R&D capacity. It integrates econometric and ML methods to uncover both linear and nonlinear effects, accounting for heterogeneity across development levels. In doing so, it adds to ongoing efforts to design trade and innovation strategies that are both evidence-based and context-sensitive.
This study contributes to the literature in two key ways: methodologically, by integrating supervised ML algorithms with classical econometric techniques; and empirically, by uncovering new insights into the nonlinear relationships among trade openness, innovation, and sustainability indicators. Departing from the dominant use of standard panel techniques, we rely instead on Robust Least Squares (RLS), Generalized Linear Model (GLM), and quantile regression to estimate structural associations across countries, capturing heterogeneity across the distribution of trade integration. To complement and extend these econometric findings, we employ a suite of supervised ML techniques – namely, Gradient Boosting Machine (GBM), bagging, and a stacked ensemble learner. These models are used to enhance predictive accuracy and to identify high-dimensional interaction effects that standard regression techniques may overlook. The analysis draws on data from the Global Innovation Index (GII) for 135 countries from 2013 to 2022, focusing on the role of Business Environment (BE), High-Tech Exports (HTE), ICT Services Imports (ICTSI), ICT Use (ICTU), and R&D intensity in shaping export competitiveness. By leveraging the strengths of both econometric and ML approaches, this study offers a more robust understanding of how innovation capabilities and trade performance co-evolve, yielding novel insights for context-specific policy design.
2. Literature review
Recent research emphasizes the ever-evolving dynamics between trade, innovation, and sustainability, showcasing both opportunities and constraints across diverse economic contexts. Innovation – particularly when supported by targeted R&D investment – emerges as a critical determinant of international competitiveness and long-run productivity growth. As Fernández (2023) shows in a systematic literature review, innovation plays a central part in determining the strategies of multinational enterprises and national innovation systems, especially through cross-border knowledge flows and integration into GVCs.
Trade, in turn, acts as a conduit for technology transfer and knowledge diffusion, but its effects are mediated by structural and policy factors. For instance, Lee (2020) finds that R&D investment and intermediate input trade with technologically advanced partners promote productivity growth in non-frontier economies, especially when domestic absorptive capacities are strong. Similarly, in the context of China’s Belt and Road Initiative (BRI), Li et al. (2022) demonstrate that expanded foreign trade under the BRI affects firms’ sustainable innovation capabilities, primarily via outward engagement (“going-out” effect), rather than through inward technology adoption.
Innovation is also increasingly viewed through the lens of environmental sustainability. Ali et al. (2021) provide evidence from the world’s top ten emitters that environmental innovation, trade, and renewable energy consumption are long-run determinants of both consumption-based and territorial carbon emissions. In the Group of Seven (G7) countries, Khan et al. (2020) confirm that while imports and income increase consumption-based emissions, exports, environmental innovation, and renewable energy use exert mitigating effects. Liu et al. (2023) reinforce this pattern, showing that green trade and green technological innovation reduce ecological footprints in South Asian economies.
However, these outcomes are not guaranteed. Usman et al. (2020), examining the US, find that higher renewable energy consumption reduces environmental degradation by lowering the ecological footprint, while trade policy also exerts a mitigation effect. Similarly, Wahab et al. (2021) reveal that energy productivity and innovation reduce consumption-based emissions in the G7, while imports and GDP raise them. These findings align with those of Essandoh et al. (2020), who note that FDI and trade tend to displace emission-intensive activities toward developing countries, enabling emission reductions in the former while heightening environmental pressures in the latter.
Institutional factors, such as fiscal decentralization and higher education, also influence the trade–innovation–sustainability nexus. In China, Ma (2024) finds that the large-scale expansion of college education significantly boosted firm-level R&D and export upgrading. Likewise, Cheng et al. (2021) report that innovation and decentralized fiscal governance jointly reduce CO2 emissions by enabling localized environmental governance. In the context of policy experimentation, Lei and Xie (2024) show that China’s Free Trade Zones (FTZs) fostered firm-level innovation, especially in technology-intensive and state-owned enterprises, by reducing financial constraints and enhancing policy responsiveness.
The role of SMEs in innovation-driven trade is equally prominent. Hilmersson et al. (2023) show that a fast pace of technological innovation accelerates international expansion among Swedish small and medium enterprises (SMEs), while Idris et al. (2022) find that simultaneous inward and outward internationalization significantly increases both product and process innovation in UK SMEs. Smallbone et al. (2022) further highlight that SMEs in developing countries benefit more from foreign technology licensing and export activities, although these effects are shaped by the country’s level of economic development.
Digital transformation also plays a growing role in reshaping trade and innovation strategies. According to Singh and Siddiqui (2023), ICT penetration influences the link between trade and economic growth in developing countries, particularly when paired with innovation-led strategies. Azmeh et al. (2020) argue that digital trade is outpacing current governance structures, calling for international trade rules to adapt to new patterns of cross-border data flows and digital services. Yet, without environmental safeguards, digital expansion may exacerbate emissions, as shown by Zameer et al. (2020) in India, where innovation and FDI help alleviate the negative environmental effects of trade and growth.
The strategic alignment of trade and environmental policy is further illustrated in studies on energy and resource efficiency. Wen et al. (2022) demonstrate that renewable energy and energy efficiency significantly enhance technological innovation, particularly when mediated by trade, investment, and human capital development. Su et al. (2023) add that in Brazil, Russia, India, China, and South Africa (BRICS), digitalization improves sustainability in the natural resource sector, while oil rents still pose challenges for sustainable trade. Within the OECD, Osabuohien-Irabor and Drapkin (2022) find that trade openness and FDI moderate the energy-saving effects of technological innovation, following an inverted U-shape relationship over time.
Finally, the shift toward circular economy models introduces new frameworks for reconciling trade with environmental objectives. Polyakov et al. (2023) propose integrated production systems, such as flexible manufacturing and distributed production, that align with trade in secondary raw materials and R&D services, thereby reinforcing circularity in international trade.
Despite the breadth of research on trade, innovation, and environmental sustainability, methodological approaches remain largely conventional. Most studies rely on linear econometric models and overlook advanced tools that could uncover subtler empirical regularities. In particular, there is a notable absence of ML applications in this domain, even as such techniques have shown strong potential in related areas like economic forecasting, environmental modeling, and digital trade analysis. While existing evidence provides important insights into structural relationships, less attention has been paid to predictive performance, model generalizability, or the ability to detect higher-order interactions. This study addresses that shortcoming by incorporating ML methods, alongside standard econometric models, into a unified analytical framework. In doing so, it opens new pathways for identifying data-driven patterns and offers an enhanced empirical lens on how innovation and sustainability interact with trade across diverse national contexts.
3. Data and empirical framework
To examine the relationship between trade levels and technological innovation, we construct a panel dataset covering 135 countries over the period 2013–2022. The data, drawn from the World Bank [1] and from World Intellectual Property Organization (WIPO)’s GII [2] databases, include a consistent and fully balanced panel across all variables and years (see Table 1).
Variables’ description
| Variable | Acronym | Definition | Source | Unit |
|---|---|---|---|---|
| Trade as a percentage of GDP | Trade | Measures the relative importance of international trade in a country’s economy. It is calculated by summing the value of a nation’s exports and imports of goods and services, then dividing this total by its GDP | World Bank | % |
| Business environment | BE | Captures institutional, legal, and macroeconomic conditions affecting firms, including infrastructure, regulatory quality, and political stability | Global Innovation Index | Composite Index |
| High-tech exports, % total trade | HTE | Share of a country’s total exports classified as high-technology products, including aerospace, biotechnology, pharmaceuticals, and ICT goods | Global Innovation Index | % of total trade |
| ICT services imports, % total trade | ICTSI | Measures the share of trade accounted for by imported ICT services such as software development, telecommunications, and data processing | Global Innovation Index | % of total trade |
| ICT use | ICTU | Composite index reflecting the degree of ICT diffusion across households, firms, and public institutions | Global Innovation Index | Composite Index |
| Global R&D companies, average expenditure top 3 | RDC | Average R&D spending of the top three global firms operating in each country | Global Innovation Index | USD (Millions) |
| Variable | Acronym | Definition | Source | Unit |
|---|---|---|---|---|
| Trade as a percentage of GDP | Trade | Measures the relative importance of international trade in a country’s economy. It is calculated by summing the value of a nation’s exports and imports of goods and services, then dividing this total by its GDP | World Bank | % |
| Business environment | BE | Captures institutional, legal, and macroeconomic conditions affecting firms, including infrastructure, regulatory quality, and political stability | Global Innovation Index | Composite Index |
| High-tech exports, % total trade | HTE | Share of a country’s total exports classified as high-technology products, including aerospace, biotechnology, pharmaceuticals, and ICT goods | Global Innovation Index | % of total trade |
| ICT services imports, % total trade | ICTSI | Measures the share of trade accounted for by imported ICT services such as software development, telecommunications, and data processing | Global Innovation Index | % of total trade |
| ICT use | ICTU | Composite index reflecting the degree of ICT diffusion across households, firms, and public institutions | Global Innovation Index | Composite Index |
| Global R&D companies, average expenditure top 3 | RDC | Average R&D spending of the top three global firms operating in each country | Global Innovation Index | USD (Millions) |
Figure 1 reports descriptive statistics for each variable, visualized through histograms, density plots, and box plots. Trade displays a positively skewed distribution, with many countries reporting relatively low trade-to-GDP ratios, and a few small economies appearing as outliers. BE is more symmetrically distributed, though variation persists across institutional contexts. Both HTE and ICTSI show considerable left skew, with only a minority of countries exhibiting high shares. Average spending by global R&D companies (RDC) is highly concentrated at the lower end, reflecting limited global R&D investment in most countries. ICTU presents a more even distribution, with a wide interquartile range, consistent with diverse levels of digital infrastructure. These characteristics point out the need for modeling approaches that are both robust to distributional irregularities and sensitive to heterogeneity.
The illustration consists of 18 plots, arranged in six rows and three columns, each representing a different type of plot for each variable. The first column contains histograms, the second column contains boxplots, and the third column contains density plots. At the top, the text reads, “Descriptive Statistics for Selected Variables.” Row 1: “Trade” Histogram of trade: The horizontal axis represents the “Trade” and spans from 0 to 900 with an interval of 300. The vertical axis is labeled “Count,” ranging from 0 to 75 with an interval of 25. The histogram shows a distribution, with very small bars nearly of zero height between the trade values of 300 and 600. A vertical line is positioned at zero trade. Boxplot of trade: The horizontal axis spans from negative 0.4 to 0.4 with an interval of 0.2. The vertical axis is labeled “Trade,” ranging from 0 to 900 with an interval of 300. Two boxes are stacked vertically at the center and elongated from left to right. A vertical line is drawn above and below the boxes at 0 value on the horizontal axis. Density Plot of Trade: The horizontal axis represents the “Trade” and spans from 0 to 900 with an interval of 300. The vertical axis is labeled “density,” ranging from 0 to 0.00075 with an interval of 0.00025. The curve starts from the zero trade at a higher density and runs towards the right with a slight curve profile, which slightly bends downward just before ending after the trade of 900. Row 2: “B E” Histogram of B E: The horizontal axis represents the “B E” and spans from 0 to 100 with an interval of 25. The vertical axis is labeled “Count,” ranging from 0 to 50 with an interval of 10. The histogram reveals a multimodal distribution with multiple peaks, indicating a mixed or complex data pattern where values are clustered between the trades of 50 and 90. Boxplot of B E: The horizontal axis spans from negative 0.4 to 0.4 with an interval of 0.2. The vertical axis is labeled “B E,” ranging from 0 to 100 with an interval of 25. Two boxes are stacked vertically at the center and elongated from left to right. The boxes are comparatively thinner than the boxes in the trade plot. A vertical line is drawn above and below the boxes at 0 value on the horizontal axis, and some black dots are clustered at the lower end of the vertical line. Density Plot of B E: The horizontal axis represents the “B E” and spans from 0 to 100 with an interval of 25. The vertical axis is labeled “density,” ranging from 0 to 0.02 with an interval of 0.01. The curve starts from the zero trade at a zero density and increases slowly towards the right with a slight curve profile and peaks nearly at 62.5 and then decreases towards the right. Row 3: “R D C” Histogram of R D C: The horizontal axis represents the “R D C” and spans from 0 to 200 with an interval of 50. The vertical axis is labeled “Count,” ranging from 0 to 1000 with an interval of 250. The histogram demonstrates a concentration of values on the left side, which shows a thick vertical line at the zero value of R D C. Boxplot of R D C: The horizontal axis spans from negative 0.4 to 0.4 with an interval of 0.2. The vertical axis is labeled “R D C,” ranging from 0 to 200 with an interval of 50. It shows a thin black horizontal line at the zero value of R D C, and a thick black vertical line is placed at the zero value on the horizontal axis. Density Plot of R D C: The horizontal axis represents the “R D C” and spans from 0 to 200 with an interval of 50. The vertical axis is labeled “density,” ranging from 0 to 0.03 with an interval of 0.01. The density curve starts from the zero R D C with a high value of density, and then it decreases gradually towards the right and runs almost horizontally from 35 to the bottom right side. Row 4: “H T E” Histogram of H T E: The horizontal axis represents the “H T E” and spans from 0 to 100 with an interval of 25. The vertical axis is labeled “Count,” ranging from 0 to 200 with an interval of 50. The histogram has a sharp peak near zero, with most of the values concentrated at the lower end, indicating a skewed distribution that decreases towards the right. Boxplot of H T E: The horizontal axis spans from negative 0.4 to 0.4 with an interval of 0.2. The vertical axis is labeled “H T E,” ranging from 0 to 100 with an interval of 25. It shows two horizontal boxes positioned between 0 and 25 of H T E. The bottom box is comparatively thinner than the upper box. A vertical line is positioned above the upper box at zero value on the horizontal axis, and dots are clustered at the top end of the line. Density Plot of H T E: The horizontal axis represents the “H T E” and spans from 0 to 100 with an interval of 25. The vertical axis is labeled “density,” ranging from 0 to 0.05 with an interval of 0.01. The density curve starts from the zero R D C with a high value of density and slightly increases initially, then it decreases gradually towards the right, which ends at the bottom right side. Row 5: “I C T S I” Histogram of I C T S I: The horizontal axis represents the “I C T S I” and spans from 0 to 400 with an interval of 100. The vertical axis is labeled “Count,” ranging from 0 to 400 with an interval of 100. The histogram shows very small, thin bars at the lower and higher values of H T E. Boxplot of I C T S I: The horizontal axis spans from negative 0.4 to 0.4 with an interval of 0.2. The vertical axis is labeled “I C T S I,” ranging from 0 to 400 with an interval of 100. It shows two horizontal boxes positioned between 0 and 250 of I C T S I. The bottom box is comparatively thinner than the upper box, but both are wider than the boxes shown in the plot of H T E. A vertical line is positioned above the upper box at zero value on the horizontal axis. Density Plot of I C T S I: The horizontal axis represents the “I C T S I” and spans from 0 to 400 with an interval of 100. The vertical axis is labeled “density,” ranging from 0 to 0.004 with an interval of 0.002. The density curve starts from the zero I C T S I with a high value of density and slightly increases initially, then it decreases gradually towards the right, which ends at the bottom right side. Row 6: “I C T U” Histogram of I C T U: The horizontal axis represents the “I C T U” and spans from 0 to 75 with an interval of 25. The vertical axis is labeled “Count,” ranging from 0 to 75 with an interval of 25. The histogram shows small bars under 25 counts, and these run from left to right. A single bar peaks at the far left side at the zero value of I C T U. Boxplot of I C T U: The horizontal axis spans from negative 0.4 to 0.4 with an interval of 0.2. The vertical axis is labeled “I C T U,” ranging from 0 to 75 with an interval of 25. It shows two horizontal boxes positioned between 12.5 and 75 of I C T U. The top box is slightly thinner than the lower box and very similar to the boxes shown in the trade plot. A vertical line is positioned above and below the boxes at zero value on the horizontal axis. Density Plot of I C T U: The horizontal axis represents the “I C T U” and spans from 0 to 75 with an interval of 25. The vertical axis is labeled “density,” ranging from 0 to 0.012 with an interval of 0.003. The density curve starts from the zero I C T U with a high value of density and slightly increases initially, then it decreases and increases slowly and bends downward before ending on the right side. Note: All numerical data values are approximated.Descriptive statistics. Source: Authors’ own creation
The illustration consists of 18 plots, arranged in six rows and three columns, each representing a different type of plot for each variable. The first column contains histograms, the second column contains boxplots, and the third column contains density plots. At the top, the text reads, “Descriptive Statistics for Selected Variables.” Row 1: “Trade” Histogram of trade: The horizontal axis represents the “Trade” and spans from 0 to 900 with an interval of 300. The vertical axis is labeled “Count,” ranging from 0 to 75 with an interval of 25. The histogram shows a distribution, with very small bars nearly of zero height between the trade values of 300 and 600. A vertical line is positioned at zero trade. Boxplot of trade: The horizontal axis spans from negative 0.4 to 0.4 with an interval of 0.2. The vertical axis is labeled “Trade,” ranging from 0 to 900 with an interval of 300. Two boxes are stacked vertically at the center and elongated from left to right. A vertical line is drawn above and below the boxes at 0 value on the horizontal axis. Density Plot of Trade: The horizontal axis represents the “Trade” and spans from 0 to 900 with an interval of 300. The vertical axis is labeled “density,” ranging from 0 to 0.00075 with an interval of 0.00025. The curve starts from the zero trade at a higher density and runs towards the right with a slight curve profile, which slightly bends downward just before ending after the trade of 900. Row 2: “B E” Histogram of B E: The horizontal axis represents the “B E” and spans from 0 to 100 with an interval of 25. The vertical axis is labeled “Count,” ranging from 0 to 50 with an interval of 10. The histogram reveals a multimodal distribution with multiple peaks, indicating a mixed or complex data pattern where values are clustered between the trades of 50 and 90. Boxplot of B E: The horizontal axis spans from negative 0.4 to 0.4 with an interval of 0.2. The vertical axis is labeled “B E,” ranging from 0 to 100 with an interval of 25. Two boxes are stacked vertically at the center and elongated from left to right. The boxes are comparatively thinner than the boxes in the trade plot. A vertical line is drawn above and below the boxes at 0 value on the horizontal axis, and some black dots are clustered at the lower end of the vertical line. Density Plot of B E: The horizontal axis represents the “B E” and spans from 0 to 100 with an interval of 25. The vertical axis is labeled “density,” ranging from 0 to 0.02 with an interval of 0.01. The curve starts from the zero trade at a zero density and increases slowly towards the right with a slight curve profile and peaks nearly at 62.5 and then decreases towards the right. Row 3: “R D C” Histogram of R D C: The horizontal axis represents the “R D C” and spans from 0 to 200 with an interval of 50. The vertical axis is labeled “Count,” ranging from 0 to 1000 with an interval of 250. The histogram demonstrates a concentration of values on the left side, which shows a thick vertical line at the zero value of R D C. Boxplot of R D C: The horizontal axis spans from negative 0.4 to 0.4 with an interval of 0.2. The vertical axis is labeled “R D C,” ranging from 0 to 200 with an interval of 50. It shows a thin black horizontal line at the zero value of R D C, and a thick black vertical line is placed at the zero value on the horizontal axis. Density Plot of R D C: The horizontal axis represents the “R D C” and spans from 0 to 200 with an interval of 50. The vertical axis is labeled “density,” ranging from 0 to 0.03 with an interval of 0.01. The density curve starts from the zero R D C with a high value of density, and then it decreases gradually towards the right and runs almost horizontally from 35 to the bottom right side. Row 4: “H T E” Histogram of H T E: The horizontal axis represents the “H T E” and spans from 0 to 100 with an interval of 25. The vertical axis is labeled “Count,” ranging from 0 to 200 with an interval of 50. The histogram has a sharp peak near zero, with most of the values concentrated at the lower end, indicating a skewed distribution that decreases towards the right. Boxplot of H T E: The horizontal axis spans from negative 0.4 to 0.4 with an interval of 0.2. The vertical axis is labeled “H T E,” ranging from 0 to 100 with an interval of 25. It shows two horizontal boxes positioned between 0 and 25 of H T E. The bottom box is comparatively thinner than the upper box. A vertical line is positioned above the upper box at zero value on the horizontal axis, and dots are clustered at the top end of the line. Density Plot of H T E: The horizontal axis represents the “H T E” and spans from 0 to 100 with an interval of 25. The vertical axis is labeled “density,” ranging from 0 to 0.05 with an interval of 0.01. The density curve starts from the zero R D C with a high value of density and slightly increases initially, then it decreases gradually towards the right, which ends at the bottom right side. Row 5: “I C T S I” Histogram of I C T S I: The horizontal axis represents the “I C T S I” and spans from 0 to 400 with an interval of 100. The vertical axis is labeled “Count,” ranging from 0 to 400 with an interval of 100. The histogram shows very small, thin bars at the lower and higher values of H T E. Boxplot of I C T S I: The horizontal axis spans from negative 0.4 to 0.4 with an interval of 0.2. The vertical axis is labeled “I C T S I,” ranging from 0 to 400 with an interval of 100. It shows two horizontal boxes positioned between 0 and 250 of I C T S I. The bottom box is comparatively thinner than the upper box, but both are wider than the boxes shown in the plot of H T E. A vertical line is positioned above the upper box at zero value on the horizontal axis. Density Plot of I C T S I: The horizontal axis represents the “I C T S I” and spans from 0 to 400 with an interval of 100. The vertical axis is labeled “density,” ranging from 0 to 0.004 with an interval of 0.002. The density curve starts from the zero I C T S I with a high value of density and slightly increases initially, then it decreases gradually towards the right, which ends at the bottom right side. Row 6: “I C T U” Histogram of I C T U: The horizontal axis represents the “I C T U” and spans from 0 to 75 with an interval of 25. The vertical axis is labeled “Count,” ranging from 0 to 75 with an interval of 25. The histogram shows small bars under 25 counts, and these run from left to right. A single bar peaks at the far left side at the zero value of I C T U. Boxplot of I C T U: The horizontal axis spans from negative 0.4 to 0.4 with an interval of 0.2. The vertical axis is labeled “I C T U,” ranging from 0 to 75 with an interval of 25. It shows two horizontal boxes positioned between 12.5 and 75 of I C T U. The top box is slightly thinner than the lower box and very similar to the boxes shown in the trade plot. A vertical line is positioned above and below the boxes at zero value on the horizontal axis. Density Plot of I C T U: The horizontal axis represents the “I C T U” and spans from 0 to 75 with an interval of 25. The vertical axis is labeled “density,” ranging from 0 to 0.012 with an interval of 0.003. The density curve starts from the zero I C T U with a high value of density and slightly increases initially, then it decreases and increases slowly and bends downward before ending on the right side. Note: All numerical data values are approximated.Descriptive statistics. Source: Authors’ own creation
Figure 2 provides a correlation heatmap, illustrating moderate correlations between Trade and BE, and between Trade and HTE. ICTU and ICTSI exhibit a stronger correlation, suggesting that digital adoption and digital trade are interlinked. The correlation involving RDC is weaker with Trade, suggesting the potential for a complex, nonlinear relationship.
The diagram shows a correlation matrix heatmap arranged in a square grid. The rows from left to right and columns from bottom to top are each labeled with the categories “Trade,” “B E,” “R D C,” “H T E,” “I C T S I,” and “I C T U.” The heatmap’s color gradient ranges from light blue to dark orange, with dark orange indicating a strong positive correlation, while lighter shades represent weaker positive correlations. Each square within the grid represents the correlation value between the corresponding row and column labels. The cells along the diagonal, where the row and column labels match (like “Trade” with “Trade”), represent the correlation of each variable with itself, which is always 1. The color bar on the right side of the heatmap illustrates the correlation scale, ranging from negative 1 (light blue) to 1 (dark orange). The bottom row and left column are shown in a very light shade, while moving upward, the shade increases.Correlation heatmap. Source: Authors’ own creation
The diagram shows a correlation matrix heatmap arranged in a square grid. The rows from left to right and columns from bottom to top are each labeled with the categories “Trade,” “B E,” “R D C,” “H T E,” “I C T S I,” and “I C T U.” The heatmap’s color gradient ranges from light blue to dark orange, with dark orange indicating a strong positive correlation, while lighter shades represent weaker positive correlations. Each square within the grid represents the correlation value between the corresponding row and column labels. The cells along the diagonal, where the row and column labels match (like “Trade” with “Trade”), represent the correlation of each variable with itself, which is always 1. The color bar on the right side of the heatmap illustrates the correlation scale, ranging from negative 1 (light blue) to 1 (dark orange). The bottom row and left column are shown in a very light shade, while moving upward, the shade increases.Correlation heatmap. Source: Authors’ own creation
The dataset used in this analysis is fully balanced across selected countries and years, a necessary condition driven by the requirements of most ML algorithms, which cannot accommodate missing values. While completeness was ensured for modeling purposes, it is key to note that secondary sources such as the GII may contain limitations, including missing or interpolated data and potential measurement inconsistencies arising from national reporting practices; these issues, although mitigated, remain relevant when interpreting the results.
To investigate how trade openness responds to the innovation and institutional factors described above, we estimate the following empirical model:
where i represents the 135 countries and t spans the years from 2013 to 2022.
To ensure robust and flexible modeling of the trade–innovation relationship in a cross-country context, we adopt a two-pronged empirical strategy combining standard econometric estimators with ML algorithms. The econometric component addresses well-known limitations of Ordinary Least Squares (OLS) by implementing three complementary estimators that improve model robustness under different types of statistical irregularities.
First, RLS is used to limit the influence of extreme values and uncertain input data. As shown by El Ghaoui and Lebret (1996, 1997), RLS generally minimizes the worst-case residual error by treating the design matrix and response vector as uncertain-but-bounded inputs. The resulting optimization problem, solved through second-order cone programming, yields robust solutions with a formally defined regularization mechanism.
Second, GLM is implemented to relax the OLS assumptions of normality and homoskedasticity. Nelder and Wedderburn (1972) introduced GLM as an extension of linear models to accommodate outcomes from the exponential family, such as binomial or Poisson distributions, via link functions. As further emphasized by Neuhaus and McCulloch (2011), GLM supports a wide range of non-normally distributed dependent variables – including binary and count outcomes – making it suitable for heterogeneous international datasets.
Third, quantile regression is applied to investigate whether covariate effects differ across the conditional distribution of trade. Koenker and Hallock (2001) characterize quantile regression as a generalization of mean regression that estimates conditional quantiles, such as the median, by minimizing asymmetrically weighted absolute deviations. According to Koenker (2017), this approach is particularly effective for capturing heterogeneous effects and for providing robustness when conventional statistical procedures fall short.
Although the Generalized Method of Moments (GMM) is a common approach for addressing endogeneity in panel data, it is not employed in this context. Wooldridge (2001) notes that while GMM is indispensable in complex applications, it does not consistently outperform simpler estimators like OLS or Two-Stage Least Squares in typical empirical settings. In our case, the absence of strong internal instruments and the lack of significant dynamic dependence in the trade variable render GMM both unnecessary and potentially unreliable.
To capture potential nonlinearities and interactions among covariates, the analysis is extended with a suite of ML models. GBM, introduced by Friedman (2001), is used to construct additive models through stagewise gradient descent in function space. When applied with decision trees as base learners, GBM produces robust and interpretable predictions and is well-suited for regression tasks involving noisy or structurally complex data.
We also integrate bagging (bootstrap aggregating) to reduce model variance. As defined by Breiman (1996), bagging generates multiple versions of a predictor by training on bootstrap samples and then aggregates their predictions. The method is especially effective for unstable learners, such as decision trees, and empirical tests show significant accuracy improvements in both regression and classification tasks. To operationalize bagging within our ensemble framework, we adopt a Random Forest (RF) model, a tree-based method developed by Breiman (2001) that extends bagging by introducing random feature selection at each node. RF is, therefore, included as the bagging component in the stacked model.
The ensemble model is constructed using the Super Learner framework (see, inter alia, van der Laan et al., 2007). This approach combines predictions from a library of candidate algorithms into a single meta-learner optimized via cross-validation [3]. The final ensemble integrates GLM, GBM, and RF, with the latter serving as a bagging-based component. This allows us to leverage the individual strengths of each learner. We implement the ensemble through the H2O platform, which provides an efficient computational environment for stacking, as described in Boehmke and Greenwell (2019). This strategy ensures a comprehensive treatment of nonlinearities, model instability, and distributional heterogeneity while enhancing predictive performance in a multi-method setting.
4. Empirical analysis
The results are presented across Tables 2 and 4, each corresponding to a distinct econometric model: RLS, GLM, and quantile regression. These methods allow us to assess the stability, heterogeneity, and distributional dynamics of trade determinants.
Results of RLS
| Variable | Coefficient | Std. error | z statistic | p-value |
|---|---|---|---|---|
| BE | 0.1413 | 0.0207 | 6.8212 | 0.0000*** |
| HTE | 0.2560 | 0.0199 | 12.8588 | 0.0000*** |
| ICTSI | 0.0864 | 0.0189 | 4.5626 | 0.0000*** |
| ICTU | 0.0499 | 0.0219 | 2.2785 | 0.0227** |
| RDC | −0.1831 | 0.0221 | −8.2829 | 0.0000*** |
| Constant | −0.1647 | 0.0172 | −9.5543 | 0.0000*** |
| R2 | 0.8388 | Adjusted R2 | 0.8353 | |
| SER | 0.9379 | Scale | 0.5815 |
| Variable | Coefficient | Std. error | z statistic | p-value |
|---|---|---|---|---|
| BE | 0.1413 | 0.0207 | 6.8212 | 0.0000*** |
| HTE | 0.2560 | 0.0199 | 12.8588 | 0.0000*** |
| ICTSI | 0.0864 | 0.0189 | 4.5626 | 0.0000*** |
| ICTU | 0.0499 | 0.0219 | 2.2785 | 0.0227** |
| RDC | −0.1831 | 0.0221 | −8.2829 | 0.0000*** |
| Constant | −0.1647 | 0.0172 | −9.5543 | 0.0000*** |
| R2 | 0.8388 | Adjusted R2 | 0.8353 | |
| SER | 0.9379 | Scale | 0.5815 |
Note(s): Method: M-estimation; M settings: weight = Bisquare, tuning = 4.685, scale = MAD (median centered). Huber Type I Standard Errors and Covariance. ***p < 0.01, **p < 0.05, *p < 0.10
Results of GLM
| Variable | Coefficient | Std. error | z statistic | p-value |
|---|---|---|---|---|
| BE | 0.1778 | 0.0563 | 3.1571 | 0.0016*** |
| HTE | 0.2600 | 0.0854 | 3.0446 | 0.0023*** |
| ICTSI | 0.1018 | 0.0647 | 1.5742 | 0.1154 |
| ICTU | 0.1591 | 0.0714 | 2.2274 | 0.0259** |
| RDC | −0.1621 | 0.0570 | −2.8431 | 0.0045 |
| Constant | −0.0034 | 0.0609 | −0.0552 | 0.9560 |
| HQIC | 2.6663 | RMSE | 0.9091 | |
| Dispersion | 0.8305 |
| Variable | Coefficient | Std. error | z statistic | p-value |
|---|---|---|---|---|
| BE | 0.1778 | 0.0563 | 3.1571 | 0.0016*** |
| HTE | 0.2600 | 0.0854 | 3.0446 | 0.0023*** |
| ICTSI | 0.1018 | 0.0647 | 1.5742 | 0.1154 |
| ICTU | 0.1591 | 0.0714 | 2.2274 | 0.0259** |
| RDC | −0.1621 | 0.0570 | −2.8431 | 0.0045 |
| Constant | −0.0034 | 0.0609 | −0.0552 | 0.9560 |
| HQIC | 2.6663 | RMSE | 0.9091 | |
| Dispersion | 0.8305 |
Note(s): Newton-Raphson/Marquardt steps. Dispersion computed using Pearson’s Chi-Square. Coefficient covariance computed using the Huber-White method with observed Hessian (Bartlett kernel). ***p < 0.01, **p < 0.05, *p < 0.10
Results of quantile regressions
| Variable | Quantile regressions | |||
|---|---|---|---|---|
| 0.20 | 0.40 | 0.60 | 0.80 | |
| BE | 0.1356*** (0.0282) | 0.1334*** (0.0305) | 0.1331*** (0.0349) | 0.1803*** (0.0287) |
| HTE | 0.1228*** (0.0281) | 0.2453*** (0.0339) | 0.3225*** (0.0498) | 0.4567*** (0.0644) |
| ICTSI | 0.0397* (0.0206) | 0.1068*** (0.0227) | 0.0676** (0.0315) | 0.0833 (0.0509) |
| ICTU | 0.0788*** (0.0288) | 0.0609** (0.0303) | 0.1045*** (0.0355) | 0.0460 (0.0469) |
| RDC | −0.0889*** (0.0263) | −0.1634*** (0.0277) | −0.2103*** (0.0282) | −0.1827*** (0.0708) |
| Constant | −0.6120*** (0.0189) | −0.3280*** (0.0212) | 0.0186 (0.0302) | 0.4308*** (0.0384) |
| Pseudo R2 | 0.8062 | 0.8321 | 0.9099 | 0.9336 |
| SER | 1.1128 | 0.9773 | 0.9194 | 1.0267 |
| Sparsity | 1.5111 | 1.4027 | 1.7025 | 2.7089 |
| Variable | Quantile regressions | |||
|---|---|---|---|---|
| 0.20 | 0.40 | 0.60 | 0.80 | |
| BE | 0.1356*** (0.0282) | 0.1334*** (0.0305) | 0.1331*** (0.0349) | 0.1803*** (0.0287) |
| HTE | 0.1228*** (0.0281) | 0.2453*** (0.0339) | 0.3225*** (0.0498) | 0.4567*** (0.0644) |
| ICTSI | 0.0397* (0.0206) | 0.1068*** (0.0227) | 0.0676** (0.0315) | 0.0833 (0.0509) |
| ICTU | 0.0788*** (0.0288) | 0.0609** (0.0303) | 0.1045*** (0.0355) | 0.0460 (0.0469) |
| RDC | −0.0889*** (0.0263) | −0.1634*** (0.0277) | −0.2103*** (0.0282) | −0.1827*** (0.0708) |
| Constant | −0.6120*** (0.0189) | −0.3280*** (0.0212) | 0.0186 (0.0302) | 0.4308*** (0.0384) |
| Pseudo R2 | 0.8062 | 0.8321 | 0.9099 | 0.9336 |
| SER | 1.1128 | 0.9773 | 0.9194 | 1.0267 |
| Sparsity | 1.5111 | 1.4027 | 1.7025 | 2.7089 |
Note(s): Huber Sandwich Standard Errors and Covariance; Sparsity method: Kernel (Epanechnikov) using residuals; Bandwidth method: Hall-Sheather. ***p < 0.01, **p < 0.05, *p < 0.10
Table 2 reports the results from the RLS model, which is particularly resilient to outliers and misspecification.
It reveals statistically significant and positive coefficients for BE, HTE, ICTSI, and ICTU, while RDC shows a significant negative effect. The adjusted R2 of 0.8353 confirms a strong fit, supporting the relevance of these predictors in explaining trade variations.
Table 3 displays the GLM results.
This model is suitable for capturing relationships under flexible assumptions about distribution and link functions. The findings largely confirm the RLS estimates: BE, HTE, and ICTU retain positive and significant associations with trade, though ICTSI’s effect becomes insignificant, possibly due to distributional sensitivity. The GLM thus offers a robustness check for structural consistency across specifications.
Table 4 introduces the results from quantile regressions across the 20th, 40th, 60th, and 80th percentiles of the trade distribution.
This analysis reveals considerable heterogeneity in trade determinants across different regimes. HTE shows a progressively stronger positive association with trade from lower to upper quantiles, suggesting that high-tech exports become increasingly central to trade dynamics as countries move up the development ladder. This finding is consistent with evidence from OECD countries, where the export performance of high-tech sectors is closely linked to digital infrastructure and GDP growth. Gürler (2023) demonstrates a significant positive correlation between HTE and ICT service imports, further reinforcing the role of digital capacity in expanding trade potential.
By contrast, RDC exhibits an overall stable and negative impact across all quantiles, pointing to its structural nature. Countries with chronic underinvestment in R&D may face persistent constraints in upgrading export capabilities, especially in high-value-added sectors. This is clearly noticeable in the case of China, where HTE expansion remains heavily reliant on foreign inputs. Kim (2011) shows that China’s exports in high-tech industries are largely driven by processing trade for foreign firms, limiting domestic value capture and innovation growth (see Figure A1 in the Appendix).
Consistent with the conceptual framework in Figure A2 in the Appendix, ICTSI demonstrates a non-linear influence, peaking at the median quantile, indicating that digital service imports contribute more substantially to trade in mid-level economies where absorptive capacities and infrastructure may be optimally aligned. ICTU, on the other hand, shows a declining impact in higher quantiles, implying that digital infrastructure plays a more decisive role in trade facilitation at earlier stages of integration (see Figure A3 in the Appendix). Dash and Parida (2012) highlight how ICT-enhanced service exports supported post-reform growth in India, while Yousefi (2018) finds that internet penetration significantly increases service trade, especially for exports in developed economies.
Nevertheless, greater trade openness – reflected in trade as a percentage of GDP – can bring new challenges. Liberalization without regulatory preparedness may lead to economic vulnerabilities and business environment deterioration (see Figure A4 in the Appendix). Gani and Clemes (2013) show that weak legal enforcement and institutional inefficiencies hinder growth in services exports, particularly in lower-income economies. Moreover, the environmental impacts of trade liberalization are especially pronounced in developing countries. Bekmez and Ozsoy (2016) provide panel data evidence that trade openness correlates with higher CO2 emissions in these contexts, suggesting a trade-environment conflict that supports the Pollution Haven Hypothesis.
Taken together, these findings highlight the dual nature of trade-enhancing factors. While ICTSI, ICTU, and HTE are positively linked to trade performance, their benefits are contingent on supportive structural conditions. In contrast, persistent RDC, weak regulatory systems, and dependence on foreign technology undermine long-term competitiveness. A strategic balance between digital integration, institutional development, and domestic innovation is therefore essential for trade-led growth to be both inclusive and sustainable.
To complement econometric findings, we implement ML techniques – linear model, GBM, and bagging with RF [4] – to explore complex patterns, interactions, and confirm variable relevance. Figure 3 shows coefficients from the linear model and feature importance from GBM and bagging.
The diagram features three horizontal bar charts displayed vertically, each illustrating a different analysis method. The left chart, labeled “Linear Model Coefficients,” shows the coefficients for five features on the vertical axis from top to bottom as “H T E,” “I C T U,” “I C T S I,” “B E,” and “R D C.” The vertical axis in each chart is labeled “Explanatory Variable.” The horizontal axis is labeled “Coefficient” and ranges from negative 0.1 to 0.3 with an interval of 0.1. The bars are positioned horizontally, with the length of each bar representing the coefficient value. The bars are ordered from top to bottom, with “H T E” having the highest positive coefficient, extending from 0 to 0.3, and “R D C” showing the lowest negative coefficient, extending from 0 to negative 0.2. The size of the bars decreases from top to bottom, and the bars are colored in light blue. The middle chart, titled “G B M Feature Importance,” shows the coefficients for five features on the vertical axis from top to bottom as “H T E,” “B E,” “I C T U,” “I C T S I,” and “R D C.” The vertical axis in each chart is labeled “Explanatory Variable.” The horizontal axis is labeled “Importance” and ranges from negative 0 to 0.15 with an interval of 0.05. The importance values for the feature “H T E” have the largest bar, indicating the highest importance, and “R D C” the smallest. The bars are again arranged in descending order of importance, and the bars are colored pink. The right chart, “Bagging using RF importance,” shows the coefficients for five features on the vertical axis from top to bottom as “B E,” “H T E,” “I C T U,” “ICTSI,” and “R D C.” The vertical axis in each chart is labeled “Explanatory variable.” The horizontal axis is labeled “Importance” and ranges from negative 0 to 150 with an interval of 50. The importance values for the features “B E” and “H T E” show the longest bars, indicating the highest importance. “I C T U” and “I C T S I” are next in importance, with “R D C” at the lowest. The size of the bars decreases from top to bottom, and the bars are colored in orange.Models’ coefficients and importance scores. Source: Authors’ own creation
The diagram features three horizontal bar charts displayed vertically, each illustrating a different analysis method. The left chart, labeled “Linear Model Coefficients,” shows the coefficients for five features on the vertical axis from top to bottom as “H T E,” “I C T U,” “I C T S I,” “B E,” and “R D C.” The vertical axis in each chart is labeled “Explanatory Variable.” The horizontal axis is labeled “Coefficient” and ranges from negative 0.1 to 0.3 with an interval of 0.1. The bars are positioned horizontally, with the length of each bar representing the coefficient value. The bars are ordered from top to bottom, with “H T E” having the highest positive coefficient, extending from 0 to 0.3, and “R D C” showing the lowest negative coefficient, extending from 0 to negative 0.2. The size of the bars decreases from top to bottom, and the bars are colored in light blue. The middle chart, titled “G B M Feature Importance,” shows the coefficients for five features on the vertical axis from top to bottom as “H T E,” “B E,” “I C T U,” “I C T S I,” and “R D C.” The vertical axis in each chart is labeled “Explanatory Variable.” The horizontal axis is labeled “Importance” and ranges from negative 0 to 0.15 with an interval of 0.05. The importance values for the feature “H T E” have the largest bar, indicating the highest importance, and “R D C” the smallest. The bars are again arranged in descending order of importance, and the bars are colored pink. The right chart, “Bagging using RF importance,” shows the coefficients for five features on the vertical axis from top to bottom as “B E,” “H T E,” “I C T U,” “ICTSI,” and “R D C.” The vertical axis in each chart is labeled “Explanatory variable.” The horizontal axis is labeled “Importance” and ranges from negative 0 to 150 with an interval of 50. The importance values for the features “B E” and “H T E” show the longest bars, indicating the highest importance. “I C T U” and “I C T S I” are next in importance, with “R D C” at the lowest. The size of the bars decreases from top to bottom, and the bars are colored in orange.Models’ coefficients and importance scores. Source: Authors’ own creation
The consistency in top predictors (BE, HTE, and ICTU) across all ML approaches reinforces their structural significance in explaining trade performance and broadly validates the findings presented in Tables 2–4.
Figure 4 compares model performance. GBM outperforms linear and bagging models in terms of mean absolute error (MAE), mean squared error (MSE), root mean squared error (RMSE), and R2.
The horizontal axis represents the metrics: “M A E,” “M S E,” “R-squared,” and “R M S E.” Each metric is represented by three bars. The first bar in each group is colored blue, representing the “Linear Model”; the second bar is yellow for “Bagging using R F,” and the third bar is pink for “G B M.” The vertical axis shows the value scale, ranging from 0.00 to 0.75 with an interval of 0.25. For “M A E,” the Linear Model (blue bar) has the highest value, slightly above 0.50, followed by Bagging using R F (yellow bar) at just below 0.50, and G B M (pink bar) below 0.375. In “M S E,” the Linear Model (blue bar) has the highest value, slightly above 0.75, followed by Bagging using R F (yellow bar) at just above 0.375, and G B M (pink bar) below 0.25. For “R-squared,” the Linear Model (blue bar) has the lowest value, slightly below 0.25, followed by Bagging using R F (yellow bar) at just above 0.625, and G B M (pink bar) below 0.75. Finally, for “R M S E,” the Linear Model (blue bar) has the highest value, slightly above 0.875, followed by Bagging using R F (yellow bar) at just below 0.625, and G B M (pink bar) below 0.50.Models’ performance metrics. Source: Authors’ own creation
The horizontal axis represents the metrics: “M A E,” “M S E,” “R-squared,” and “R M S E.” Each metric is represented by three bars. The first bar in each group is colored blue, representing the “Linear Model”; the second bar is yellow for “Bagging using R F,” and the third bar is pink for “G B M.” The vertical axis shows the value scale, ranging from 0.00 to 0.75 with an interval of 0.25. For “M A E,” the Linear Model (blue bar) has the highest value, slightly above 0.50, followed by Bagging using R F (yellow bar) at just below 0.50, and G B M (pink bar) below 0.375. In “M S E,” the Linear Model (blue bar) has the highest value, slightly above 0.75, followed by Bagging using R F (yellow bar) at just above 0.375, and G B M (pink bar) below 0.25. For “R-squared,” the Linear Model (blue bar) has the lowest value, slightly below 0.25, followed by Bagging using R F (yellow bar) at just above 0.625, and G B M (pink bar) below 0.75. Finally, for “R M S E,” the Linear Model (blue bar) has the highest value, slightly above 0.875, followed by Bagging using R F (yellow bar) at just below 0.625, and G B M (pink bar) below 0.50.Models’ performance metrics. Source: Authors’ own creation
This validates its strength in capturing non-linearities and variable interactions often missed by parametric models. Lastly, Figure 5 displays partial dependence plots (PDPs) [5] for BE, HTE, ICTSI, ICTU, and RDC. These reveal marginal effects net of all other variables.
The illustration displays a grid of ten partial dependence plots (P D P), arranged in five rows and two columns. The left column represents the “Boosting Model P D P,” and the right column represents the “Bagging Model P D P.” In each plot, the horizontal axis represents the feature values, and the vertical axis represents the predicted outcome (y hat). Left Group (Boosting Model P D P): Row 1, Column 1: The horizontal axis is labeled “B E” and ranges from negative 4 to 2 with an interval of 2. The vertical axis is labeled “y hat” and ranges from 0 to 1.5 with an interval of 0.5. The curve starts at a very low value from the bottom left side, then increases slowly with small variation, and shows a rapid increase between 0 and 2 on the horizontal axis, and ends at the top right side. Row 2, Column 1: The horizontal axis is labeled “H T E” and ranges from 0 to 3 with an interval of 1. The vertical axis is labeled “y hat” and ranges from 0 to 0.8 with an interval of 0.2. The curve begins just above 0.2 on the vertical axis and then decreases and increases with multiple small peaks and troughs and then increases towards the right, which ends at the top right side. Row 3, Column 1: The horizontal axis is labeled “I C T S I” and ranges from negative 1 to 4 with an interval of 1. The vertical axis is labeled “y hat” and ranges from 0 to 0.3 with an interval of 0.1. The curve begins nearly at 0.1 on the vertical axis and then decreases and increases after value 1 on the horizontal axis and increases towards the right, which ends at the top right side. Row 4, Column 1: The horizontal axis is labeled “I C T U” and ranges from negative 1 to 2 with an interval of 1. The vertical axis is labeled “y hat” and ranges from 0 to 1 with an interval of 0.5. The curve begins at just above 0.25 on the vertical axis and then decreases and increases just after value 1 on the horizontal axis, and then it increases rapidly, showing two small peaks just before ending at the top right side. Row 5, Column 1: The horizontal axis is labeled “R D C” and ranges from 0 to 3 with an interval of 1. The vertical axis is labeled “y hat” and ranges from negative 0.2 to 0.1 with an interval of 0.1. The curve begins nearly at 0.05 on the vertical axis and then increases and reaches a peak nearly at value 1 on the horizontal axis, then it decreases with small variation and ends at the bottom right side. Right Group (Bagging Model P D P): Row 1, Column 2: The horizontal axis is labeled “B E” and ranges from negative 4 to 2 with an interval of 2. The vertical axis is labeled “y hat” and ranges from 0 to 0.8 with an interval of 0.4. The curve starts at a very low value from the bottom left side, then increases slowly with small variation, and shows a rapid increase between 0 and 2 on the horizontal axis, and ends at the top right side. Row 2, Column 2: The horizontal axis is labeled “H T E” and ranges from 0 to 3 with an interval of 1. The vertical axis is labeled “y hat” and ranges from 0 to 0.6 with an interval of 0.2. The curve begins just above 0 on the vertical axis and then decreases and then increases with multiple small peaks and troughs and then increases towards the right, which ends at the top right side. Row 3, Column 2: The horizontal axis is labeled “I C T S I” and ranges from negative 1 to 4 with an interval of 1. The vertical axis is labeled “y hat” and ranges from 0 to 0.6 with an interval of 0.2. The curve begins nearly at 0.1 on the vertical axis and then decreases and then increases after value 1.5 on the horizontal axis and increases towards the right, showing rapid increases before ending at the top right side. Row 4, Column 2: The horizontal axis is labeled “I C T U” and ranges from negative 1 to 2 with an interval of 1. The vertical axis is labeled “y hat” and ranges from 0 to 0.75 with an interval of 0.25. The curve begins just above 0 on the vertical axis and then decreases and increases after the value 0 on the horizontal axis, and then it increases rapidly, showing two small peaks just before ending at the top right side. Row 5, Column 1: The horizontal axis is labeled “R D C” and ranges from 0 to 3 with an interval of 1. The vertical axis is labeled “y hat” and ranges from negative 0.2 to 0.2 with an interval of 0.1. The curve begins nearly at 0.05 on the vertical axis and then increases and reaches a peak nearly at value 1 on the horizontal axis, then it decreases with small variation and ends at the bottom right side. Note: All numerical data values are approximated.PDPs. Source: Authors’ own creation
The illustration displays a grid of ten partial dependence plots (P D P), arranged in five rows and two columns. The left column represents the “Boosting Model P D P,” and the right column represents the “Bagging Model P D P.” In each plot, the horizontal axis represents the feature values, and the vertical axis represents the predicted outcome (y hat). Left Group (Boosting Model P D P): Row 1, Column 1: The horizontal axis is labeled “B E” and ranges from negative 4 to 2 with an interval of 2. The vertical axis is labeled “y hat” and ranges from 0 to 1.5 with an interval of 0.5. The curve starts at a very low value from the bottom left side, then increases slowly with small variation, and shows a rapid increase between 0 and 2 on the horizontal axis, and ends at the top right side. Row 2, Column 1: The horizontal axis is labeled “H T E” and ranges from 0 to 3 with an interval of 1. The vertical axis is labeled “y hat” and ranges from 0 to 0.8 with an interval of 0.2. The curve begins just above 0.2 on the vertical axis and then decreases and increases with multiple small peaks and troughs and then increases towards the right, which ends at the top right side. Row 3, Column 1: The horizontal axis is labeled “I C T S I” and ranges from negative 1 to 4 with an interval of 1. The vertical axis is labeled “y hat” and ranges from 0 to 0.3 with an interval of 0.1. The curve begins nearly at 0.1 on the vertical axis and then decreases and increases after value 1 on the horizontal axis and increases towards the right, which ends at the top right side. Row 4, Column 1: The horizontal axis is labeled “I C T U” and ranges from negative 1 to 2 with an interval of 1. The vertical axis is labeled “y hat” and ranges from 0 to 1 with an interval of 0.5. The curve begins at just above 0.25 on the vertical axis and then decreases and increases just after value 1 on the horizontal axis, and then it increases rapidly, showing two small peaks just before ending at the top right side. Row 5, Column 1: The horizontal axis is labeled “R D C” and ranges from 0 to 3 with an interval of 1. The vertical axis is labeled “y hat” and ranges from negative 0.2 to 0.1 with an interval of 0.1. The curve begins nearly at 0.05 on the vertical axis and then increases and reaches a peak nearly at value 1 on the horizontal axis, then it decreases with small variation and ends at the bottom right side. Right Group (Bagging Model P D P): Row 1, Column 2: The horizontal axis is labeled “B E” and ranges from negative 4 to 2 with an interval of 2. The vertical axis is labeled “y hat” and ranges from 0 to 0.8 with an interval of 0.4. The curve starts at a very low value from the bottom left side, then increases slowly with small variation, and shows a rapid increase between 0 and 2 on the horizontal axis, and ends at the top right side. Row 2, Column 2: The horizontal axis is labeled “H T E” and ranges from 0 to 3 with an interval of 1. The vertical axis is labeled “y hat” and ranges from 0 to 0.6 with an interval of 0.2. The curve begins just above 0 on the vertical axis and then decreases and then increases with multiple small peaks and troughs and then increases towards the right, which ends at the top right side. Row 3, Column 2: The horizontal axis is labeled “I C T S I” and ranges from negative 1 to 4 with an interval of 1. The vertical axis is labeled “y hat” and ranges from 0 to 0.6 with an interval of 0.2. The curve begins nearly at 0.1 on the vertical axis and then decreases and then increases after value 1.5 on the horizontal axis and increases towards the right, showing rapid increases before ending at the top right side. Row 4, Column 2: The horizontal axis is labeled “I C T U” and ranges from negative 1 to 2 with an interval of 1. The vertical axis is labeled “y hat” and ranges from 0 to 0.75 with an interval of 0.25. The curve begins just above 0 on the vertical axis and then decreases and increases after the value 0 on the horizontal axis, and then it increases rapidly, showing two small peaks just before ending at the top right side. Row 5, Column 1: The horizontal axis is labeled “R D C” and ranges from 0 to 3 with an interval of 1. The vertical axis is labeled “y hat” and ranges from negative 0.2 to 0.2 with an interval of 0.1. The curve begins nearly at 0.05 on the vertical axis and then increases and reaches a peak nearly at value 1 on the horizontal axis, then it decreases with small variation and ends at the bottom right side. Note: All numerical data values are approximated.PDPs. Source: Authors’ own creation
Notably, RDC displays negative slopes, corroborating previous regression findings, while the rest of the variables show a positive marginal effect across the data range.
Taken together, the findings provide a multidimensional perspective on the drivers of trade, particularly emphasizing the synergistic roles of technological infrastructure, innovation capacity, and structural development.
First, the consistently positive coefficient for ICTSI supports the argument that robust ICT services facilitate cross-border transactions, reduce search costs, and improve trade efficiency. This is aligned with the findings of Azmeh et al. (2020), who emphasize the rising centrality of digital infrastructure and cross-border data flows in global trade dynamics. Similarly, Singh and Siddiqui (2023) show that ICT penetration significantly correlates with trade and innovation in developing economies, particularly through enhanced connectivity and economic reorganization.
Second, the positive association between BE and trade reflects the growing influence of firm-level and national innovation environments on global integration. Fernández (2023) identifies, among the themes linking innovation and international business, GVCs and cross-border knowledge flows, noting how these dimensions shape firms’ competitiveness and trade performance. Additionally, Anand et al. (2021) describe how emerging economies reconfigure institutional and organizational frameworks to enhance innovation-led growth. These findings suggest that BE contributes by promoting absorptive capacities and dynamic learning across firms.
Third, the positive relationship between HTE and trade is consistent with empirical evidence showing that high-technology exports are associated with R&D-driven productivity gains and technological spillovers. Lee (2020) finds that trade in intermediate inputs and R&D expenditures foster productivity in both frontier and non-frontier economies. Ma (2024) similarly documents that educational expansion boosts innovation in exporting firms, demonstrating that HTE may act as both a cause and consequence of skill-biased trade expansion.
Fourth, ICTU shows a robust positive correlation with trade, which likely reflects the embedded role of upstream ICT usage in enabling efficient logistics, data management, and service integration. According to Hilmersson et al. (2023), firms with accelerated technological adoption tend to internationalize faster. This is also supported by Lei and Xie (2024), who highlight how policy-driven ICT integration in FTZs enhances innovation outputs and firm-level trade capacity.
Finally, the negative and significant impact of RDC points toward a structural disconnect between trade participation and long-term R&D commitments. As noted by Belazreg and Mtar (2020), innovation and trade may at times follow neutral or weakly related paths, stressing the need for greater investment in human capital and creativity to strengthen innovation potential. Lu et al. (2022) further caution that FDI-induced trade may suppress domestic value-added unless accompanied by strong innovation capabilities. Thus, RDC’s negative coefficient may signal an overreliance on short-term export expansion at the expense of sustained technological upgrading.
In sum, the evidence suggests that while trade benefits from ICT-related infrastructure (ICTSI, ICTU), innovation ecosystems (BE, HTE) are critical complements. However, persistent underinvestment in R&D could undermine the sustainability of trade-induced growth, echoing calls by Liu et al. (2022) and Zameer et al. (2020) for stronger innovation–trade alignment in climate-resilient and competitive development models. These findings extend economic theory by showing nonlinear and distributional effects of development indicators on trade. The ML models reinforce and clarify relationships detected in regression outputs, enhancing confidence in the robustness of our results.
5. Robustness checks
To validate the consistency of the findings, a stacking ensemble model was applied, combining the predictive strengths of previously estimated learners. The ensemble incorporates GBM, RF, and GLM, each offering distinct modeling advantages [6]. By aggregating their outputs through a meta-learner, the ensemble captures complex interactions and nonlinearities more effectively than any individual model.
Figure 6 illustrates the performance metrics (MAE, MSE, RMSE, and R2) for the ensemble model. The ensemble achieves strong predictive accuracy, with an R2 exceeding 0.70, and low error values across MAE, MSE, and RMSE. These results confirm the ensemble’s capacity to model the complex relationship between trade and the selected predictors effectively.
The vertical axis is labeled “Value” and ranges from 0 to 0.6 with an interval of 0.2. The horizontal axis is labeled “Metric,” representing four features labeled from left to right as “M A E,” “M S E,” “R-squared,” and “R M S E.” The legend, positioned to the right of the chart, clearly distinguishes the colors for each metric: “M A E” in red, “M S E” in green, “R-squared” in light blue, and “R M S E” in purple. The height of the bars varies, with the “M A E” bar having 0.329, the “M S E” bar having 0.208, the “R-squared” bar having 0.72, and the “R M S E” bar having 0.458. Note: All numerical data values are approximated.Performance metrics of the stacking ensemble model. Source: Authors’ own creation
The vertical axis is labeled “Value” and ranges from 0 to 0.6 with an interval of 0.2. The horizontal axis is labeled “Metric,” representing four features labeled from left to right as “M A E,” “M S E,” “R-squared,” and “R M S E.” The legend, positioned to the right of the chart, clearly distinguishes the colors for each metric: “M A E” in red, “M S E” in green, “R-squared” in light blue, and “R M S E” in purple. The height of the bars varies, with the “M A E” bar having 0.329, the “M S E” bar having 0.208, the “R-squared” bar having 0.72, and the “R M S E” bar having 0.458. Note: All numerical data values are approximated.Performance metrics of the stacking ensemble model. Source: Authors’ own creation
When compared to the individual models displayed in Figure 4, namely, the linear, bagging, and boosting models, the ensemble’s performance remains competitive and close to the top-performing learners. However, its slightly lower R2 (with respect to boosting) can be attributed to the inclusion of less predictive models, such as GLM, in the Super Learner library. Despite this, the ensemble still benefits from the strengths of more accurate base learners (e.g. GBM and RF) and achieves balanced generalization performance, demonstrating the value of model aggregation even in the presence of weaker components.
To better understand the internal structure of the ensemble, Figure 7 reports the coefficients estimated by the GLM meta-learner. These values reflect the relative contribution of each base learner. GBM receives the highest weight, indicating its dominant role in the ensemble’s overall predictive capacity. RF also contributes meaningfully, while GLM shows smaller coefficients, suggesting a complementary rather than central function.
The horizontal axis, labeled “Coefficient,” spans from negative 0.3 to 0.6 with an interval of 0.3. The vertical axis lists the models from top to bottom as “G B M model,” “R F model,” “Intercept,” and “G L M model.” The “GBM model” is represented by a long horizontal bar, extending from the 0 mark to 0.834. The “R F model” extends from 0 to 0.45. The “Intercept” has a very narrow bar, positioned near the 0 mark and extending from 0 to negative 0.014. Finally, the “G L M model” has a bar that starts at 0 and extends slightly to the left, reaching negative 0.277. Each bar is filled with light purple and outlined in black, and they are all aligned horizontally with equal spacing between each model on the vertical axis. Note: All numerical data values are approximated.Meta-learner coefficients in the stacking ensemble. Source: Authors’ own creation
The horizontal axis, labeled “Coefficient,” spans from negative 0.3 to 0.6 with an interval of 0.3. The vertical axis lists the models from top to bottom as “G B M model,” “R F model,” “Intercept,” and “G L M model.” The “GBM model” is represented by a long horizontal bar, extending from the 0 mark to 0.834. The “R F model” extends from 0 to 0.45. The “Intercept” has a very narrow bar, positioned near the 0 mark and extending from 0 to negative 0.014. Finally, the “G L M model” has a bar that starts at 0 and extends slightly to the left, reaching negative 0.277. Each bar is filled with light purple and outlined in black, and they are all aligned horizontally with equal spacing between each model on the vertical axis. Note: All numerical data values are approximated.Meta-learner coefficients in the stacking ensemble. Source: Authors’ own creation
Unlike single models, the stacking approach does not produce variable-level importance scores. Instead, interpretability is shifted to the model level through the meta-learner’s coefficients. The prominence of GBM confirms its robustness, mirroring results obtained in earlier performance comparisons.
6. Conclusions and policy implications
This study investigated the links between trade openness and innovation-related drivers, showing how HTE, ICTSI, ICTU, and RDC influence trade intensity. Using both econometric models (RLS, GLM, quantile) and supervised ML techniques (GBM, bagging via RF, ensemble stacking), we conduct a thorough evaluation of linear and nonlinear effects. The empirical findings show that BE and HTE are consistently strong predictors of trade intensity, especially in RLS and GLM estimates. Quantile regression further uncovers heterogeneity: HTE and RDC play a greater role in economies with already high trade integration, while BE and ICTSI matter more at lower quantiles, suggesting different policy levers across development levels.
The ML results reinforce these understandings, with GBM emerging as the most accurate base learner within the stacking ensemble, followed by RF. The ensemble model exhibits strong overall performance (high R2, low MAE and RMSE), validating the econometric results while improving prediction in more complex scenarios. Although ML models do not yield easily interpretable coefficients, their capability to seize interactions and nonlinearities makes them valuable forecasting tools for trade and innovation policy. National governments and international organizations could apply these techniques to monitor trade performance, run policy simulations, and prioritize interventions.
The broader literature confirms the differentiated trade–innovation dynamics identified in our analysis. In developing countries, technological innovation is a key driver of export development, with knowledge management, manufacturing capability, and service innovation identified as foundational dimensions that enhance competitiveness (Moughari and Daim, 2023). In China, a quasi-natural experiment based on World Trade Organization (WTO) accession shows that reducing trade policy uncertainty significantly increased invention patent applications, particularly in sectors facing larger declines in uncertainty, with firm responses varying by productivity, ownership, and export orientation (Liu and Ma, 2020). Broader structural factors also matter: in OECD manufacturing sectors, technological intensity, especially in research-intensive industries, has a stronger influence on specialization and competitiveness than unit labor costs, confirming the primacy of innovation in sustaining trade performance (Frantzen, 2008). In African economies, the link between innovation and environmental outcomes is nonlinear: innovation initially increases CO2 emissions but leads to reductions over time as cleaner technologies diffuse, while renewable energy and human capital contribute to long-run environmental improvements (Dauda et al., 2021). Persistence in both innovation and internationalization is also shown to yield stronger productivity gains, as firms continuously engaging in export markets and R&D are better positioned to internalize external knowledge flows (Iandolo and Ferragina, 2019). Finally, in China, the expansion of higher education significantly increased firms’ innovation output and export skill intensity, accounting for much of the observed growth in R&D intensity over the last two decades (Ma, 2024).
Policy recommendations are context-specific but aligned in their contribution to inclusive, sustainable, and innovation-driven growth. In developing economies, narrowing technology gaps and enhancing ICTU through targeted investments in infrastructure and education can foster digital inclusion and broader participation in trade. For BRICS countries, the focus should be on leveraging existing industrial capabilities to scale up innovation systems, reduce trade policy uncertainty, and promote skill-intensive exports, consistent with the observed effects of institutional reforms and education expansion on innovation performance. In OECD countries, where digital and R&D systems are already mature, policy should prioritize green innovation, support the clean energy transition, and enhance the resilience of GVCs, particularly in high-tech and research-intensive sectors. Across all contexts, coordinated trade and innovation strategies, such as reducing trade frictions, protecting intellectual property rights, and integrating SMEs into GVCs, are essential for advancing economic competitiveness, environmental sustainability, and social equity. These recommendations are grounded in both the empirical results of this study and the broader literature, which confirms the differentiated dynamics of trade and innovation across regions.
This study also contributes methodologically by integrating classical econometrics and ML, demonstrating how the two approaches can be used in a complementary manner. While econometric models provide clarity in terms of causal inference, hypothesis testing, and parameter interpretability, ML (particularly ensemble techniques such as boosting, bagging, and stacking) offers enhanced predictive performance, robustness to nonlinearities, and the ability to detect complex interactions that traditional methods may overlook. Importantly, we explicitly discuss how ML methods can be operationalized by policymakers and international institutions, especially through the use of ensemble models that deliver accurate, data-driven insights without relying on strong functional form assumptions.
Future research should incorporate real-time data and extend the analysis to subnational or sectoral levels. Evaluating how trade and innovation policies interact over longer horizons would also refine policy design, particularly for countries transitioning toward digital and green economies.
All potential errors and opinions are ours. Standard disclaimers apply.
Appendix
The diagram is a structured mind map centered around a bright green rectangular node labeled “Economic Factors,” from which multiple branches radiate to explain its influence on trade and high-tech exports. Extending upward from “Economic Factors” is a direct connection to “High-Tech Exports as percentage of Total Trade” in a similarly shaped box. This node continues upward to a wider rectangular node labeled “Structural Factors.” From “Structural Factors,” two separate upward branches point toward “Global Competition” and “Dependency on Foreign Markets,” respectively, both in rectangular green boxes. To the right of “High-Tech Exports as percentage of Total Trade,” another box in blue is labeled “Trade and High-Tech Exports Relationship,” which is further connected to a yellow box labeled “Trade as percentage of G D P.” To the right of “Economic Factors,” a horizontal link connects to “Increased Competition,” which further extends to “Wage Inequality.” Just below that, a downward-angled branch from “Economic Factors” leads to “Protectionist Policies,” and continuing downward is “Decline in Innovation.” Toward the lower side of “Economic Factors,” another branch leads to “Reliance on Traditional Exports,” and towards the lower left side of “Economic Factors,” another branch leads to “Trade Vulnerabilities,” and that node further links to “Economic Instability.” A separate leftward link from “Economic Factors” points toward “Environmental Degradation.” At the top left side of “Economic Factors,” a diagonal line connects to “R and D Investment Challenges,” which then connects to a different node also labeled “Economic Instability.”The relationship between Trade and HTE. Source: Authors’ own creation
The diagram is a structured mind map centered around a bright green rectangular node labeled “Economic Factors,” from which multiple branches radiate to explain its influence on trade and high-tech exports. Extending upward from “Economic Factors” is a direct connection to “High-Tech Exports as percentage of Total Trade” in a similarly shaped box. This node continues upward to a wider rectangular node labeled “Structural Factors.” From “Structural Factors,” two separate upward branches point toward “Global Competition” and “Dependency on Foreign Markets,” respectively, both in rectangular green boxes. To the right of “High-Tech Exports as percentage of Total Trade,” another box in blue is labeled “Trade and High-Tech Exports Relationship,” which is further connected to a yellow box labeled “Trade as percentage of G D P.” To the right of “Economic Factors,” a horizontal link connects to “Increased Competition,” which further extends to “Wage Inequality.” Just below that, a downward-angled branch from “Economic Factors” leads to “Protectionist Policies,” and continuing downward is “Decline in Innovation.” Toward the lower side of “Economic Factors,” another branch leads to “Reliance on Traditional Exports,” and towards the lower left side of “Economic Factors,” another branch leads to “Trade Vulnerabilities,” and that node further links to “Economic Instability.” A separate leftward link from “Economic Factors” points toward “Environmental Degradation.” At the top left side of “Economic Factors,” a diagonal line connects to “R and D Investment Challenges,” which then connects to a different node also labeled “Economic Instability.”The relationship between Trade and HTE. Source: Authors’ own creation
The map starts with the central node labeled “I C T Services Imports as percentage of Total Trade,” positioned in a light green rectangular box. This central node connects to several other nodes that represent key concepts and relationships related to trade and I C T services. From the central node, lines lead to multiple light green rectangular boxes. At the top of “I C T Services Imports as percentage of Total Trade,” a line connects to a light green box labeled “Enhanced Efficiency,” which is connected to the other four light green boxes labeled “Streamlined Supply Chain,” “Reduced Transaction Costs,” “Data Exchange,” and “Electronic Payments.” At the top right of “I C T Services Imports as percentage of Total Trade,” a line connects to a light green box labeled “Enhanced Productivity and Innovation,” followed by another box labeled “Economic Growth.” To the right of the central box, a line connects to the box labeled “Better connectivity.” To the bottom right side of the central box, a line leads to the box labeled “Improved Competitiveness,” which is followed by another box labeled “Economic Growth.” At the bottom of the central box, a line leads to a box labeled “Improved Communication.” At the bottom left side of the central box, a line leads to a box labeled “Increased Internet Penetration,” followed by a box labeled “Growth in Service Trade,” which further branches into two boxes labeled “Increased Imports” and “Increased Exports.” On the left side of the central box, a line leads to a blue box labeled “Trade and I C T Services Relationship,” which is connected to a yellow box labeled “Trade as percentage of G D P.” At the top left of the central box, a line leads to a box labeled “Economic Diversification,” which is further connected to another box labeled “Economic Growth.”Relationship between Trade and ICTSI. Source: Authors’ own creation
The map starts with the central node labeled “I C T Services Imports as percentage of Total Trade,” positioned in a light green rectangular box. This central node connects to several other nodes that represent key concepts and relationships related to trade and I C T services. From the central node, lines lead to multiple light green rectangular boxes. At the top of “I C T Services Imports as percentage of Total Trade,” a line connects to a light green box labeled “Enhanced Efficiency,” which is connected to the other four light green boxes labeled “Streamlined Supply Chain,” “Reduced Transaction Costs,” “Data Exchange,” and “Electronic Payments.” At the top right of “I C T Services Imports as percentage of Total Trade,” a line connects to a light green box labeled “Enhanced Productivity and Innovation,” followed by another box labeled “Economic Growth.” To the right of the central box, a line connects to the box labeled “Better connectivity.” To the bottom right side of the central box, a line leads to the box labeled “Improved Competitiveness,” which is followed by another box labeled “Economic Growth.” At the bottom of the central box, a line leads to a box labeled “Improved Communication.” At the bottom left side of the central box, a line leads to a box labeled “Increased Internet Penetration,” followed by a box labeled “Growth in Service Trade,” which further branches into two boxes labeled “Increased Imports” and “Increased Exports.” On the left side of the central box, a line leads to a blue box labeled “Trade and I C T Services Relationship,” which is connected to a yellow box labeled “Trade as percentage of G D P.” At the top left of the central box, a line leads to a box labeled “Economic Diversification,” which is further connected to another box labeled “Economic Growth.”Relationship between Trade and ICTSI. Source: Authors’ own creation
At the top, the box labeled “Trade as Percentage of G D P” is positioned centrally, with an arrow pointing to another box, “I C T Use,” placed directly below it. The box then branches into two main sections: “Structural Factors,” positioned on the far left side, and “Economic Factors,” positioned just below the box of “I C T Use.” The box of “Economic Factors” is further divided into six boxes at the bottom, which are arranged horizontally, labeled from left to right as “Reliance on Traditional Exports,” “Policy Focus on Traditional Industries,” “Economic Dependencies,” “Regulatory Challenges,” “Economic Volatility,” and “Focus on Short-Term Gains.” These six boxes and the far left box, “Structural Factors,” lead to the final box at the bottom center, which is labeled “Limits I C T Development.” A double-headed arrow is connected between the boxes “I C T Use” and “Limits I C T Development.” The entire flowchart uses boxes colored in soft purple.The relationship between Trade and ICTU. Source: Authors’ own creation
At the top, the box labeled “Trade as Percentage of G D P” is positioned centrally, with an arrow pointing to another box, “I C T Use,” placed directly below it. The box then branches into two main sections: “Structural Factors,” positioned on the far left side, and “Economic Factors,” positioned just below the box of “I C T Use.” The box of “Economic Factors” is further divided into six boxes at the bottom, which are arranged horizontally, labeled from left to right as “Reliance on Traditional Exports,” “Policy Focus on Traditional Industries,” “Economic Dependencies,” “Regulatory Challenges,” “Economic Volatility,” and “Focus on Short-Term Gains.” These six boxes and the far left box, “Structural Factors,” lead to the final box at the bottom center, which is labeled “Limits I C T Development.” A double-headed arrow is connected between the boxes “I C T Use” and “Limits I C T Development.” The entire flowchart uses boxes colored in soft purple.The relationship between Trade and ICTU. Source: Authors’ own creation
The flowchart begins at the top with a central box labeled “Trade as Percentage of G D P.” From this central box, three arrows branch out to the left, middle, and right. On the left side, the arrow leads to a box labeled “Rapid Trade Liberalization,” which further points to another box below it labeled “Regulatory Inefficiencies.” The middle arrow leads to a box labeled “Deterioration of Business Environment.” On the right side of the top box, another arrow leads to “Economic Vulnerabilities,” which splits into two main paths: one path leads downward to “Deterioration of Business Environment,” and another path leads to “Governance Issues,” which then points to a box labeled “Corruption,” and the other path continues downward to “Deterioration of Business Environment.” This box has three arrows branching out below it. The left arrow points to a box labeled “Heightened Competition,” the middle to “Regulatory Challenges,” and the right to “Environmental Impacts.” Under “Environmental Impacts,” two branches are leading to “Environmental Degradation” and “Race to the Bottom.” From these boxes, the arrows lead to the final box labeled “Increased C O 2 Emissions.” The entire flowchart uses boxes colored in soft purple.The relationship between Trade and BE. Source: Authors’ own creation
The flowchart begins at the top with a central box labeled “Trade as Percentage of G D P.” From this central box, three arrows branch out to the left, middle, and right. On the left side, the arrow leads to a box labeled “Rapid Trade Liberalization,” which further points to another box below it labeled “Regulatory Inefficiencies.” The middle arrow leads to a box labeled “Deterioration of Business Environment.” On the right side of the top box, another arrow leads to “Economic Vulnerabilities,” which splits into two main paths: one path leads downward to “Deterioration of Business Environment,” and another path leads to “Governance Issues,” which then points to a box labeled “Corruption,” and the other path continues downward to “Deterioration of Business Environment.” This box has three arrows branching out below it. The left arrow points to a box labeled “Heightened Competition,” the middle to “Regulatory Challenges,” and the right to “Environmental Impacts.” Under “Environmental Impacts,” two branches are leading to “Environmental Degradation” and “Race to the Bottom.” From these boxes, the arrows lead to the final box labeled “Increased C O 2 Emissions.” The entire flowchart uses boxes colored in soft purple.The relationship between Trade and BE. Source: Authors’ own creation
Notes
Available at: https://data.worldbank.org/indicator/NE.TRD.GNFS.ZS downloaded as bulk data in July 2024.
Available at: https://www.wipo.int/global_innovation_index/en/, downloaded as bulk data in July 2024. The GII aggregates indicators from various underlying sources, some of which provide information for countries and years not included in the official GII rankings. This may explain discrepancies in country coverage between our dataset and the published GII reports.
The dataset was split into training (70%) and testing (30%) sets for the main ML analysis, and into 80%/20% for the ensemble model to ensure a larger training base for optimizing multiple base learners and improving generalization performance during stacking.
All models are trained on the training dataset and validated using the test dataset. Linear regression is estimated using the lm() function, boosting is implemented via xgboost() with cross-validated tuning, and bagging is performed using randomForest() for bagging with 100 trees, using the RF model serving as a proxy.
Each PDP was generated for individual features (excluding fixed effects).
The stacking ensemble was implemented using the H2O framework via h2o.stackedEnsemble() on normalized data, combining RF, GBM, and GLM base learners trained with five-fold cross-validation. Model performance was evaluated on a held-out test set, and both base model feature importance and meta-learner (GLM) coefficients were extracted to assess variable relevance.
Acronym and Definition
- B2B
Business to Business
- B2C
Business to Customer
- B2G
Business to Government
- BE
Business Environment
- BRICS
Brazil, Russia, India, China, South Africa
- BRI
China’s Belt and Road Initiative
- CO2
Carbon Dioxide
- FDI
Foreign Direct Investment
- FTZs
Free Trade Zones
- G7
Group of Seven
- GBM
Gradient Boosting Machine
- GDP
Gross Domestic Product
- GII
Global Innovation Index
- GLM
Generalized Linear Model
- GMM
Generalized Method of Moments
- GVCs
Global Value Chains
- HTE
High-Tech Exports
- ICT
Information and Communication Technology
- ICTSI
ICT Services Imports
- ICTU
ICT Use
- MAE
Mean Absolute Error
- ML
Machine Learning
- MSE
Mean Squared Error
- OECD
Organization for Economic Co-operation and Development
- OLS
Ordinary Least Squares
- PDPs
Partial Dependence Plots
- R&D
Research and Development
- RDC
Global R&D Companies
- RF
Random Forest
- RLS
Robust Least Squares
- RMSE
Root Mean Squared Error
- SMEs
Small and Medium Enterprises
- WIPO
World Intellectual Property Organization
- WTO
World Trade Organization

