Six machine learning methods (linear regression, logistic regression, extreme gradient boosting (XGBoost), support vector machine, K-nearest neighbours and artificial neural network) were used to predict/classify the hydraulic conductivity of conventional sodium bentonite (Na-B) geosynthetic clay liners (GCLs) to saline solutions or leachates. Data were collected from the literature and randomly divided into two groups – that is, 80% of the data were used to train machine learning models and the rest, 20%, were applied to evaluate model performance. Features that are known to affect the hydraulic conductivity of Na-B GCLs (e.g. mass per unit area of GCLs, monovalent and divalent cations, ionic strength (I), relative abundance of monovalent to divalent cations (RMD), swell index and effective stress) were employed to predict/classify the hydraulic conductivity of Na-B GCLs. Comparative analyses were conducted with seven subsets corresponding to the combination of different features, and the best model was determined through cross-validation. The results showed that XGBoost consistently had the best performance among all methods over all subsets of features for both regression and classification analyses. Subset 4, using the swell index, I, RMD, I2 × RMD, monovalent cations, divalent cations, effective stress and mass per unit area as features, outperformed all other six subsets in both regression analysis (R2 = 0.826) and classification analysis (accuracy = 0.887) in the out-of-sample tests.

CD

concentration of divalent cations (mM)

CM

concentration of monovalent cations (mM)

I

ionic strength (mM)

K

hydraulic conductivity of sodium bentonite (Na-B) geosynthetic clay liners (GCLs) (m/s)

KDI

hydraulic conductivity of Na-B GCLs to deionised water (m/s)

Km'¯

mean value of Km' (log(m/s))

Kmi'

measured hydraulic conductivity (log(m/s))

Kp'¯

mean value of Kpi' (log(m/s))

Kpi'

predicted hydraulic conductivity using machine learning models (log(m/s))

N

total number of observations in the data set

R2

coefficient of determination

xmax

maximum value of the variable

xmin

minimum value of the variable

Zi

valence of the ith ion in solution

σ

effective stress (kPa)

Geosynthetic clay liners (GCLs) consist of a thin layer (5–10 mm) of sodium bentonite (Na-B) sandwiched between two layers of geotextiles. GCLs are widely used as hydraulic barriers in waste containment facilities due to their low permeability; thinness, which saves airspace; and ease of installation (Chen et al., 2018; Jo et al., 2001, 2005; Li et al., 2021a, 2023; Shackelford et al., 2000; Zainab et al., 2021). The primary mineral in Na-B GCLs is sodium (Na) montmorillonite, which has a large surface area, a high cation-exchange capacity and a high swelling potential (Bradshaw and Benson, 2014; Jo et al., 2001, 2005; Kolstad et al., 2004a; Lee et al., 2005; Li and Tian, 2022; Shackelford et al., 2000). Osmotic swelling of sodium montmorillonite reduces the pore size, resulting in a tortuous pathway for water flow and the low hydraulic conductivity of GCLs (e.g. <1 × 10−10 m/s) (Ashmawy et al., 2002; Athanassopoulos et al., 2015; Chen et al., 2018, 2019a; Gastelo et al., 2023; Guyonnet et al., 2005; Setz et al., 2017; Tian et al., 2016, 2019). However, the swelling of Na-B can be inhibited when exposed to leachates with high ionic strength (e.g. I > 300 mM), predominantly polyvalent cations and/or extreme pH conditions, leading to a large intergranular pore space and a high hydraulic conductivity of GCLs (>1 × 10−10 m/s) (Chen et al., 2019a; Di Emidio et al., 2015; Jo et al., 2001, 2005; Shackelford et al., 2000; Tian et al., 2016, 2019). It is important to determine whether a Na-B GCL can maintain low hydraulic conductivity to landfill leachates before applying it in the field.

Tremendous research has been conducted to investigate the hydraulic conductivity of Na-B GCLs to saline solutions or leachates from solid waste disposal facilities (Bradshaw et al., 2013, Chen et al., 2019a; Jo et al., 2001, 2005; Li et al., 2021b; Tian and Li, 2021; Tian et al., 2016; Zainab et al., 2021; Zhao et al., 2023). Tests are required to meet the termination criteria (e.g. hydraulic and chemical equilibria) according to ASTM D 6766 (ASTM, 2020) to reflect the long-term hydraulic conductivity of GCLs, which typically takes a very long time (e.g. months) (Jo et al., 2001, 2005; Scalia et al., 2014). Jo et al. (2005) reported that it took approximately 215 days for the hydraulic conductivity test of Na-B GCLs to 5 mM calcium chloride (CaCl2) solution to reach both hydraulic and chemical equilibria.

Alternatively, the swell index (SI) test (ASTM, 2019) is used as a quick index test to evaluate the chemical compatibility of Na-B GCLs with leachates. The hydraulic conductivity of Na-B GCLs is inversely related to the SI of bentonite – that is, Na-B GCL with a high SI in permeant solution can maintain a low hydraulic conductivity (Chen et al., 2018; Jo et al., 2001, 2005; Katsumi et al., 2008; Kolstad et al., 2004a; Lee et al., 2005). However, the relationship between the hydraulic conductivity of Na-B GCLs and the SI value is not unique. Na-B GCLs with the same SI may have different hydraulic conductivities. For example, Cruz (2021) reported that the hydraulic conductivity of Na-B GCLs to 10 mM calcium chloride was 5.8 × 10−11 m/s with a SI of 24.0 ml/2 g. Chen et al. (2018) reported that Na-B GCLs showed a high hydraulic conductivity (e.g. 2.0 × 10−9 m/s) when permeated with synthetic coal combustion product (CCP) leachate (I = 48 mM), although the SI of Na-B was 24.0 ml/2 g.

Statistical models were also developed to predict the hydraulic performance of GCLs based on laboratory measurements. Kolstad et al. (2004a) proposed a linear regression equation for predicting the hydraulic conductivity of Na-B GCLs using ionic strength (I) and the relative abundance of monovalent to divalent cations (RMD) in the leachates as predictors. The proposed equation was valid for predicting the hydraulic conductivity of GCLs to leachates with I = 0.05–0.50 M and RMD < 2.0 mM1/2 at an effective stress of 20 kPa. Katsumi et al. (2007, 2008) proposed an equation for predicting the hydraulic conductivity of Na-B GCLs as a function of the SI. As discussed previously, the application of the SI as the sole factor may not accurately predict the hydraulic conductivity of Na-B GCLs. Empirical models were developed based on limited laboratory measurements, valid only to a limited extent (Katsumi et al., 2007, 2008; Kolstad et al., 2004a). Other factors such as effective stress and mass per unit area (MPUA) of Na-B GCLs also influence the hydraulic conductivity of Na-B GCLs. The empirical models may not provide accurate predictions of the hydraulic conductivity of Na-B GCLs permeated with saline solutions and leachates considering only the leachate properties or SI.

To date, machine learning (ML) approaches have been widely adopted in different domains in civil engineering (Chao et al., 2023a; Elbisy, 2015; Goli et al., 2022; Jiang et al., 2021; Kardani et al., 2022; Ozcoban et al., 2018; Rogiers et al., 2012; Salami et al., 2022; Samui, 2008; Shahin et al., 2002; Zhang et al., 2021). Many ML algorithms, such as artificial neural networks (ANNs), support vector machines (SVMs) and K-nearest neighbours (KNN), have been used to estimate the hydraulic conductivity of soil. Rogiers et al. (2012) applied multiple linear regression (MLR) and ANN models to estimate the saturated hydraulic conductivity of soil using grain-size distribution data. The ANN model proved to be more accurate than the MLR model, with greater accuracy in predicted hydraulic conductivity (e.g. R2: 0.93 against 0.89). Das et al. (2012) applied ANN and SVM models to predict the field hydraulic conductivity of clay liners based on in situ test results. The input of the models included the liquid limit, plastic limit, percentage fines content, moisture content, dry density, weight of the compactor, lift thickness and number of lifts. Raja and Shukla (2021) combine the grey-wolf optimisation (GWO) and ANN algorithms (ANN–GWO) to predict the settlement of a geosynthetic-reinforced soil foundation (GRSF). The results showed that the proposed model (ANN–GWO) predicts the settlement of a GRSF with higher accuracy compared with an ANN model (root mean square error: 0.612 against 0.774; R2: 0.962 against 0.943).

The objective of this paper is to explore and develop an ML-based framework to predict the hydraulic conductivity of Na-B GCLs to leachates. Six ML algorithms – namely, linear regression, logistic regression, extreme gradient boosting (XGBoost), SVM, KNN and ANN – are applied and compared in this study. Two types of target (output) variables are investigated according to two tasks: (a) prediction of the hydraulic conductivity of GCLs and (b) classification of the status of hydraulic conductivity based on the threshold value (i.e. ‘success’ if ≤1 × 10−10 m/s; ‘failure’, otherwise). The input features of the models include leachate properties (I, RMD, concentrations of monovalent and divalent cations), MPUA of GCLs, SI and effective stress. The model developed in this study can be used to select appropriate Na-B GCLs to manage leachates from solid waste disposal facilities.

Swell index

Numerous studies have shown that there exist relationships between the SI and hydraulic conductivity of Na-B GCLs (Chen et al., 2018; Guyonnet et al., 2005; Jo et al., 2001, 2005; Katsumi et al., 2008; Kolstad et al., 2004a; Lee et al., 2005; Wireko et al., 2020). For Na-B GCLs, a hydraulic conductivity <1.0 × 10−10 m/s generally correlates with SI >20 ml/2 g at low effective stress and without prehydration (Scalia et al., 2018).

However, the relationship between the hydraulic conductivity of Na-B GCLs and the SI is not trivial at low effective stress, as shown in Figure 1 (which compiles all available data in the literature). The hydraulic conductivity of Na-B GCLs to saline solutions or leachates is higher than 1 × 10−10 m/s when the swelling of bentonite is < 12 ml/2 g, whereas Na-B GCLs can maintain a low hydraulic conductivity (e.g. <1 × 10−10 m/s) when the SI is higher than 30 ml/2 g. When the SI of Na-B falls into the range between 12 and 30 ml/2 g, no conclusion can be made on whether Na-B GCLs can maintain a hydraulic conductivity lower than the threshold or not.

Empirical model

Kolstad et al. (2004a) developed an empirical equation for predicting the hydraulic conductivity of Na-B GCLs (MPUA = 4.3 kg/m2) based on the ionic strength (I) and RMD of the inorganic solution. The equation was developed based on 31 hydraulic conductivity test observations (data set A from Table 1) using stepwise regression, as shown in the following equation:

1
Table 1.

Empirical models for predicting the hydraulic conductivity of Na-B GCLs

Empirical modelData setNumber of dataPermeant solutionsEffective stress: kPaMPUA: kg/m2abcdR2
logK/logKDI=a+b×I+c×RMD+d(I2)×RMD
(Kolstad et al., 2004a)
Aa31Saline solutions204.30.965−0.9760.07970.2510.967
Bb82Saline solutions, synthetic CCP leachates, MSW leachates203.6–5.10.880−0.7400.1030.2210.514
log(K/c)=exp[a(SI+B)]
(Katsumi et al., 2007)
Cc40DI water, saline solutions, MSW leachate20–304.73−0.3108.6903.09 × 10−11N/A0.830
 Dd248DI water, saline solution, MSW, CCP, LLW leachates, bauxite liquor20–303.6–6.4−0.07028.2205.20 × 10−12N/A0.595

Equation 1 is linear in I and RMD, and the product I2 × RMD reflects that the sensitivity to RMD varies non-linearly with ionic strength. The equation is valid for I = 0.05–0.50 M and RMD < 2.0 M1/2 and at 20 kPa (Kolstad et al., 2004a).

A new data set B (containing 82 hydraulic conductivity tests) was created by combining the data reported by Kolstad et al. (2004a) (31 tests) and additional data set (51 tests) from the literature (Table 1). The fitting parameters of Equation 1 for data set B were obtained by regression analysis, as shown in Table 1. R2 decreased to 0.514. Most of the predicted hydraulic conductivity of Na-B GCLs for data set B fell within two orders of magnitude in comparison with the measured value (Figure 2(a)). The low prediction accuracy reflected the limitation of Equation 1, indicating that factors other than ionic strength and RMD might affect the hydraulic conductivity of GCLs.

Figure 2.

Predicted hydraulic conductivity plotted against measured hydraulic conductivity (a) as a function of ionic strength and RMD using the method proposed by Kolstad et al. (2004a) and (b) as a function of the SI using the method proposed by Katsumi et al. (2007) 

Figure 2.

Predicted hydraulic conductivity plotted against measured hydraulic conductivity (a) as a function of ionic strength and RMD using the method proposed by Kolstad et al. (2004a) and (b) as a function of the SI using the method proposed by Katsumi et al. (2007) 

Close modal

Katsumi et al. (2007, 2008) proposed another empirical equation for predicting the hydraulic conductivity Na-B GCLs as a function of the SI. The equation was developed based on 40 hydraulic conductivity tests at 29.4 kPa effective stress (Table 1, data set C) (Katsumi et al., 2007, 2008), as shown in the following equation:

2

Equation 2 can be used under an effective stress of 20–30 kPa. A new data set D (248 tests) was compiled including the hydraulic conductivity of Na-B GCLs to saline solutions and leachates at 20–30 kPa effective confining stress (Table 1). The R2 of Equation 2 using data set D decreased to 0.595. The predicted hydraulic conductivity of Na-B GCLs primarily fell within 1000 times (three orders of magnitude) to the measured value, as shown in Figure 2(b). The low R2 of Equation 2 illustrates that the prediction of hydraulic conductivity of the Na-B GCLs as a function of the SI is insufficient and inaccurate, as discussed in the previous section.

The two empirical models showed poor performance in predicting the hydraulic conductivity of Na-B GCLs permeated with saline solutions and leachates and thus may not provide accurate predictions for selecting appropriate Na-B GCLs to manage landfill leachates. It is important to develop a more accurate model for predicting the hydraulic conductivity of Na-B GCLs. Therefore, this study developed an ML-based framework for predicting the hydraulic conductivity of Na-B GCLs to leachates using MPUA, leachate properties (e.g. I and RMD), effective stress (σ) and SI.

ML has been widely applied in geotechnical engineering due to its capability and flexibility to characterise complex relationships between variables. Supervised ML approaches can be broadly categorised into two classes – regression and classification, where the regression model aims to predict the numerical value of the target variable and the classification model aims to predict the categorical level of the target variable. In this study, six popular ML approaches were adopted and compared – namely, linear regression, logistic regression, XGBoost, SVM, KNN and ANN. Linear regression can be used only for regression tasks, logistic regression is used only for classification tasks, while the remaining four can be used for both regression and classification tasks. To be self-contained, a brief introduction to XGBoost, SVM, KNN and ANN are described in the following sections, while the descriptions of linear and logistic regression models are skipped since they are well known.

Extreme gradient boosting

XGBoost is an ensemble learning algorithm based on classification and regression trees (Carts). A single Cart is considered a weak ML model – that is, it is not a competitive model with a high prediction accuracy (James et al., 2013). To improve the model performance, many ensemble ML models are developed based on Carts using the bagging or boosting approach. For instance, a random forest is developed based on the bagging approach and consists of multiple Carts. Further developed from random forests, XGBoost is an ensemble learning algorithm that uses a gradient boosting framework with higher computational efficiency than a random forest model with a weighted quantile design (Chen and Guestrin, 2016). The idea of XGBoost is to add trees continuously and perform feature splitting to grow a tree. The new tree learns a new Cart function to fit the residual of the last prediction. When the training process is completed, a specified number of trees are obtained as outcomes. Each of the learned trees corresponds with a leaf node and a prediction score. The final predicted value is calculated as the weighted sum of the scores across all trees. The XGBoost method shows advantages over logistic regression in terms of accuracy and predictive power. As an ensemble method that combines multiple decision trees, XGBoost captures non-linear patterns in the data that a logistic regression model may overlook. Furthermore, through gradient boosting, XGBoost reduces its error rate with each iteration, leading to more precise predictions. Additionally, XGBoost efficiently handles large, high-dimensional data sets without feature selection, making it suitable for demanding ML tasks. Although its black-box nature may be perceived as a drawback, the exceptional accuracy and predictive power of XGBoost outweigh this lack of interpretability. Relative influence was employed to obtain the importance for each predictor. At each split in each tree, the improvement in the split criterion (mean squared error (MSE) for regression, accuracy for classification) was computed. Then, the average of the improvement made by each variable across all the trees can be calculated. The variables with the largest average decrease in MSE or increase in accuracy were considered most important.

Support vector machine

An SVM is an ML model that can be utilised for classification and regression analysis (Cortes and Vapnik, 1995). The basic idea of SVMs comes from the maximum margin classifier, which is a classification method based on selecting the hyperplane that can maximise the interval between each class. More specifically, for two-class classification analysis, the SVM algorithm maps training observations to points in space and then maximises the hyperplanes between the two categories through a fixed kernel function. The new input observations are mapped into the same space, and which side of the gap the new mapped points belong to is analysed. SVMs can also be applied in regression analysis resulting from the support vector regression (SVR) model. Unlike linear regression models that minimise the error between the actual and predicted values, SVR can fit the optimal line within a threshold of values called the epsilon-insensitive tube. SVMs can capture non-linear relationships among the variables by adopting different kernel functions, such as linear, polynomial and radial basis kernels (Chao et al., 2023b).

K-nearest neighbours

KNN is a non-parametric ML approach that can be used for regression and classification analysis, where K is a positive integer specified by the user (Anava and Levy, 2016). The idea behind KNN is to predict/classify the observation according to its most similar neighbours. For both cases, the input training data are assumed to contain K-closest neighbours. For the case of classification, the output object is classified by the plurality vote of the neighbours. The output case is predicted by the average value of its KNN in the training data set for the regression case. The number of K is the only hyperparameter that needs to tune in the model-training process.

Artificial neural networks

ANNs are a supervised learning algorithm that is the basis of deep learning research (Chen et al., 2019b). The name of ANNs are inspired by the human brain, mimicking the way that biological neurons signal to one another. ANNs comprise node layers, including an input layer, at least one hidden layer and an output layer. Each node in the network has a threshold value and weight. If the output of a node is over the given threshold value, the node become active and passes the data to the next layer. If the node cannot be activated by the output, the node will not send data to the next layer. The output of a node is obtained by the activation function based on the inputs of the node. For regression analysis, the output layer provides continuous results based on one node. For classification analysis, the output is classified in one of the classes.

The proposed framework for predicting hydraulic conductivity is shown in Figure 3. Two different approaches were considered in the proposed procedure: one was regression and the other one was classification. The two approaches were conducted independently. The procedure consists of three modules – namely, data source integration, model selection and out-of-sample performance evaluation. The tasks to be completed in each module are summarised as follows.

  • Data source integration. Multiple data sets were collected from the existing literature and were integrated according to the feature variables.

  • Model selection. Different ML models were constructed using training data, and the best model (with features) was selected through a k-fold cross-validation (CV) procedure.

  • Performance evaluation. The value (or success/failure status) of the hydraulic conductivity of Na-B GCLs to saline solution or leachate was predicted using the holdout testing data to evaluate the proposed model performance with the selected model.

Figure 3.

Procedure for predicting hydraulic conductivity (K)

Figure 3.

Procedure for predicting hydraulic conductivity (K)

Close modal

The data used in this study are collected from 22 papers from the years 2000 to 2022, including those on CCP, municipal solid waste (MSW), low-level radioactive waste and MSW incineration ash landfill leachates, which represent the typical leachates in landfills (Benson and Meer, 2009; Bradshaw and Benson, 2014; Bradshaw et al., 2016; Chen et al., 2018; Cruz, 2021; EPRI, 2014; Jo et al., 2001; Katsumi et al., 2007; Kolstad et al., 2004a; Lee et al., 2005; Li et al., 2021a, 2021b; Meer and Benson, 2007; Salihoglu, 2015; Scalia et al., 2014; Setz et al., 2017; Tian et al., 2016, 2019; Wang et al., 2019; Wireko and Abichou, 2021; Wireko et al., 2022; Zainab et al., 2021). Table 2 summarises the minimal, median, mean and maximal values of each variable. The hydraulic conductivity of Na-B GCLs in the data set ranged from 1.8 × 10−12 to 5.2 × 10−6 m/s. The SI in the data set ranged from 2.0 to 35.6 ml/2 g, and the MPUA of Na-B GCLs ranged from 3.6 to 6.4 kg/m2. The effective stress in the data set ranged from 20 to 500 kPa, and most of the effective stresses were lower than 40 kPa (63.9%).

Table 2.

Summary of data variables

VariableUnitMin.MedianMeanMax.
Hydraulic conductivity, Km/s1.8 × 10−121.1 × 10−91.8 × 10−75.2 × 10−6
Monovalent cationsmM0.062.6249.54638.7
Divalent cationsmM0.018.589.21000.0
Ionic strength, ImM0.02171.7496.24676.0
RMDM0.50.000.6514.59100.00a
Effective stress, σkPa20.020.097.6500.0
MPUAkg/m23.64.04.26.4
SIml/2 g2.013.514.335.6
a

100 indicates infinity observations of RMD (∞)

I, ionic strength; MPUA, mass per unit area; RMD, relative abundance of monovalent to divalent cations; SI, swell index; σ, effective stress

The entire data set consisting of 308 observations from field experiments of infiltration process were separated randomly into two groups, which were denoted as training and testing data sets. For the train–test split ratio, there was no clear guidance on what ratio was best or optimal for a given data set (Joseph, 2022). A widely used train–test split ratio was 80:20, which meant 80% of the data were for training and 20% are for testing. Other train–test split ratios such as 70:30, 60:40 and 50:50 were also used in practice. The 80:20 split draws its justification from the well-known Pareto principle (Sanders, 1987). The selection of the size of the training data set primarily depended on the complexity of the model. Previous research has shown that the number of the data sets should be at least ten times that of the features for small models (Harrell et al., 1996; Peduzzi et al., 1996). In the present study, the training data set was composed of 246 observations, and the number of features ranged from 2 to 8. The size of the training data set in this study met this rule of thumb. In this sense, the 80:20 ratio was adopted in this study. The training data set contained 246 observations (∼80% of the total data), while the testing data had 62 observations (the remaining ∼20% of the total data) (Sanders, 1987). To validate the model generalisation ability, the 70:30 ratio was also compared in this study. The input parameters (features) included monovalent cations, divalent cations, I, RMD, I2 × RMD, effective stress, MPUA and SI. Both regression and classification models were investigated. For the regression model, the numerical value of hydraulic conductivity of GCL was used as the target variable. For the classification model, the status of hydraulic conductivity – that is, success (when the hydraulic conductivity was ≤1 × 10−10 m/s) or failure (when the hydraulic conductivity was >1 × 10−10 m/s) – was employed as the target variable.

Data preprocessing was conducted before the training procedure to improve the performance of ML models. When the concentration of monovalent cations in calcium chloride solution and that of divalent cations in sodium chloride (NaCl) solution equal zero, the zero values of monovalent cations and divalent cations in the leachate composition were replaced by 0.001. Logarithm transformation of hydraulic conductivity, monovalent cations, divalent cations and ionic strength was performed taken to convert the data to be more normally distributed. Then, to rescale all independent variables into the same range, min–max normalisation was performed using the following equation:

3

where xmin and xmax were the minimum and maximum values of the variable, respectively. Finally, data sets were randomly divided into two subsets: 80% as the training set and 20% as the holdout testing set. The training set was used to build and tune ML models, and the testing set was used to evaluate the performance of the built models.

The optimal model for each ML algorithm was studied through k-fold CV during the training process. At the beginning of the k-fold CV approach, the training data were randomly and evenly divided into k folds. Then, the first fold was considered the validation set and the remaining k − 1 folds were used to build a learning model. The validation fold was employed to evaluate the performance of the obtained model. The process was repeated k times. Each time, a different fold of data was considered a validation set and a model was trained with the remaining folds. For each round, the MSE (or accuracy) was calculated in the validation data. Then, the k-fold CV estimate was obtained by taking the average over the k MSEs (or accuracy). The value of k was set as 5 in this study (i.e. fivefold CV).

Two assignments were considered in the model-training process: feature selection and parameter tuning. Feature selection could improve the model prediction accuracy and model interpretability. In this study, the subset selection approach was used to conduct feature selection progress. Based on the previous literature, key variables that are known to affect the hydraulic conductivity of Na-B GCLs to saline solution and leachate were divided and combined to create feature subsets, as listed in Table 3. The variables can be classified into four categories: (a) GCL properties (MPUA); (b) leachate properties (e.g. monovalent and divalent cations, I, RMD); (c) effective confining stress; and (d) SI. Overall, seven subsets corresponding to different combinations of features were evaluated in this study, which can be divided into three comparative categories. Subsets 1 and 2 consider MPUA effective confining stress with various combinations of leachate properties (Table 3). The comparison between subsets 1 and 2 can illustrate the significance of parameters within leachate properties. In comparison with subsets 1 and 2, one additional parameter (e.g. SI) was evaluated in subsets 3–5. The comparison between these two sets can reveal the importance of SI on predicting the hydraulic conductivity of Na-B GCLs. Subsets 6 and 7 were compared to evaluate the accuracy of the prediction of the hydraulic conductivity of Na-B GCLs based on the three parameters of SI, MPUA and effective confining stress.

Table 3.

Subsets of the predictors

Subset 1MPUA, σ, monovalent cations, divalent cations, I, RMD, I2 × RMD
Subset 2MPUA, σ, I, RMD
Subset 3MPUA, σ, monovalent cations, divalent cations, I, RMD, SI
Subset 4MPUA, σ, monovalent cations, divalent cations, I, RMD, I2 × RMD, SI
Subset 5MPUA, σ, I, RMD, SI
Subset 6MPUA, σ, SI
Subset 7σ, SI

I, ionic strength; MPUA, mass per unit area; RMD, relative abundance of monovalent to divalent cations; SI, swell index; σ, effective stress

Moreover, when training the model, some hyperparameters of the ML models cannot be directly estimated through an algorithm. A commonly used method to obtain the optimal hyperparameters is grid search, which can exhaustively search a designated subset of the hyperparameter space of the ML model. In the training process, a grid search algorithm was used to find the optimal hyperparameters for each ML model through fivefold CV over each proposed feature subset (James et al., 2013). The detailed ML model parameter settings and employed packages are listed in the  Appendix (which also includes Tables 8–12).

Evaluation metric for the regression model

The performance of regression models was evaluated using five metrics: MSE, coefficient of determination (R2), mean absolute error (MAE), mean absolute percentage error (MAAPE) and refined Willmott index (RWI). Denoting Kpi'=log10(Kpi) and Kmi'=log10(Kmi), MSE and R2 are obtained as follows:

4
5
6
7
8

where N is the number of observations; Kpi' is the predicted result using ML models; Kp'¯ represents the mean value of Kpi'; Kmi' is the measured hydraulic conductivity; and Km'¯ represents the mean value of Kmi'. The best model performance is marked by high R2 and RWI and low MSE, MAE and MAAPE values.

Evaluation metric for the classification model

Two metrics – accuracy and area under the curve (AUC) – were used to evaluate the performance of the classification models. Accuracy denotes the overall proportion of observations that are correctly classified. The accuracy metric is primarily used to evaluate performance and select the best model. The mathematical expression for accuracy is as follows:

9

where N represents the total number of observations in the data set. The true positive (TP) expresses the number of positive class observations correctly classified into the positive class by the ML model (Kmeasured, Kpredicted ≤ 1 × 10−10 m/s). The true negative (TN) indicates negative class observations correctly classified into the negative class (Kmeasured, Kpredicted > 10−10 m/s).

The AUC value is also reported to measure the overall performance of the classification model, which represents the area under the receiver operating characteristic curve. The AUC is an evaluation metric that works only for two-class classification analysis. For binary samples (positive and negative), the AUC value represents the probability that the chance that a positive observation is classified into a positive class is larger than the chance that this observation is classified as a negative class. The range of the AUC is between 0.5 and 1. When the AUC value is equal to 1, the model can perfectly classify the given samples. When the AUC value is equal to 0.5, the model has no power to conduct prediction; and therefore, has no application value.

The model performance of all training models (e.g. MSE, R2, MAE, MAAPE, RWI) across subsets 1–7 is shown in Figure 4. The models with advanced algorithms such as XGBoost, SVM, KNN and ANN outperformed the linear regression model. For each subset, the XGBoost model had consistently the best performance by giving the smallest MSE, MAE and MAAPE and highest RWI and R2 in comparison with those of the other four models. For instance, the MSE of XGBoost for subset 3 was 0.126, which was 1127.8, 761.9, 438.9% and 513.5% lower than those of linear regression, SVM, KNN and ANN, respectively. The R2 of subset 3 by the XGBoost method was 0.962, which was 17.5–40.8% higher than those of the other four methods. The RWI of subset 3 was also 18.5–63.9% higher than those of the other four methods. Therefore, these results showed that XGBoost was the most accurate method in predicting the hydraulic conductivity of Na-B GCLs to saline solutions or leachates. Hereafter, the discussion focuses on the comparison among subsets 1–7 using XGBoost methods. The results for the 70:30 ratio show similar (although slightly poorer) performance in the training data set as those for the 80:20 ratio. The subsequent discussion is based on the 80:20 ratio.

Figure 4.

Model performance on the training data set by regression analysis using linear regression, XGB, SVM, KNN and ANN: (a) MSE; (b) R2; (c) MAE; (d) MAAPE; (e) RWI

Figure 4.

Model performance on the training data set by regression analysis using linear regression, XGB, SVM, KNN and ANN: (a) MSE; (b) R2; (c) MAE; (d) MAAPE; (e) RWI

Close modal

The model performance of the training data set with subsets 1–7 by the XGBoost method is shown in Table 4. Subsets 1 and 2, considering MPUA, leachate properties and effective stress factor, illustrated great performance in terms of low MSE, MAE and MAAPE and high R2 and RWI. Subsets 3–5 contained the same features as subsets 1 and 2 correspondingly but with one additional feature - the SI. With the SI, the performance of subsets 3–5 improved on that of subsets 1 and 2 with MSE ranging from 0.126 to 0.186 and R2 from 0.944 to 0.962. Subsets 6 and 7, ignoring leachate properties, had the worst performance in predicting the hydraulic conductivity of Na-B GCLs – that is, in terms of high MSE, MAE and MAAPE and low R2 and RWI – in comparison with those of subsets 1–5.

Table 4.

Model performance for training data using regression analysis (XGBoost algorithm)

SubsetMSE: log2(m/s)R2MAE: log(m/s)MAAPERWI
10.1830.9440.2640.0320.840
20.2440.9260.3400.0410.793
30.1260.9620.2270.0270.862
40.1410.9570.2510.0300.848
50.1860.9440.3100.0380.812
60.3460.8950.4110.0490.751
70.6990.7880.6370.0750.614

Empirical equations were developed using the linear regression method based on subsets, as shown in Table 10. The model with subset 4 (Equation 10) by the linear regression method outperformed the other six Subsets (Table 11), considering the leachate properties, GCL properties, SI and effective stress.

10

where K is the hydraulic conductivity (m/s); σ is the effective stress (kPa); MPUA is the mass per unit area (kg/m2); CM is the concentration of monovalent cations (mM); and CD is the concentration of divalent cations (mM).

The R2 of Equation 10 was 0.531, which was much lower than that developed by the XGBoost method (e.g. 0.531 against 0.957). Equation 10 was supposed to be linear in all eight parameters (e.g. σ, MPUA, SI, CM, CD, I, RMD, I2RMD) in the linear regression method. The logarithm of hydraulic conductivity was reported to be not linear with parameters such as SI and effective stress (e.g. log K ∼ exp(SI), log Kσ0.5), which resulted in poorer model performance (Bradshaw et al., 2013; Katsumi et al., 2007, 2008). Equation 10 could be calculated also to classify the status of the hydraulic conductivity of Na-B GCLs (e.g. ≤10−10 or >10−10 m/s) but with a lower accuracy (0.842). It is important to develop a tool (e.g. classification model) to expedite the screening process for identifying qualified Na-B GCLs to manage specific leachates (e.g. ≤1 × 10−10 m/s).

XGBoost is an ensemble ML method that can evaluate the comparative contribution (or relative importance) of each feature in predicting the target variable. A feature with higher relative importance indicates a greater contribution to predicting the target value. This section analyses the feature importance in XGBoost models across different subsets. The importance of features of subsets 1, 2, 4 and 6 are shown in Figure 5.

Figure 5.

Importance of input parameters on the training data set by regression analysis (XGBoost algorithm): (a) subset 1; (b) subset 2; (c) subset 4; (d) subset 6

Figure 5.

Importance of input parameters on the training data set by regression analysis (XGBoost algorithm): (a) subset 1; (b) subset 2; (c) subset 4; (d) subset 6

Close modal

The ionic strength (I) played a predominant role in subset 1 (Figure 5(a)). Previous studies concluded that the osmotic swelling of Na-B is sensitive to the ionic strength (e.g. >300 mM) in the surrounding environment – that is, osmotic swelling of Na-B can occur in a dilute solution, whereas a solution with high ionic strength can suppress the osmotic swelling of Na-B (Chen et al., 2018; Di Emidio et al., 2015; Jo et al., 2001; Shackelford et al., 2000; Zainab et al., 2021). The hydraulic conductivity is inversely related to the swelling of Na-B – that is, Na-B with a higher swelling in permeant solution can maintain a low hydraulic conductivity (Chen et al., 2018; Guyonnet et al., 2005; Jo et al., 2001, 2005; Katsumi et al., 2008; Kolstad et al., 2004b; Zainab et al., 2021). Divalent cations ranked at second place among all features because the divalent cations in the solution can replace the original sodium in the exchangeable complex of Na-B, resulting in inhibition of the osmotic swelling of Na-B and high hydraulic conductivity (Jo et al., 2001; Kolstad et al., 2004a).

The MPUA and effective stress also affected the hydraulic conductivity of Na-B GCLs, which ranked third and fourth places in subset 1, respectively. Lee et al. (2005) reported that the hydraulic conductivity of Na-B GCLs with an MPUA of 5.1 kg/m2 to 20 mM calcium chloride was 8.8 × 10−11 m/s, whereas Na-B GCLs with an MPUA of 4.0 kg/m2 showed a higher hydraulic conductivity (e.g. 2.0 × 10−8 m/s) under the same testing conditions (Cruz, 2021). Therefore, Na-B GCL with a higher MPUA might have higher chemical compatibility than Na-B GCL with a lower MPUA under the same testing conditions. Increasing the effective confining stress applied on a GCL results in a decrease in hydraulic conductivity (Bradshaw et al., 2016; Li et al., 2021a; Petrov and Rowe, 1997; Petrov et al., 1997). Li et al. (2021a) reported that the hydraulic conductivity of Na-B GCLs permeated with TFGDS-473 leachate (I = 473 mM) decreased from 3.1 × 10−7 to 2.1 × 10−9 m/s as the effective confining stress increased from 20 to 500 kPa.

The concentration of monovalent cation was less important in comparison with the divalent cation concentration (e.g. ranked fifth), because monovalent cations were less aggressive in suppressing the swelling of Na-B than divalent cations at the same concentrations. For example, the SI of Na-B was 23.1 ml/2 g in 100 mM sodium chloride, whereas the SI of Na-B in 100 mM calcium chloride was much lower (e.g. 8.6 ml/2 g) (Jo et al., 2001). The impact of RMD (including I2 × RMD) was at least partially offset by the synergy effect – that was, RMD is calculated using divalent cations and monovalent cations. In subset 2, the ionic strength was also the most important factor followed by effective stress, MPUA and RMD (Figure 5(b)). Without considering monovalent and divalent cations, the importance of ionic strength and RMD increased significantly.

The importance of each feature in subset 4 is shown in Figure 5(c). SI was added as an additional feature. As expected, the SI was the predominant factor that contributed to the model. This result was consistent with the literature and practice that swelling of Na-B could be used as a quick index test to predict the serviceability of Na-B GCLs to leachates (Chen et al., 2018; Jo et al., 2001, 2005; Katsumi et al., 2008; Kolstad et al., 2004a; Lee et al., 2005; Zainab et al., 2021). The ranking of other features was mostly consistent with that in subset 1, except for RMD and I2 × RMD. In subset 6, the SI was also ranked at first place, followed by MPUA and effective stress (Figure 5(d)).

The XGBoost model also consistently had the best performance for testing data set (Table 11). The model performance for testing data using the XGBoost method is shown in Table 5. The performance of subsets 3–5 was better than that of subsets 1 and 2 (e.g. lower MSE, MAE and MAAPE and higher RWI and R2). The results illustrated the significant contribution of SI in predicting the hydraulic conductivity of Na-B GCLs. Subsets 6 and 7 showed relatively poor performance due to ignoring the leachate properties as features (e.g. higher MSE, MAE and MAAPE and lower RWI and R2). Similar to the observed patterns in training data, subsets 3 and 4 showed great results. Subset 4 had the best performance in terms of the smallest MSE, MAE and MAAPE and highest RWI and R2, followed by subset 3. In comparison with subset 3, subset 4 consists of one additional feature - I2 × RMD, which reflects that the sensitivity to RMD varies non-linearly with ionic strength (Kolstad et al., 2004a). Subset 6 had better performance than subset 7 (e.g. MSE: 0.776 against 1.113; R2: 0.747 against 0.637), illustrating the importance of MPUA in predicting the hydraulic conductivity of Na-B GCLs. The testing results of the 70:30 ratio exhibit comparable performance (slightly lower) with those of the 80:20 ratio.

Table 5.

Model performance for testing data using regression analysis (XGBoost algorithm)

SubsetMSE: log2(m/s)R2MAE: log(m/s)MAAPERWI
10.5780.8110.5600.0690.635
20.5800.8110.5830.0710.620
30.5550.8190.5520.0660.640
40.5350.8260.5300.0640.654
50.5720.8130.5730.0680.626
60.7760.7470.6490.0770.577
71.1130.6370.8290.0950.460

A comparison of the predicted hydraulic conductivity and the measured hydraulic conductivity of Na-B GCLs (subsets 4 and 6) is plotted in Figure 6. The results showed that subset 4 led to more accurate prediction than subset 6 with a higher R2 (e.g. 0.826 against 0.747). The majority of predicted hydraulic conductivities using subset 4 were within ten times the measured values. In contrast, more data points fell between ten and 100 times using subset 6.

Figure 6.

Comparison between the predicted and measured hydraulic conductivities on the testing data set by regression analysis (XGBoost algorithm): (a) subset 4; (b) subset 6

Figure 6.

Comparison between the predicted and measured hydraulic conductivities on the testing data set by regression analysis (XGBoost algorithm): (a) subset 4; (b) subset 6

Close modal

In this section, the results of five classification models (logistic regression, XGBoost, SVM, KNN and ANN) are discussed. The accuracy of all five models across all seven subsets are shown in Figure 7. The classification model based on logistic regression showed the worst model performance compared with those with advanced algorithms (e.g. XGBoost, SVM, KNN and ANN). In each subset, the XGBoost method obtained higher accuracy than the other four methods and E (Figure 7(a)). For example, in subset 4, the accuracy of XGBoost was 0.976, which was 14.6, 8.4, 10.5 and 8.4% higher than the accuracy of the logistic regression, SVM, KNN and ANN methods, respectively. The AUC value also supported that the XGBoost method showed the best prediction performance in all seven Subsets, as shown in Figure 7(b). For example, the AUC of subset 4 by the XGBoost method was 0.994, which was 2.8–9.5% higher than those of the other four methods. Herein, the further analysis of classification results thus would be focused on the XGBoost method.

Figure 7.

Model performances on training data set by classification analysis using logistic regression, XGBoost, SVM, KNN and ANN: (a) accuracy; (b) AUC

Figure 7.

Model performances on training data set by classification analysis using logistic regression, XGBoost, SVM, KNN and ANN: (a) accuracy; (b) AUC

Close modal

The performance of training models (e.g. seven subsets) using the XGBoost method are shown in Table 6. The accuracy of all seven subsets ranged from 0.866 to 0.976, and the AUCs were between 0.936 and 0.994. Subsets 1 and 2 (leachate properties, effective stress, MPUA) displayed good performance with the accuracy ranging from 0.939 to 0.963 and AUCs ranging from 0.981 to 0.994. By adding an additional feature – the SI – the accuracy of subsets 3–5 increased from 0.939 to 0.976. Subset 6, ignoring leachate properties, yielded lower accuracy (0.919) and AUC (0.962) in comparison with subsets 1–5. Subset 7, further excluding MPUA from subset 6, showed the worst performance among all models, with the lowest accuracy (0.866) and AUC (0.936). The classification model developed in this study did not directly provide hydraulic conductivity prediction, which could be obtained from regression model.

Table 6.

Model performance for training data by classification analysis (XGBoost algorithm)

SubsetAccuracyAUC
10.9630.994
20.9390.981
30.9670.993
40.9760.994
50.9390.989
60.9190.962
70.8660.936

Feature importance in XGBoost models across different subsets was analysed in this section. The importance of features for subsets 1, 2, 4 and 6 is plotted in Figure 8. The ionic strength played a predominant role with importance factors of 0.423, 0.589 and 0.344 in subsets 1, 2 and 4, respectively. This observation was consistent with the previous results that ionic strength significantly affected the hydraulic conductivity of Na-B GCLs (Chen et al., 2018; Jo et al., 2001; Li et al., 2021b; Shackelford et al., 2000; Zainab et al., 2021). Divalent cations ranked second place among all features of leachate properties In subsets 1 and 4, illustrating the adverse effect of divalent cations in solutions on the hydraulic conductivity of Na-B GCLs. In subset 4, SI was added as an additional feature and ranked third place (e.g. 0.176), showing its significant contribution in classifying the hydraulic conductivity. Effective stress and MPUA were ranked after ionic strength and divalent cations (and SI in subset 4). Monovalent cations, RMD and I2 × RMD, however, were the three least important to the classification of Na-B GCLs in subsets 1 and 4.

Figure 8.

Importance of input parameters on the training data set by classification analysis (XGBoost algorithm): (a) subset 1; (b) subset 2; (c) subset 4; (d) subset 6

Figure 8.

Importance of input parameters on the training data set by classification analysis (XGBoost algorithm): (a) subset 1; (b) subset 2; (c) subset 4; (d) subset 6

Close modal

In subset 2, the ionic strength ranked at first place, followed by RMD. The RMD in subset 2 had a higher contribution in classifying the hydraulic conductivity than the RMD (I2 × RMD) in subsets 1 and 4. The reason that the contribution of RMD (and I2 × RMD) was weakened in subsets 1 and 4 was because of the synergy effect among monovalent cations, divalent cations and RMD – that is, RMD was derived from monovalent and divalent cations (Equation 2).

In subset 6, SI showed a dominating importance factor of 0.633 (Figure 8(d)), which is smaller than that of the corresponding case in regression analysis, 0.748 (Figure 5(d)). In subset 4 with leachate properties, SI became less important, which was ranked at third place after ionic strength and divalent cations, with an importance factor of 0.176 (Figure 8(c)). This observation was different from the regression results of subset 4. In the regression case, SI outperformed other features in subset 4 with an importance factor of 0.434 (Figure 5(c)). These observations indicated that SI played a more important role in predicting than classifying the hydraulic conductivity of Na-B GCLs.

The performance of XGBoost models across all subsets in testing data is summarised in Table 7. The accuracy of all seven subsets ranged from 0.742 to 0.887, with AUCs from 0.825 to 0.931. The model developed from subset 4 was considered the best model since it showed dominating performance in both evaluation metrics. Subsets 1 and 3 displayed great performance, ranked after subset 6. Subsets 1 and 3 had the same accuracy value (0.871), while subset 3 showed a slightly higher AUC value (0.927 against 0.926) than subset 1. Subsets 6 and 7, without considering leachate properties, displayed relatively poor performance, whose accuracy values were below 0.80 (i.e. 0.790 and 0.742, respectively). These results showed that subset 4 provided the best classification of the hydraulic conductivities of Na-B GCLs permeated with saline solution and leachates, which can be an effective method to screen Na-B GCLs in practice.

Table 7.

Model performance for testing data by classification analysis (XGBoost algorithm)

SubsetAccuracyAUC
10.8710.926
20.8550.895
30.8710.927
40.8870.931
50.8550.890
60.7900.872
70.7420.825

The ML models were used to predict and classify the hydraulic conductivity of Na-B GCLs to leachates. The data set (including 308 observations) was compiled from 22 papers published from 2000 to 2022. Features such as leachate properties (I, RMD, monovalent and divalent cations), MPUA, SI and effective stress were used as input features for the ML models. ML models, including linear regression, logistic regression, XGBoost, SVM, KNN and ANN, were applied and compared in this study. The following conclusions and recommendations were drawn.

  • The XGBoost method outperformed all the other models (linear, Logistic, SVM, KNN and ANN) over all of feature subsets for both regression and classification analyses. The XGBoost method showed a lower MSE value and a higher R2 than other algorithms in regression analysis and higher accuracy and AUC value in classification analysis.

  • Subset 4, including MPUA, effective stress, monovalent and divalent cations, I, RMD, I2 × RMD and SI as features, showed the best performance in testing data, with the lowest MSE and highest R2 in regression analysis and the highest accuracy and AUC value in classification analysis.

  • The XGBoost model using subset 4 as features provided good fitting for both regression and classification tasks. The classification model could be used as a tool to screen qualified Na-B GCLs to specific leachates (e.g. ≤1 × 10−10 m/s), while the regression model could provide good estimation of the hydraulic conductivity of Na-B GCLs.

It is recommended to apply the XGBoost model using subset 4 (e.g. MPUA, σ, monovalent cations, divalent cations, I, RMD, I2 × RMD, SI) for both regression and classification. The regression model can be used to predict the hydraulic conductivity of Na-B GCLs, and the classification model can classify the status of hydraulic conductivity based on the threshold value (e.g. ‘success’ if ≤1 × 10−10 m/s; ‘failure’, otherwise). If leachate properties are unknown, other subsets (e.g. subsets 6 and 7) also can be used by the XGBoost model to obtain a rough estimation of the hydraulic conductivity of Na-B GCLs. Additional laboratory hydraulic conductivity tests should be conducted to validate the results. To evaluate the long-term hydraulic performance of Na-B GCLs, laboratory hydraulic conductivity tests are highly recommended. The cation-exchange process also affected the long-term hydraulic conductivity of Na-B GCLs (Jo et al., 2005). The different rankings of the importance of input parameters were observed in regression and classification analyses, which could not be explained only by statistical analysis in this study. Laboratory experiments should be conducted to investigate the different importance of input parameters observed in regression and classification analyses. In the future, more data are desired to enrich the database to modify and validate the models. A hybrid ML model combining metaheuristic algorithms (e.g. GWO) can be developed and validated to improve further the prediction accuracy and generalisation ability.

Tables 8 and 9 show the hyperparameters tuned for regression and classification models, respectively. In each table, the first and second columns list the model and parameter names, respectively, while the third column illustrates the candidate values for grid search. The last column shows the optimal value for the selected model in subset 6 (the subset with best performance).

Table 8.

Parameter settings for regression models

ModelParameter nameCandidate valuesOptimal value
XGBoostnrounds15, 20, 25, 30, 3530
Eta0.3, 0.4, 0.50.4
Gamma0.2, 0.3, 0.40.3
max_depth10, 12, 15, 18, 2012
min_child_weight1, 2, 3, 4, 53
SVMDegree1, 2, 3, 52
C2, 3, 5, 8, 103
Scale0.1, 0.3, 0.50.3
KNNK2, 3, 4, 5, 6, 7, 8, 9, 103
ANNSize1, 2, 3, 4, 5, 6, 7, 86
Decay0.0001, 0.001, 0.01, 0.05, 0.1, 0.50.1
Table 9.

Parameter settings for classification models

ModelParameter nameCandidate valuesOptimal value
XGBoostnrounds3, 5, 7, 10, 15, 2015
Eta0.5, 0.7, 0.9, 11
Gamma0, 0.3, 0.6, 10.3
max_depth1, 5, 10, 15, 205
min_child_weight1, 3, 5, 8, 123
SVMSigma0.1, 0.3, 0.50.3
C3, 5, 8, 108
KNNK2, 3, 4, 53
ANNSize5, 10, 15, 20, 25, 3025
Decay0.001, 0.005, 0.01, 0.05, 0.10.05

The models are obtained using the R 4.1.2 software with packages caret 6.0, xgboost 1.5.0.2, e1071 1.7 and nnet 7.3.

Table 10 shows the equations for predicting the hydraulic conductivity of Na-B GCLs using the linear regression method based on subsets. Tables 11 and 12 show model performance on training data set and testing data set for regression analysis. The data are provided in the Microsoft Excel file ‘Data-ENGE-2022-181.xlsx’ in the online supplementary material.

Table 10.

Equations for predicting the hydraulic conductivity of Na-B GCLs using the linear regression method based on subsets

SubsetEmpirical equation (linear regression method)
1logK=11.790.003σ0.19MPUA0.559logCM0.086logCD+1.86logI0.001RMD+0.18log(I2RMD)
2logK=12.780.003σ+0.14MPUA+1.67logI0.001RMD
3logK=8.640.003σ0.31MPUA0.08SI0.23logCM+0.127logCD+1.13logI+0.01RMD
4logK=8.560.002σ0.31MPUA0.08SI0.18logCM+0.15logCD+1.11logI+0.01RMD0.03log(I2RMD)
5logK=9.600.003σ0.029MPUA0.089SI+1.02logI0.0004RMD
6logK=5.940.002σ0.05MPUA0.16SI
7logK=6.180.002σ0.16SI
Table 11.

Model performance of training data using regression analysis

SubsetMSE: log2(m/s)R2MAE: log(m/s)MAAPERWI
LinearXGBSVMKNNANNLinearXGBSVMKNNANNLinearXGBSVMKNNANNLinearXGBSVMKNNANNLinearXGBSVMKNNANN
11.6910.1831.2480.9680.9560.4880.9440.6220.7070.7101.0960.2640.8580.7320.7800.1290.0320.1050.0860.0920.3350.8400.4790.5560.527
21.8670.2441.5870.6970.9370.4350.9260.5190.7890.7161.1680.3400.9990.6180.7840.1380.0410.1200.0730.0940.2910.7930.3940.6250.524
31.5470.1261.0860.6790.7730.5310.9620.6710.7940.7661.0530.2270.7960.5960.6980.1230.0270.0950.0700.0830.3610.8630.5170.6380.577
41.5470.1411.0870.7290.7940.5310.9570.6710.7790.7601.0530.2510.7870.6150.7140.1230.0300.0940.0730.0840.3610.8480.5230.6270.567
51.6870.1861.3810.8501.1270.4890.9440.5820.7430.6591.1220.3100.9490.7140.9030.1320.0380.1130.0850.1070.3190.8120.4240.5670.452
61.9200.3461.6480.7730.8230.4190.8950.5010.7660.7511.2010.4111.0370.6400.6940.1420.0490.1210.0750.0820.2710.7510.3700.6120.579
71.9210.6991.7691.3021.1450.4180.7880.4640.6060.6531.2040.6371.0850.9270.8810.1420.0750.1280.1100.1050.2690.6140.3410.4370.465
Table 12.

Model performance of testing data using regression analysis

SubsetMSE: log2(m/s)R2MAE: log(m/s)MAAPERWI
LinearXGBSVMKNNANNLinearXGBSVMKNNANNLinearXGBSVMKNNANNLinearXGBSVMKNNANNLinearXGBSVMKNNANN
11.5210.5781.6821.1351.0680.5030.8110.4510.6290.6521.0320.5601.0430.8320.8220.1210.0690.1290.0950.0980.3270.6350.3200.4580.464
21.5810.5801.8941.0931.9100.4840.8110.3820.6430.3771.0380.5831.1120.7971.0480.1210.0710.1340.0940.1260.3240.6200.2750.4810.317
31.4310.5551.4301.0490.8870.5330.8190.5330.6580.7111.0160.5520.9590.7580.7670.1180.0660.1170.0890.0950.3380.6400.3750.5060.500
41.4300.5351.4611.0810.9630.5330.8260.5230.6470.6861.0160.5300.9640.7790.7710.1180.0640.1170.0900.0920.3380.6540.3720.4920.498
51.4670.5721.5551.1771.5470.5210.8130.4930.6160.4951.0310.5730.9940.8211.0570.1190.0680.1200.0950.1280.3280.6260.3520.4650.311
61.6740.7761.5761.1691.7310.4540.7470.4860.6180.4351.1310.6491.0180.8281.0200.1320.0770.1180.0990.1230.2630.5770.3360.4600.335
71.6811.1131.7461.8831.4850.4510.6370.4300.3860.5151.1360.8291.0531.1151.0250.1320.0950.1230.1310.1230.2590.4600.3140.2730.332
Anava
O
,
Levy
K
2016
k*-nearest neighbors: from global to local
Advances in Neural Information Processing Systems 29
Lee
D
,
Sugiyama
M
,
Luxburg
U
,
Guyon
I
,
Garnett
R
Neural Information Processing Systems
La Jolla, CA, USA
4923
-
4931
Ashmawy
AK
,
El-Hajji
D
,
Sotelo
N
,
Muhammad
N
2002
Hydraulic performance of untreated and polymer-treated bentonite in inorganic landfill leachates
Clays and Clay Minerals
50
5
546
-
552
ASTM
2019
D 5890: Standard test method for swell index of clay mineral component of geosynthetic clay liners
ASTM International
West Conshohocken, PA, USA
ASTM
2020
D 6766: Standard test method for evaluation of hydraulic properties of geosynthetic clay liners permeated with potentially incompatible aqueous solutions
ASTM International
West Conshohocken, PA, USA
Athanassopoulos
C
,
Benson
CH
,
Donovan
M
,
Chen
JN
2015
Hydraulic conductivity of a polymer-modified GCL permeated with high-pH solutions
Proceedings of the Geosynthetics Conference 2015
Portland, OR, USA
Benson
CH
,
Meer
SR
2009
Relative abundance of monovalent and divalent cations and the impact of desiccation on geosynthetic clay liners
Journal of Geotechnical and Geoenvironmental Engineering
135
3
349
-
358
Bradshaw
SL
,
Benson
CH
2014
Effect of municipal solid waste leachate on hydraulic conductivity and exchange complex of geosynthetic clay liners
Journal of Geotechnical and Geoenvironmental Engineering
140
4
04013038
Bradshaw
SL
,
Benson
CH
,
Scalia
J
2013
Hydration and cation exchange during subgrade hydration and effect on hydraulic conductivity of geosynthetic clay liners
Geotextile and Geomembranes
139
4
526
-
538
Bradshaw
SL
,
Benson
CH
,
Rauen
TL
2016
Hydraulic conductivity of geosynthetic clay liners to recirculated municipal solid waste leachates
Journal of Geotechnical and Geoenvironmental Engineering
142
2
04015074
Chao
Z
,
Dang
Y
,
Pan
Y
, et al
2023a
Prediction of the shale gas permeability: a data mining approach
Geomechanics for Energy and the Environment
33
article 100435
Chao
Z
,
Shi
D
,
Fowmes
G
, et al
2023b
Artificial intelligence algorithms for predicting peak shear strength of clayey soil–geomembrane interfaces and experimental validation
Geotextiles and Geomembranes
51
1
179
-
198
Chen
TQ
,
Guestrin
C
2016
XGBoost: a scalable tree boosting system
KDD ’16: Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining
Association for Computing Machinery
New York, NY, USA
785
-
794
Chen
JN
,
Benson
CH
,
Edil
TB
2018
Hydraulic conductivity of geosynthetic clay liners with sodium bentonite to coal combustion product leachates
Journal of Geotechnical and Geoenvironmental Engineering
144
3
04018008
Chen
JN
,
Salihoglu
H
,
Benson
CH
,
Likos
WJ
,
Edil
TB
2019a
Hydraulic conductivity of bentonite–polymer composite geosynthetic clay liners permeated with coal combustion product leachates
Journal of Geotechnical and Geoenvironmental Engineering
145
9
04019038
Chen
M
,
Challita
U
,
Saad
W
,
Yin
C
,
Debbah
M
2019b
Artificial neural networks-based machine learning for wireless networks: a tutorial
IEEE Communications Surveys & Tutorials
21
4
3039
-
3071
Cortes
C
,
Vapnik
V
1995
Support-vector networks
Machine Learning
20
3
273
-
297
Cruz
AJ
2021
Effect of Elevated Temperature on Hydraulic Conductivity of Bentonite–Polymer Composite Geosynthetic Clay Liner to Saline Solutions. MS thesis
George Mason University
Fairfaz, VA, USA
Das
SK
,
Samui
P
,
Sabat
AK
2012
Prediction of field hydraulic conductivity of clay liners using an artificial neural network and support vector machine
International Journal of Geomechanics
12
5
606
-
611
Di Emidio
G
,
Mazzieri
F
,
Verastegui-Flores
RD
,
Van Impe
W
,
Bezuijen
A
2015
Polymer-treated bentonite clay for chemical-resistant geosynthetic clay liners
Geosynthetics International
22
1
125
-
137
Elbisy
MS
2015
Support vector machine and regression analysis to predict the field hydraulic conductivity of sandy soil
KSCE Journal of Civil Engineering
19
7
2307
-
2316
EPRI (Electric Power Research Institute)
2014
Engineering Properties of Geosynthetic Clay Liners Permeated with Coal Combustion Product Leachate
EPRI
Palo Alto, CA, USA
Report No. 3002003770
Gastelo
J
,
Li
D
,
Tian
K
,
Tanyu
BF
,
Guler
FE
2023
Hydraulic conductivity of GCL overlap permeated with saline solutions
Waste Management
157
348
-
356
Goli
VSNS
,
Paleologos
EK
,
Farid
A
, et al
2022
Extraction, characterisation and remediation of microplastics from organic solid matrices
Environmental Geotechnics
Guyonnet
D
,
Gaucher
E
,
Gaboriau
H
, et al
2005
Geosynthetic clay liner interaction with leachate: correlation between permeability, microstructure, and surface chemistry
Journal of Geotechnical and Geoenvironmental Engineering
131
6
740
-
749
Harrell
FE
 Jr
,
Lee
KL
,
Mark
DB
1996
Multivariable prognostic models: issues in developing models, evaluating assumptions and adequacy, and measuring and reducing errors
Statistics in Medicine
15
4
361
-
387
James
G
,
Witten
D
,
Hastie
T
,
Tibshirani
R
2013
An Introduction to Statistical Learning: With Applications in R
Springer
New York, NY, USA
Jiang
NJ
,
Hanson
JL
,
Della Vecchia
G
, et al
2021
Geotechnical and geoenvironmental engineering education during the pandemic
Environmental Geotechnics
8
3
233
-
243
Jo
HY
,
Katsumi
T
,
Benson
CH
,
Edil
TB
2001
Hydraulic conductivity and swelling of nonprehydrated GCLs permeated with single-species salt solutions
Journal of Geotechnical and Geoenvironmental Engineering
127
7
557
-
567
Jo
HY
,
Benson
CH
,
Shackelford
CD
,
Lee
JM
,
Edil
TB
2005
Long-term hydraulic conductivity of a geosynthetic clay liner permeated with inorganic salt solutions
Journal of Geotechnical and Geoenvironmental Engineering
131
4
405
-
417
Joseph
VR
2022
Optimal ratio for data splitting
Statistical Analysis and Data Mining the ASA Data Science Journal
15
4
531
-
538
Kardani
N
,
Aminpour
M
,
Raja
MNA
, et al
2022
Prediction of the resilient modulus of compacted subgrade soils using ensemble machine learning methods
Transportation Geotechnics
36
article 100827
Katsumi
T
,
Ishimori
H
,
Ogawa
A
, et al
2007
Hydraulic conductivity of nonprehydrated geosynthetic clay liners permeated with inorganic solutions and waste leachates
Soils and Foundations
47
1
79
-
96
Katsumi
T
,
Ishimori
H
,
Onikata
M
,
Fukagawa
R
2008
Long-term barrier performance of modified bentonite materials against sodium and calcium permeant solutions
Geotextiles and Geomembranes
26
1
14
-
30
Kolstad
DC
,
Benson
CH
,
Edil
TB
2004a
Hydraulic conductivity and swell of nonprehydrated geosynthetic clay liners permeated with multispecies inorganic solutions
Geotextiles and Geomembranes
130
12
1236
-
1249
Kolstad
DC
,
Benson
CH
,
Edil
TB
,
Jo
HY
2004b
Hydraulic conductivity of a dense prehydrated GCL permeated with aggressive inorganic solutions
Geosynthetics International
11
3
233
-
241
Lee
JM
,
Shackelford
CD
,
Benson
CH
,
Jo
HY
,
Edil
TB
2005
Correlating index properties and hydraulic conductivity of geosynthetic clay liners
Journal of Geotechnical and Geoenvironmental Engineering
131
11
1319
-
1329
Li
D
,
Tian
K
2022
Effects of prehydration on hydraulic conductivity of bentonite–polymer geosynthetic clay liner to coal combustion product leachates
Geo-Congress 2022: Soil Improvement, Geosynthetics, and Innovative Geomaterials
Lemnitzer
A
,
Stuedlein
AW
American Society of Civil Engineers
Reston, VA, USA
568
-
577
Li
D
,
Zainab
B
,
Tian
K
2021a
Effect of effective stress on hydraulic conductivity of bentonite–polymer geosynthetic clay liners to coal combustion product leachates
Environmental Geotechnics
40
Li
Q
,
Chen
JN
,
Benson
CH
,
Peng
D
2021b
Hydraulic conductivity of bentonite–polymer composite geosynthetic clay liners permeated with bauxite liquor
Geotextiles and Geomembranes
49
2
420
-
429
Li
D
,
Tian
K
,
Gorakhki
R
,
Donovan
M
2023
Hydraulic conductivity of lighter bentonite–polymer geosynthetic clay liners
Proceedings of Geosynthetics Conference 2023
Kansas, MO, USA
231
-
238
Lin
L
,
Katsumi
T
,
Kamon
M
, et al
2000
Evaluation of chemical-resistant bentonite for landfill barrier application
Disaster Prevention Research Institute Annuals
43
B-2
525
-
533
Meer
SR
,
Benson
CH
2007
Hydraulic conductivity of geosynthetic clay liners exhumed from landfill final covers
Journal of Geotechnical and Geoenvironmental Engineering
133
5
550
-
563
Ozcoban
MS
,
Isenkul
ME
,
Güneş-Durak
S
, et al
2018
Predicting permeability of compacted clay filtrated with landfill leachate by k-nearest neighbors modelling method
Water Science and Technology
77
8
2155
-
2164
Peduzzi
P
,
Concato
J
,
Kemper
E
,
Holford
TR
,
Feinstein
AR
1996
A simulation study of the number of events per variable in logistic regression analysis
Journal of Clinical Epidemiology
49
12
1373
-
1379
Petrov
RJ
,
Rowe
RK
1997
Geosynthetic clay liner (GCL)–chemical compatibility by hydraulic conductivity testing and factors impacting its performance
Canadian Geotechnical Journal
34
6
863
-
885
Petrov
RJ
,
Rowe
RK
,
Quigley
RM
1997
Comparison of laboratory measured GCL hydraulic conductivity based on three permeameter types
Geotechnical Testing Journal
20
1
49
-
62
Raja
MNA
,
Shukla
SK
2021
Predicting the settlement of geosynthetic-reinforced soil foundations using evolutionary artificial intelligence technique
Geotextiles and Geomembranes
49
5
1280
-
1293
Rogiers
B
,
Mallants
D
,
Batelaan
O
, et al
2012
Estimation of hydraulic conductivity and its uncertainty from grain-size data using GLUE and artificial neural networks
Mathematical Geosciences
44
6
739
-
763
Salami
BA
,
Iqbal
M
,
Abdulraheem
A
, et al
2022
Estimating compressive strength of lightweight foamed concrete using neural, genetic and ensemble machine learning approaches
Cement and Concrete Composites
133
article 104721
Salihoglu
H
2015
Behavior of Polymer-modified Bentonites with Aggressive Leachates. MS thesis
University of Wisconsin–Madison
Madison, WI, USA
Samui
P
2008
Support vector machine applied to settlement of shallow foundations on cohesionless soils
Computers and Geotechnics
35
3
419
-
427
Sanders
R
1987
The Pareto principle: its use and abuse
Journal of Services Marketing
1
2
37
-
40
Scalia
J
,
Benson
CH
,
Bohnhoff
GL
,
Edil
TB
,
Shackelford
CD
2014
Long-term hydraulic conductivity of a bentonite–polymer composite permeated with aggressive inorganic solutions
Journal of Geotechnical and Geoenvironmental Engineering
140
3
04013025
Scalia
J
,
Bohnhoff
GL
,
Shackelford
CD
, et al
2018
Enhanced bentonites for containment of inorganic waste leachates by GCLs
Geosynthetics International
25
4
392
-
411
Setz
MC
,
Tian
K
,
Benson
CH
,
Bradshaw
SL
2017
Effect of ammonium on the hydraulic conductivity of geosynthetic clay liners
Geotextiles and Geomembranes
45
6
665
-
673
Shackelford
CD
,
Benson
CH
,
Katsumi
T
,
Edil
TB
,
Lin
L
2000
Evaluating the hydraulic conductivity of GCLs permeated with non-standard liquids
Geotextiles and Geomembranes
18
2–4
133
-
161
Shahin
MA
,
Maier
HR
,
Jaksa
MB
2002
Predicting settlement of shallow foundations using neural networks
Journal of Geotechnical and Geoenvironmental Engineering
128
9
785
-
793
Shan
HY
,
Lai
YJ
2002
Effect of hydrating liquid on the hydraulic properties of geosynthetic clay liners
Geotextiles and Geomembranes
20
1
19
-
38
Tian
K
,
Li
D
2021
Hydraulic conductivity of bentonite–polymer geosynthetic clay liners to coal combustion product leachates
Geosynthetics Conference 2021
Beauregard
M
,
Nicks
JE
Industrial Fabrics Association International
Roseville, MN, USA
173
-
180
Tian
K
,
Benson
CH
,
Likos
WJ
2016
Hydraulic conductivity of geosynthetic clay liners to low-level radioactive waste leachate
Journal of Geotechnical and Geoenvironmental Engineering
142
8
04016037
Tian
K
,
Likos
WJ
,
Benson
CH
2019
Polymer elution and hydraulic conductivity of bentonite–polymer composite geosynthetic clay liners
Journal of Geotechnical and Geoenvironmental Engineering
145
10
04019071
Wang
B
,
Xu
J
,
Chen
B
,
Dong
X
,
Dou
T
2019
Hydraulic conductivity of geosynthetic clay liners to inorganic waste leachate
Applied Clay Science
168
3
244
-
248
Wireko
C
,
Abichou
T
2021
Investigating factors influencing polymer elution and the mechanism controlling the chemical compatibility of GCLs containing linear polymers
Geotextiles and Geomembranes
49
4
1004
-
1018
Wireko
C
,
Zainab
B
,
Tian
K
,
Abichou
T
2020
Effect of specimen preparation on the swell index of bentonite–polymer GCLs
Geotextiles and Geomembranes
48
6
875
-
885
Wireko
C
,
Abichou
T
,
Tian
K
,
Zainab
B
,
Zhang
Z
2022
Effect of incineration ash leachates on the hydraulic conductivity of bentonite–polymer composite geosynthetic clay liners
Waste Management
139
25
-
38
Zainab
B
,
Wireko
C
,
Dong
L
,
Tian
K
,
Tarek
K
2021
Hydraulic conductivity of bentonite–polymer geosynthetic clay liners to coal combustion product leachates
Geotextiles and Geomembranes
49
5
1129
-
1138
Zhang
W
,
Wu
C
,
Zhong
H
,
Li
Y
,
Wang
L
2021
Prediction of undrained shear strength using extreme gradient boosting and random forest based on Bayesian optimization
Geoscience Frontiers
12
1
469
-
477
Zhao
HR
,
Li
D
,
Tian
K
2023
Long-term hydraulic conductivity of bentonite–polymer geosynthetic clay liner to coal combustion product leachates
Proceedings of Geosynthetics Conference 2023
Kansas, MO, USA
239
-
247

Supplementary data

or Create an Account

Close Modal
Close Modal