Purpose

Cryptocurrency markets are gaining popularity, with over 23,000 cryptocurrencies in 2023 and a total market valuation of 870.81 billion USD in 2023. With its increasing popularity, cryptocurrencies are also susceptible to volatility. Predicting the price with the least fallacy or more accuracy has become the need of the hour as it significantly influences investment decisions.

Design/methodology/approach

This study aims to create a dynamic forecasting model using the ensemble method and test the forecasting accuracy of top 15 cryptocurrencies’ prices. Statistical and econometric model prediction accuracy is examined after hyper tuning the parameters. Drawing inferences from the statistical model, an ensemble model using machine learning (ML) algorithms is developed using gradient-boosted regressor (GBR), random forest regressor (RFR), support vector regression (SVR) and multi-layer perceptron (MLP). Validation curves are utilized to optimize model parameters and boost prediction accuracy.

Findings

It is found that when the price movement exhibits autocorrelation, the autoregressive integrated moving average (ARIMA) model and the ensemble model performed better. ARIMA, simple linear regression (SLR), random forest (RF), decision tree (DT), gradient boosting (GB) and multi-model regression (MLR) ensemble models performed well with coins, showing that trends, seasonality and historical price patterns are prominent. Furthermore, the MLR approach produces more accurate predictions for coins with higher volatility and irregular price patterns.

Research limitations/implications

Although the dataset includes crisis period data, anomalies or outliers are yet to be explicitly excluded from the analysis. The models employed in this study still demonstrate high accuracy in predicting cryptocurrency prices despite these outliers, suggesting that the models are robust enough to handle unexpected fluctuations or extreme events in the market. However, the lack of specific analysis on the impact of outliers on model performance is a limitation of the study, as it needs to fully explore the resilience of the forecasting models under adverse market conditions.

Practical implications

The present study contributes to the body of literature on ensemble methods in forecasting crypto price in general, potentially influencing future studies on price forecasting. The study motivates the researchers on empirical testing of our framework on various asset classes. As a result, on the prediction ability of ensemble model, the study will significantly influence the decision-making process of traders and investors. The research benefits the traders and investors to effectively develop a model to forecast cryptocurrency price. The findings highlight the potential of ensemble model in predicting high volatile cryptocurrencies and other financial assets. Investors can design the investment strategies and asset allocation decisions by understanding the relationship between market trends and consumer behavior. Investors can enhance portfolio performance and mitigate risk by incorporating these insights into their decision-making processes. Policymakers can use this information to design more effective regulations and policies promoting economic stability and consumer welfare. The study emphasizes the need for using diversified model to understand the market dynamics and improving trading strategies.

Originality/value

This research, to the best of our knowledge, is the first to use the above models to develop an ensemble model on the data for which the outliers have not been adjusted, and the model still outperformed the other statistical, econometric, ML and deep learning (DL) models.

研究目的

加密貨幣市場越來越受歡迎; 於2023年,不同種類的加密貨幣為數已超過23,000種; 同年,它們的總市場估值為八千七百零八點壹億美元。雖然加密貨幣越來越受歡迎,但它們仍然容易受到波動性的影響。預測謬誤減至最少的價格或作出更準確的價格預測就成為某些特定時刻的首要事項,這是因為投資決策會顯著地受到這些預測的影響。

研究方法

研究人員擬以集成學習方法來創造一個動態預測模型,並以此模型測試預測15個頂尖加密貨幣價格的準確性。 研究人員調校超參數後,便審查統計及計量經濟學模式的預測準確性。研究人員基於從統計模式作出的推斷,研製一個使用機器學習算法的集成模型。研究人員在研製這個集成模型時,使用了梯度提升迴歸變量、隨機森林迴歸、支持向量迴歸和多層感知器。 驗證曲線被用來優化模型參數,以及提高預測的準確性。

研究結果

研究人員發現,當價格變動展示自相關時,差分整合移動平均自我迴歸模型和集成模型會表現得更好; 另外,若使用加密貨幣,差分整合移動平均自我迴歸模型、簡單線性迴歸、隨機森林、決策樹、梯度提升和多模型迴歸集成模型會有良好的表現。再者,就波動性較高和價格模式不規則的加密貨幣而言,採用多重線性迴歸的方法會使預測更為準確。

研究的原創性

據我們所知,這是首個研究,以上述的各個模型來研發一個集成模型,而這個集成模型,雖建基於異常值並未調整的數據,但它的表現卻比其它的統計、計量經濟學、多重線性和深度學習等的模型更為優良。

A cryptocurrency is a cryptographically secured digital currency distributed on a decentralized network using blockchain technology, making them anonymous and untraceable. The advent of cryptocurrencies and blockchain technology has sparked a revolutionary shift in the financial sector (Kayani, 2023; Kayani and Hasan, 2024).

Cryptocurrencies offer substantial returns on investments besides possessing a higher risk factor. A classic example would be the Bitcoin index (in EUR), which stands at a 108.27% compound annual growth rate and a standard deviation of 156.99% between 2011 and 2024. Figure 1 shows the abnormal spike in Bitcoin return during 2021, with significantly higher fluctuations indicating heightened volatility. This calls for accurate forecasting models to combat the large fluctuations in non-stationary cryptocurrencies (Bouteska et al., 2024). Prediction of cryptocurrency prices is also important for several other reasons, such as building trading strategies, risk management, price discovery, market sentiment analysis and business applications.

Figure 1

Historical returns of cryptocurrencies

Figure 1

Historical returns of cryptocurrencies

Close modal

Fundamentalists and chartists rely on price prediction to make an informed decision. Chartists looking for price patterns and trends can refine their technical analysis based on forecasting. While fundamentalist’s primary focus remains on long-term value assessment, the model’s predictions could potentially serve as an additional data point for considering short-term market sentiment that might influence their investment decisions (Soltani et al., 2023).

The early-stage researchers employed statistical and econometrics models to forecast the price of cryptocurrencies. To predict Bitcoin’s short-term prices, the autoregressive integrated moving average (ARIMA) model was applied by Wirawan et al. (2019). The progressive growth in machine learning (ML) also motivated researchers to compare the forecasting accuracy of various models. Lyu (2022) compared the accuracy of various ML models and found gradient boosting (GB) to be the most efficient in predicting most major cryptocurrencies.

Recent studies on forecasting started applying deep learning (DL) and hybrid models, combining classical models such as the ARIMA model and artificial neural networks (Suhartono et al., 2017). Li et al. (2022) found that the hybrid model improved the forecasting accuracy of foreign currency and demonstrated that the hybrid data decomposition model outperforms econometric, ML and deep-learning models. Chen (2023) predicted the Bitcoin price of the next day and found that random forest (RF) gives better accuracy than long short-term memory (LSTM). An extensive literature review revealed limited studies using ensemble methods to predict the prices of multiple cryptocurrencies and hyper-tune the parameters to improve forecasting accuracy (Derbentseva et al., 2021).

Furthermore, the existing literature predominantly focuses mostly on Bitcoin. Empirical research is necessary to test the forecasting models in a broader range of cryptocurrencies that display diverse price dynamics and volatility, which have evolved over the past few years (Piryonesi and El-Diraby, 2020). ML has the power to identify complex patterns in price movements and overcome the limitations of traditional models to improve forecasting accuracy. In contrast, traditional models like ARIMA may exhibit reliability in some cases (Nakano et al., 2018). Ensemble model combine the output from multiple models, hyper-tune the parameters and improve the forecasting accuracy. A systematic, three-stage approach is followed to create an ensemble model for predicting cryptocurrency prices. First, a dataset of historical prices of the 15 cryptocurrencies selected based on their market capitalization. Then, the base models, random forest regressor (RFR), gradient boosting regressor (GBR), support vector regression (SVR) and MLP regressor, are trained on the training set, optimizing their hyper-parameters. With the base models trained, predictions on the test set are made using each model. Finally, these individual predictions are combined using stacking to create the ensemble model. Root mean squared error (RMSE), mean squared error (MSE), mean absolute error (MAE) and R2 are used to evaluate the ensemble model’s performance on the 15 cryptocurrencies.

The study addresses the critical gaps in the existing literature. First, the existing studies have explored the statistical or ML models in silos. The traditional statistical models fail to capture the non-linear and rapid price movement of cryptocurrencies. Further it carries limitations such as hyperparameter tuning and lacks computational efficiency. The standalone ML models suffer from issues such as overfitting or underfitting adversely affecting the forecasting accuracy. This study bridges the gap by applying a robust ensemble machine leveraging the potential of ML to improve forecasting accuracy. Further the existing studies on forecasting, focus mostly on a single cryptocurrency, Bitcoin in general and few others like Ethereum. The broader cryptocurrency market remains largely unresearched. The study provides a methodological framework to empirically test the forecasting accuracy of ensemble model on a broader set of cryptocurrencies. Ensemble models are flexible in capturing non-linear relationships, complex patterns in the data and provide superior forecasting performance as compared to the standalone models. This paper contributes to the existing theories on predictive analytics, providing empirical evidence by creating an ensemble ML model. In addition, it contributes to the body of knowledge on how to leverage bias-variance trade-offs in forecasting. To fill the research gap, the study aims to address the following questions: (1) How to develop an ensemble model to accurately predict the prices of multiple cryptocurrencies? (2) Does the ensemble model outperform the statistical and econometric model in improving forecasting accuracy? (3) Does the ensemble model provide a better estimate of future price on a broader set of multiple cryptocurrencies that differ in return and volatility dynamics?

The rest of the paper is structured as follows. A brief literature review is detailed in Section 2, outlining the previous papers that used various models to predict the price of cryptocurrencies. Section 3 explores the methodology and the ensemble prediction approach of this study, whereas Section 4 presents the empirical findings, including a comparison of the proposed ensemble method with econometric and ML models. The paper concludes with Section 5, providing avenues for future research in forecasting.

Over the last decade, cryptocurrencies have gone up from an obscure asset to a preferred investment due to millennials' popularity, increased risk-taking propensity, perceived profitability, decentralization and acceptance of cryptocurrency as an investment option.

Cryptocurrency price forecasting is considered one of the financial domain’s most difficult predictions (Livieris et al., 2021). Challenges, complexity and interpretability of models have been classified as a problem of time series forecasting by most successful researchers (Adegboruwa et al., 2019; Hamayel and Owda, 2021; Lahmiri and Bekiros, 2019; Patel et al., 2020; Tandon et al., 2019; Wirawan et al., 2019).

This literature review examines the existing research studies on cryptocurrency price forecasting and the various possible methodologies for the same. Traditional statistical and econometric models have been considered for the analysis, which goes on to assess promising ML techniques. Finally, the potential of combining these methodologies to create hybrid models is analyzed.

2.1.1 Traditional statistical, econometric and ML models

The early research on cryptocurrency price forecasting relied on traditional econometric and statistical models. Efficient and immediate solutions for time series problems can be achieved through statistical models (Meenu et al., 2020). Econometric approaches that combine statistical and economic principles, according to Alahmari (2019) are crucial to estimating and predicting economic variables relevant to cryptocurrency prices. Time-series analysis often employs the ARIMA model, which, as discussed above, is a widespread technique (El-Bannany et al., 2020). Yet, the statistical and econometric models in forecasting come along with limitations. The statistical models rely on specific assumptions such as the normal distribution of data, the linear relationship between the variables and the stability of parameters over time. Thus, statistical models perform poorly either with noisy data or with changes in the trends and patterns.

Nikou et al. (2019) found that the ARIMA model, while suitable for price prediction within specific periods, struggles with capturing sharp price fluctuations. Similarly, Greaves and Au (2015) also observed significant prediction errors with ARIMA, particularly in the cases pertaining to volatile markets. Moreover, statistical and econometric models work within a specific framework, thus limiting the model’s ability to handle complexities in real-world data. The presence of non-linear relationships between variables adversely affects the forecasting accuracy.

ML techniques have gained significant traction in cryptocurrency price forecasting as a consequence of advancements in big data technology and artificial intelligence. These techniques analyze complex data patterns and make predictions powerfully (Chevallier et al., 2021; Derbentseva et al., 2021).

2.1.2 Integration of techniques (ensemble models)

There is a solid foundation through traditional statistical and econometric models to understand trends and relationships of financial data. However, cryptocurrency markets present the complexity of non-linear patterns that traditional models might not capture. ML techniques have the ability to identify such patterns with ease, only when provided with large datasets and careful tuning to avoid overfitting.

The benefits of statistical models lie in providing a foundational framework and distinguishing patterns and trends. For ML, the same reflects in uncovering the intricacies of complex non-linear relationships within the data. Through a combined approach, it is possible to tackle the limitations of individual models and optimize prediction accuracy (Madan et al., 2015).

Hybrid models that combine traditional and ML approaches can outperform classical ML models in terms of accuracy (Alahmari, 2019). The strengths of both methodologies come together to optimize accuracy in such a hybrid model.

2.1.3 Effects of information shocks on cryptocurrency prices

Wang et al. (2020) studied the relationship between economic policy uncertainty (EPU) and Bitcoin prices. Their findings suggest that EPU has a negative impact on Bitcoin prices in the short term. This effect gradually diminishes over time and has no notable long-term effects. Similarly, positive information shocks lead to higher volatility in the short run, while negative shocks have left no significant impact on cryptocurrency prices (Yin et al., 2021)

Cryptocurrency markets are known for their intense fluctuations and unpredictability. A predictive model which can efficiently navigate price fluctuations is required for effective cryptocurrency price prediction. Comparative studies of statistical models and their strengths and weaknesses are lacking (Lobell and Burke, 2010). ML models are dynamic and can deal with complexities to predict the price with a high degree of accuracy. Researchers used ML models such as DT, GB, SVR and MLP. Contradicting the researchers who supported ML models for price forecasting, Cohn et al. (1996) have concluded that statistically-based learning architectures combined with a mixture of Gaussians and locally weighted regression improve forecasting accuracy. Feature engineering and feature extraction using ML models can improve forecasting accuracy.

Ensemble model synthesize the output from multiple ML models to make an accurate forecast. It works like seeking advice from many sources thus improving the forecasting accuracy. The model improves the forecasting accuracy by merging predictions from multiple models. The learning rate of ensemble model can be effectively improved through optimization and hyperparameter tuning. By combining diverse models, ensemble forecasting accuracy can be improved using features such as maximum vote, stacking, blending, bagging and boosting. The combination of multiple algorithms in an ensemble model makes it capable of outperforming single-base models in predicting cryptocurrency (Yang et al., 2022).

Primarily differing in their approach to value stability and underlying mechanisms, stablecoins and cryptocurrencies are quite different. Stablecoins are engineered to mitigate price volatility by pegging their worth to a stable asset or fiat currency such as the US dollar or Euro. This ensures their consistency in value. This research paper focuses on the prediction of cryptocurrency prices due to their volatility when compared to stable asset prices. By focusing solely on non-stablecoin cryptocurrencies, our research avoids potential complexities introduced by stablecoins, whose value is pegged to external assets and may exhibit less price volatility compared to other cryptocurrencies.

Daily price data of 15 coins, based on market capitalization, is collected to build the forecasting model starting from January 2018 to September 2022. The coins selected for the study are Ethereum (ETH), Binance (BNB), Ripple (XRP), Dogecoin (DOGE), TRON (TRX), Litecoin (LTC), Ethereum Classic (ETC), Stellar (XLM), Monero (XMR), File coin (FIL), Decentraland (MANA), ZCash (ZEC), Dash (DASH), NEM coin (XEM) and Mona coin (MONA), all priced in USD. Data regarding the prices is obtained from Coinmarketcap as it offers comprehensive coverage and is widely used in cryptocurrency research (Vidal-Tomás, 2022).

The research methodology is divided into four stages. The first stage is to examine the forecasting accuracy of statistical models and refine them further to improve the forecasting accuracy. We evaluate the forecasting accuracy of ARIMA in stage two. In Stage Three, ML algorithms, including GB, DT, RF, SVR and MLP, are employed.

The final stage combines the best-performing ML models to optimize the forecasting accuracy further and build the model. The three models are RF With eXtreme Gradient Boosting (XGBoost), Ensemble model with RF, GB, DT and multi-model regression (MLR). R-squared, RMSE, MSE and mean absolute error were employed to evaluate the performance of the models.

4.2.1 Simple linear regression (SLR)

A SLR model is applied to forecast future prices when there is a linear relationship between dependent and independent variables. The equation has the form Y = a + bX, where Y is the dependent variable, i.e. the price of cryptocurrencies and X is the independent variable (time), b is the slope of the line and a is the y-intercept.

4.2.2 Auto-regressive integrated moving average (ARIMA)

The ARIMA model, a stochastic process, includes sums of autoregressive and moving average components. The model is applicable when there is autocorrelation among the residuals. The model includes three parameters, i.e. p, d and q.

  • p-auto-regressive component

  • d-order of differencing

  • q-moving average component.

ARIMA model specification

(1)

4.2.3 Random forest (RF)

RF, a supervised learning algorithm, is applicable to solving classification and regression problems. The first work on RF goes back to Ho (1995). Breiman (2001) further developed the idea and presented the RF algorithm in 2001. Various DTs are created on randomly selected samples, the forecasted results are produced from each tree and then the best tree is chosen by voting. The key benefits of RF are its generalization capability and minimal sensitivity to hyperparameters (Naing and Htike, 2015). In the prediction of cryptocurrency prices, RF has been used for BTC forecasting (Chevallier et al., 2021) and BTC, ETH and XRP in (Derbentseva et al., 2021).

4.2.4 Decision tree

DT, a type of supervised ML model, made of decision nodes and leaf nodes, resembling a tree-like structure with branches and leaves. The decision node represents features and attributes, and the leaf node represents the outcome. Inducing DT is one of the oldest and most popular techniques for learning discriminatory models, which has been developed independently in the statistical (Breiman, 2017; Kass, 1980) and ML (Hunt et al., 1966; Quinlan, 1983, 1986) communities.

4.2.5 Gradient boosting (GB)

GB, an ML technique, is used to solve regression and classification issues. A collection of weak prediction models, often DT, are produced as outcomes (Madeh and El-Diraby, 2021). GB frequently outperforms RF (Madeh and El-Diraby, 2021; Piryonesi and El-Diraby, 2020), allowing optimization of any differentiable loss function.

The loss function and the base-learner models can be arbitrarily specified on demand. Given a loss function denoted by Ψ(y, f) and/or a custom base-learner h(x, θ), obtaining a solution to the parameter estimates can be difficult.

4.2.6 Support vector regressor (SVR)

SVR is used for regression tasks to minimize forecasting errors. It operates by identifying a hyperplane with maximum margin from the training data. Support vectors, kernel functions and hyperparameter tuning are employed to capture complex relationships and handle outliers. This algorithm can powerfully predict continuous numerical values. Built on support vector machines for classification, SVR enables both linear and non-linear regressions. Examples of SVR usage in forecasting cryptocurrency prices can be found in (Chevallier et al., 2021; Khedr et al., 2021).

4.2.7 Multi-layer perceptron (MLP) regressor

MLP regressor processes input data with multiple layers of interconnected nodes (neurons) to form non-linear transformations. It employs the backpropagation algorithm to adjust the weights and biases of the network, optimizing the model to minimize the error. Many complicated problems are not linearly separable, so the single hidden layer perceptron, which can only solve linearly separable problems, is not effective (Rosenblatt, 1958). These issues require a solution with multiple hidden layers (Rumelhart et al., 1986). The MPLNN, a feedforward neural network with multiple layers, comprises three layers: an input layer, one or more hidden layers and an output layer.

4.2.8 Ensemble models

Three ensemble models have been created and tested to enhance the accuracy of this study.

4.2.8.1 RF ensembles with XGBoost

GB is used in this approach to train an RF ensemble. The boosting technique is applied to the DTs in RF ensembles by updating the weights of the observations based on the ensemble residuals. Then, each tree in the ensemble is trained to correct the errors of the previous trees. This is done by using the updated weights of the observations. Such an iterative process continues as long as the ensemble’s performance continues to improve.

The final ensemble is expected to be more accurate than a simple RF, as the GB algorithm is used to train the DTs.

4.2.8.2 Averaging ensemble model with RF, GB and DT

This method integrated multiple ML model forecasts to increase the forecast’s overall accuracy. This ensemble technique is built to individually develop each of the three models, namely RF, GB and DT, which will then return the average of their predictions. Each model may cause errors, so this method can balance them. This practice increases the overall accuracy of the ensemble model’s predictions. Each model is trained with a distinct set of algorithms and parameters on the same data set.

4.2.8.3 Multi-model regression model (MLR)

This approach employs the RFR, the GB, the SVR and the MLP regressor to form an MLR. Each model’s capabilities enhance this ensemble model to increase forecast accuracy. A meta-features data frame is generated by the MLR, which trains a linear regression meta-model using the ensemble of these models. This effectively captures varied patterns and correlations in cryptocurrency data. The MLR technique’s success in lowering prediction errors and offering insights into model performance is validated through the RMSE, MSE, MAE and R2 assessment measures.

Table 1 shows the descriptive statistics of 15 cryptocurrencies. Generally, cryptocurrencies show high volatility and also create volatility clustering. A perfect normal distribution data series has a skewness value of 0, and kurtosis equals 3. The skewness, kurtosis and Jarque–Bera results indicate that cryptocurrencies’ price movements do not follow a normal distribution. This makes it evident that investors face few larger gains and many periods of loss.

Table 1

Descriptive statistics

CoinsSkewnessKurtosisJarque-Bera
ETH_USD1.1485573.094957381.8956
BNB_USD1.0732682.716866338.6924
XRP_USD2.29587612.240737692.837
DOGE_USD2.0874967.8193372937.438
TRX_USD1.2087054.660242621.3699
LTC_USD1.1596864.068617471.173
ETC_USD1.6763786.3758991635.572
XLM_USD1.1862344.21978514.1647
XMR_USD0.943093.44275271.205
FIL_USD2.7087311.431447256.654
MANA_USD2.2528847.7572343101.925
ZEC_USD2.68768613.109489471.696
DASH_USD3.35219917.5383318518.55
XEM_USD4.24598927.4081148253.59
MONA_USD2.83572813.9747411026.09

Note(s): Table showing Descriptive statistics of cryptocurrencies

Source(s): Table by authors

Table 2 shows the MAE of cryptocurrency prices. Upon examining the results, it is evident that the performance of the models varies across different coins. Starting with XRP, the RF model exhibits the lowest MAE of 0.019, outperforming other models such as ARIMA, SLR, DT, SVR and MLP. The RF model consistently demonstrates superior accuracy with comparatively lower MAE values than other models for BNB, MONA, XLM, DASH, XMR, FIL and TRX. However, the MLP model shows the lowest MAE values for ETH, LTC, ZEC, MANA and ETC, indicating its effectiveness in predicting the prices. On the other hand, SLR has high MAE values for most coins except for XRP. XRP shows a relatively low MAE compared to other models through SLR. Both RF ensemble with XGBoost and averaging ensemble of RF, GB and DT show low MAE demonstrating that the model performs well for most cryptocurrencies.

Table 2

Mean absolute error

CoinARIMASLRRFDTGBSVRMLPRF with XGBoostAvg RF, GB and DT
XRP98.2880.2640.0190.0240.0210.1080.1830.055790.008
BNB13.19297.6654.615.6715.69660.99111.3213.325644.659
MONA0.0310.7130.0540.0690.0620.3120.57925.634295.134
XLM0.0090.1130.0080.0090.0080.0590.12920.206794.259
DASH0.00395.755.8797.2116.53156.31840.015160.004
ETH5.45473938.61446.38941.215724.5775.941.76226.805
XMR2.06865.6694.9166.2575.87935.5461.330.012830.002
FIL0.0118.5031.11.3461.1679.25719.050.116220.023
LTC8.74251.083.84.4994.22935.3450.073.886211.086
DOGE2.2940.060.0040.0040.0040.070.07724.220425.544
TRX0.8350.020.0020.0020.0020.0690.0670.374560.062
XEM6.3950.1110.0060.0080.0070.0620.06523.804695.818
ZEC6.21268.5364.9345.8425.57235.4559.540.043090.009
MANA0.0060.5050.0260.0330.0310.2540.5580.087720.029
ETC0.03813.1671.0351.1620.9757.47811.626.771171.212

Note(s): Table showing the Mean Absolute Error of cryptocurrencies

Source(s): Table by authors

According to Table 3 on the MSE of cryptocurrency prices, the results vary in accordance with the performance of models across different coins. First, the RF model outperforms other models such as ARIMA, SLR, DT, GB, RF ensemble with XGBoost and averaging ensemble of RF, GB and DT for XRP, demonstrating the lowest MSE of 0.002. However, it is essential to note that it is unusual for the MLP model to show a negative MSE value for XRP. This might indicate an issue with the model or the data. The RF model consistently exhibits the lowest MSE values for BNB, MONA, XLM, DASH, XMR, FIL and TRX, indicating its superior accuracy when compared to other models. For ETH, LTC, ZEC, MANA and ETC, the MLP model, conversely, showcases the lowest MSE values among the models considered and stands out as the most accurate. It is worth mentioning that the SLR model shows a relatively low MSE value for XRP and performs poorly for most coins, resulting in a significant rise in MSE values compared to other models.

Table 3

Mean squared error

CoinARIMASLRRFDTGBSVRMLPRF with XGBoostAvg RF, GB and DT
XRP189310.1050.0020.0030.0020.58−1.670.00770.000
BNB393.715799117.392197.296213.0910.09−0.91126782.86882.085
MONA0.0030.9330.0140.030.0180.750.022553.567140.136
XLM00.0190000.62−10.21959.323261.044
DASH014219161.275239.256156.187−4.5−24.590.001270.000
ETH99.239815760.55419.3058791.7466711−23.35−0.793470.993186.544
XMR16.7056763.32477.521142.697117.923−0.23−1.270.00030.000
FIL0878.6767.57412.6039.638−1.23−3.530.026450.002
LTC191.464001.80653.79979.78269.582−2.23−3.2767.3312614.105
DOGE17.4430.0090000.4−52.481185.88891.333
TRX1.5920.001000−593−2.070.262020.017
XEM108.510.02100.00100.67−0.031243.139112.427
ZEC126.387103.62397.882139.455112.793−1.23−11.310.004090.000
MANA00.5640.0050.0080.0080.44−5.230.038880.006
ETC0.007329.03114.37915.21312.707−0.93−3.74235.60929.808

Note(s): Table showing mean squared error of cryptocurrencies

Source(s): Table by authors

Table 4 results derive variations in the performance of models among different coins. The RF model outperforms other models such as ARIMA, SLR, DT, GB, SVR and MLP and achieves the lowest RMSE for XRP at 0.052. However, it is essential to note that the SLR model presents a relatively low RMSE for XRP. The RF model consistently exhibits the lowest RMSE values for BNB, MONA, XLM, DASH, XMR, FIL, TRX and ZEC, indicating its superior accuracy compared to other models. For ETH, LTC, MANA and ETC, the MLP model conversely stands out as the most accurate. It showcases the lowest RMSE values among the models considered. The ARIMA model shows relatively high RMSE values across the board, while the SVR model performs moderately well for most coins. Additionally, except for XRP and TRX, where it performs relatively better, the SLR model exhibits higher RMSE values for most coins. RF ensemble with XGBoost and Averaging ensemble of RF, GB and DT has relatively low RMSE demonstrating that the model performs well for most cryptocurrencies.

Table 4

Root mean squared error

CoinARIMASLRRFDTGBSVRMLPRF with XGBoostAvg RF, GB and DT
XRP137.590.3230.0520.0430.0410.1710.2740.087770.020
BNB19.842125.714.04614.59810.835109.593136.629356.065782.958
MONA0.0540.9660.1730.1350.1180.4750.69550.5328311.838
XLM0.020.1370.0210.0170.0160.0750.15130.972947.813
DASH0.004119.2515.46812.49712.699100.591145.2170.035590.015
ETH9.962903.293.76481.92273.6161212.689925.48858.9151313.658
XMR4.08782.23911.94610.8598.80556.47785.6820.017240.004
FIL0.01929.6423.553.1052.75222.93129.5210.162640.044
LTC13.83763.268.9328.3427.33549.15569.1468.205563.756
DOGE4.1770.0940.0150.0140.0160.0850.10634.436739.557
TRX1.2620.0260.0040.0040.0030.0740.0720.511880.129
XEM10.4170.1430.0240.020.0190.0790.10335.2581710.603
ZEC11.24284.28311.80910.629.89459.68697.8130.063960.018
MANA0.010.7510.0890.0880.0730.5060.8060.197190.080
ETC0.08418.1393.93.5653.79215.18717.3115.349573.132

Note(s): Table showing root mean squared error of cryptocurrencies

Source(s): Table by authors

The results of Table 5 show that the RF, DT and GB models have strong predictive power, due to high R-squared values across most coins. Their ability to account for a considerable proportion of the fluctuations in cryptocurrency prices is evident in these models. The ARIMA model has limited explanatory capability for cryptocurrency price movements, as it exhibits comparatively lower R-squared values for all coins. Some coins have relatively higher R-squared values through the SLR model, while others have significantly negative values. Hence, the results are mixed. It is important to note that the SLR model performs worse than a horizontal line (mean) in fitting the data when R-squared values are negative. This implies that the relationships between those specific coins’ dependent and independent variables cannot be captured adequately by the SLR model. There is varied performance across different coins with the MLP model as well. It achieves high R-squared values for some coins, such as BNB and ETH, and presents negative R-squared values for others. This suggests that for those particular coins, the MLP model may not effectively capture the underlying patterns in cryptocurrency prices. For RF ensemble with XGBoost and averaging ensemble model with RF, GB and DT show reasonably high R2 values indicate that the model can explain a considerable percentage of the data variation. However, when certain cryptocurrencies, such as XRP, have lower R2 values for RF ensemble with XGBoost, it is inferred that the model cannot explain a large variation in the data for this coin. Thus, this model’s prediction accuracy is quite poor for this currency. The averaging ensemble model with RF, GB and DT model has R2 values are typically near one, indicating that the model can accurately predict the target variable.

Table 5

R-Squared

CoinARIMASLRRFDTGBSVRMLPRF with XG BoostAvg RF, GB and DT
XRP0−356.350.980.980.980.58−1.670.250.980
BNB0.030.140.990.9910.09−0.910.911.000
MONA0.03−1.440.980.970.990.750.020.921.000
XLM0.01−1029.320.980.980.990.62−10.210.670.980
DASH0.13−1.550.990.980.99−4.5−24.590.870.980
ETH0−0.2610.991−23.35−0.790.70.990
XMR0.01−15.390.980.980.99−0.23−1.270.450.980
FIL0.01−6.30.990.990.99−1.23−3.530.620.980
LTC0.01−71.370.980.980.99−2.23−3.270.780.960
DOGE0.08−1.890.980.980.980.4−52.480.790.990
TRX0.99−3.050.990.980.99−593.2−2.070.640.980
XEM0.01−5.60.980.970.980.67−0.030.750.980
ZEC0.01−5.440.980.980.99−1.23−11.310.70.980
MANA0.05−0.920.990.990.990.44−5.230.950.990
ETC0.02−4.20.970.960.96−0.93−3.740.680.990

Note(s): Table showing R Squared of cryptocurrencies

Source(s): Table by authors

Table 6 and Figure 2 shows the forecasting results of MLR for pre and post-COVID crisis. The results show that the model consistently achieves exceptional performance across all coins. Extremely low values of MAE, MSE and RMSE indicate its high accuracy in predicting cryptocurrency prices covering pre and post-crisis. Additionally, it can be concluded that the model can perfectly explain the variance in cryptocurrency prices, indicating its strong explanatory power, given that the R2 value of all coins is 1.00. The significant performance of the MLR in accurately forecasting cryptocurrency prices across different coins, as the table highlights, is overall indicative of its effectiveness.

Table 6

Multi-model regression model

CoinPre MAEPre MSEPre RMSEPre R2Post MAEPost MSEPost RMSEPost R2Con MAECon MSECon RMSECon R2
XRP0.020.0010.0350.9910.0070.0000.0100.9991.0358.8282.9711
BNB0.2520.1510.3880.9973.26720.1674.4910.9980.1980.3810.6171
MONA0.0690.0140.1190.9950.0060.0000.0110.9990.00100.0031
XLM0.00800.0150.9870.0020.0000.0030.9990.00200.0040.99
DASH7.662198.40414.0860.9951.1733.3451.8290.9990001
ETH7.223181.37413.4680.99724.6511107.433.2770.9990.3340.4390.6631
XMR3.37234.6075.8830.9951.9507.2812.6980.9970.2950.2660.5160.99
FIL0.190.1080.3290.9960.3750.4320.6570.999000.0021
LTC2.37113.6743.6980.9941.2313.6501.9100.9990.1240.0950.3081
DOGE0000.9910.0010.0000.0020.9990.0670.0410.2031
TRX0.00100.0030.9780.0010.0000.0010.9960.00500.0121
XEM0.0130.0010.0290.9870.0010.0000.0020.9990.4250.6260.7911
ZEC5.533118.97610.9080.9931.5255.4792.3410.9980.7861.7221.3121
MANA0.00200.0040.9890.0310.0030.0520.9980.00100.0031
ETC0.2630.3640.6030.9940.3700.3000.5501.0000.00400.0091

Note(s): Table showing multi-model regression model for all cryptocurrencies

Source(s): Table by authors

Figure 2

Figure showing the multi-modal regression model performance for all cryptocurrencies

Figure 2

Figure showing the multi-modal regression model performance for all cryptocurrencies

Close modal

The descriptive statistics analysis, helped understand the nature of cryptocurrency price distributions. Cryptocurrencies are known for their high volatility and the presence of volatility clustering, as noted in the literature review. Aligning with findings from the previous research, the skewness and kurtosis values indicate deviations from a normal distribution (Nikou et al., 2019). This supports the need for robust forecasting models capable of handling such volatile and skewed data.

ARIMA, SLR, RF, DT, GB and MLR ensemble models performed well with coins such as XLM, FIL, MANA, DASH, XMR and ZEC. These models adequately reflect the above coins' trends, seasonality and historical price patterns. Furthermore, MLR demonstrates its advantages when applied to coins like BNB and MONA. SLR has performed well for cryptocurrencies that have significant market capitalization, strong liquidity and price movements that have historically been consistent, like XRP, TRX and DOGE. The MLR ensemble approach produces more accurate predictions for coins with higher volatility and irregular price patterns, such as ETH and ETC, by combining ML models to remove errors.

It is worth noting that the intricacies of each coin’s behavior may not be captured as these are broad observations. Macroeconomic, attractiveness and demand and supply variables are all factors that can impact a model’s efficacy. The literature review expresses concern regarding cryptocurrency price forecasting challenges, echoed by the findings, which demonstrate that different models perform variably across different cryptocurrencies. The analysis of the RF ensembles with XGBoost and averaging ensemble model with RF, GB and DT demonstrates the effectiveness of ensemble models combining multiple models for improved forecasting accuracy. The results support that ensemble models can enhance predictive performance. They often outperformed individual models across various cryptocurrencies in the current study.

Cryptocurrency price forecasting carries a significant influence on investment decisions. The present study makes contribution to the theoretical and practical implications on cryptocurrency investments. On the theoretical side, the study advances the conceptual framework on building predictive models by demonstrating the efficiency of ensemble model in forecasting. It shows a new dimension to predictive analytics by empirically validating the accuracy of multiple algorithms pre-COVID and post-COVID period. On the practical side, it offers a valuable forecasting tool to the traders and investors in making informed decision. The novelty of the research is in its attempt to empirically test the forecasting ability of ensemble model across time span and its comparison with econometric model on a broad-based set of cryptocurrencies. Consistent with the expectations, ensemble models created by combining RF, XGBoost, gradient boosting and decision tree (DT) reduce the variance in predicting errors indicating the forecasting accuracy. Interestingly, it is found that when the price movement exhibits autocorrelation, the ARIMA model is best applicable. Compared to other ML models applied in isolation the ensemble model performed better.

The present study contributes to the body of literature on ensemble methods in forecasting crypto price in general, potentially influencing future studies on price forecasting. The study motivates the researchers on empirical testing of our framework on various asset classes. As a result, on the prediction ability of ensemble model, the study will significantly influence the decision-making process of traders and investors. The research benefits the traders and investors to effectively develop a model to forecast cryptocurrency price. The findings highlight the potential of ensemble model in predicting high volatile cryptocurrencies and other financial assets. Investors can design the investment strategies and asset allocation decisions by understanding the relationship between market trends and consumer behavior. Investors can enhance portfolio performance and mitigate risk by incorporating these insights into their decision-making processes. Policymakers can use this information to design more effective regulations and policies promoting economic stability and consumer welfare. The study emphasizes the need for using diversified model to understand the market dynamics and improving trading strategies.

Although the dataset includes crisis period data, anomalies or outliers are yet to be explicitly excluded from the analysis. The models employed in this study still demonstrate high accuracy in predicting cryptocurrency prices despite these outliers, suggesting that the models are robust enough to handle unexpected fluctuations or extreme events in the market. However, the lack of specific analysis on the impact of outliers on model performance is a limitation of the study, as it needs to fully explore the resilience of the forecasting models under adverse market conditions.

Cryptocurrency markets can be influenced by investor sentiment, which can be reflected in price movements. By analyzing these price movements, investors might gain insights into overall market sentiment. However, exploring this link further is certainly an area for future research. While our study focuses on price prediction using the ensemble model, we acknowledge the need for further exploration to definitively link price movements to specific investor behaviors. Complementary techniques, such as sentiment analysis of social media data or transaction volume analysis, could be valuable tools for future research in this area.

Adegboruwa
,
T.I.
,
Adeshina
,
S.A.
and
Boukar
,
M.M.
(
2019
), “
Time series analysis and prediction of bitcoin using long short term memory neural network
”,
2019 15th International Conference on Electronics, Computer and Computation (ICECCO)
, pp. 
1
-
5
, doi: .
Alahmari
,
S.A.
(
2019
), “
Using machine learning ARIMA to predict the price of Cryptocurrencies$
”, Vol. 
11
No. 
3
, p.
6
.
Bouteska
,
A.
,
Abedin
,
M.Z.
,
Hajek
,
P.
and
Yuan
,
K.
(
2024
), “
Cryptocurrency price forecasting–a comparative analysis of ensemble learning and deep learning methods
”,
International Review of Financial Analysis
, Vol. 
92
, 103055, doi: .
Breiman
,
L.
(
2001
), “
Random forests
”,
Machine Learning
, Vol. 
45
No. 
1
, pp. 
5
-
32
, doi: .
Breiman
,
L.
(
2017
),
Classification and Regression Trees
,
Routledge
,
Boca Raton
.
Chen
,
J.
(
2023
), “
Analysis of bitcoin price prediction using machine learning
”,
Journal of Risk and Financial Management
, Vol. 
16
No. 
1
, p.
51
, doi: .
Chevallier
,
J.
,
Guégan
,
D.
and
Goutte
,
S.
(
2021
), “
Is it possible to forecast the price of bitcoin?
”,
Forecasting
, Vol. 
3
No. 
2
, pp. 
377
-
420
, doi: .
Cohn
,
D.A.
,
Ghahramani
,
Z.
and
Jordan
,
M.I.
(
1996
), “
Active learning with statistical models
”,
Journal of Artificial Intelligence Research
, Vol. 
4
, pp. 
129
-
145
, doi: .
Derbentseva
,
V.
,
Khrustalovac
,
S.
,
Babenko
,
V.
,
Khrustalev
,
K.
and
Obruch
,
H.
(
2021
), “
Comparative performance of machine learning ensemble algorithms for forecasting cryptocurrency prices
”,
International Journal of Engineering
, Vol. 
34
No. 
1
, doi: .
El-Bannany
,
M.
,
Sreedharan
,
M.
and
Khedr
,
A.M.
(
2020
), “
A robust deep learning model for financial distress prediction
”,
International Journal of Advanced Computer Science and Applications (IJACSA), The Science and Information (SAI) Organization Limited
, Vol. 
11
No. 
2
, doi: .
Greaves
,
A.
and
Au
,
B.
(
2015
), “
Using the bitcoin transaction graph to predict the price of bitcoin
”, p.
8
.
Hamayel
,
M.J.
and
Owda
,
A.Y.
(
2021
), “
A novel cryptocurrency price prediction model using GRU, LSTM and bi-LSTM machine learning algorithms
”,
AI, Multidisciplinary Digital Publishing Institute
, Vol. 
2
No. 
4
, pp. 
477
-
496
, doi: .
Ho
,
T.K.
(
1995
), “
Random decision forests
”,
Proceedings of the Third International Conference on Document Analysis and Recognition (Volume 1)–Volume 1
,
USA
,
IEEE Computer Society
, p.
278
.
Hunt
,
E.B.
,
Marin
,
J.
and
Stone
,
P.J.
(
1966
),
Experiments in Induction
,
Academic Press
,
Massachusetts
.
Kass
,
G.V.
(
1980
), “
An exploratory technique for investigating large quantities of categorical data
”,
Journal of the Royal Statistical Society: Series C (Applied Statistics)
, Vol. 
29
No. 
2
, pp. 
119
-
127
, doi: .
Kayani
,
U.N.
(
2023
), “
Exploring prospects of blockchain and fintech: using SLR approach
”,
Journal of Science and Technology Policy Management
, Vol.
ahead-of-print
No.
ahead-of-print
, doi: .
Kayani
,
U.
and
Hasan
,
F.
(
2024
), “
Unveiling cryptocurrency impact on financial markets and traditional banking systems: lessons for sustainable blockchain and interdisciplinary collaborations
”,
Journal of Risk and Financial Management
, Vol. 
17
No. 
2
, p.
58
, doi: .
Khedr
,
A.M.
,
Arif
,
I.
,
El‐Bannany
,
M.
,
Alhashmi
,
S.M.
and
Sreedharan
,
M.
(
2021
), “
Cryptocurrency price prediction using traditional statistical and machine‐learning techniques: a survey
”,
Intelligent Systems in Accounting, Finance and Management
, Vol. 
28
No. 
1
, pp. 
3
-
34
, doi: .
Lahmiri
,
S.
and
Bekiros
,
S.
(
2019
), “
Cryptocurrency forecasting with deep learning chaotic neural networks
”,
Chaos, Solitons and Fractals
, Vol. 
118
, pp. 
35
-
40
, doi: .
Li
,
Y.
,
Jiang
,
S.
,
Li
,
X.
and
Wang
,
S.
(
2022
), “
Hybrid data decomposition-based deep learning for Bitcoin prediction and algorithm trading
”,
Financial Innovation
, Vol. 
8
No. 
1
, p.
31
, doi: .
Livieris
,
I.E.
,
Kiriakidou
,
N.
,
Stavroyiannis
,
S.
and
Pintelas
,
P.
(
2021
), “
An advanced CNN-LSTM model for cryptocurrency forecasting
”,
Electronics
, Vol. 
10
No. 
3
, p.
287
, doi: .
Lobell
,
D.B.
and
Burke
,
M.B.
(
2010
), “
On the use of statistical models to predict crop yield responses to climate change
”,
Agricultural and Forest Meteorology
, Vol. 
150
No. 
11
, pp. 
1443
-
1452
, doi: .
Lyu
,
H.
(
2022
), “
Cryptocurrency price forecasting: a comparative study of machine learning model in short-term trading
”,
2022 Asia Conference on Algorithms, Computing and Machine Learning (CACML)
, pp. 
280
-
288
, doi: .
Madeh
,
P.S.
and
El-Diraby
,
T.E.
(
2021
), “
Using machine learning to examine impact of type of performance indicator on flexible pavement deterioration modeling
”,
Journal of Infrastructure Systems
, Vol. 
27
No. 
2
, 04021005, doi: .
Madan
,
I.
,
Saluja
,
S.
and
Zhao
,
A.
(
2015
), “
Automated bitcoin trading via machine learning algorithms
”, p.
5
.
Meenu
,
S.
,
Khedr
,
A.M.
and
El Bannany
,
M.
(
2020
), “
A multi-layer perceptron approach to financial distress prediction with genetic algorithm
”,
Automatic Control and Computer Sciences
, Vol. 
54
No. 
6
, pp. 
475
-
482
, doi: .
Naing
,
W.Y.N.
and
Htike
,
Z.Z.
(
2015
), “
Forecasting of monthly temperature variations using random forests
”,
ARPN Journal of Engineering and Applied Sciences
, Vol. 
10
No. 
21
, pp. 
10109
-
10112
.
Nakano
,
M.
,
Takahashi
,
A.
and
Takahashi
,
S.
(
2018
), “
Bitcoin technical trading with artificial neural network
”,
Physica A: Statistical Mechanics and its Applications
, Vol. 
510
, pp. 
587
-
609
, doi: .
Nikou
,
M.
,
Mansourfar
,
G.
and
Bagherzadeh
,
J.
(
2019
), “
Stock price prediction using DEEP learning algorithm and its comparison with machine learning algorithms
”,
Intelligent Systems in Accounting, Finance and Management
, Vol. 
26
No. 
4
, pp. 
164
-
174
, doi: .
Patel
,
M.M.
,
Tanwar
,
S.
,
Gupta
,
R.
and
Kumar
,
N.
(
2020
), “
A deep learning-based cryptocurrency price prediction scheme for financial institutions
”,
Journal of Information Security and Applications
, Vol. 
55
, 102583, doi: .
Piryonesi
,
S.M.
and
El-Diraby
,
T.E.
(
2020
), “
Data analytics in asset management: cost-effective prediction of the pavement condition index
”,
Journal of Infrastructure Systems
, Vol. 
26
No. 
1
, 04019036, doi: .
Quinlan
,
J.R.
(
1983
), “
Learning efficient classification procedures and their application to chess end games
”,
Machine Learning
, pp. 
463
-
482
, doi: .
Quinlan
,
J.R.
(
1986
), “
Induction of decision trees
”,
Machine Learning
, Vol. 
1
, pp. 
81
-
106
, doi: .
Rosenblatt
,
F.
(
1958
), “
The perceptron: a probabilistic model for information storage and organization in the brain
”,
Psychological Review
, Vol. 
65
No. 
6
, pp. 
386
-
408
, doi: .
Rumelhart
,
D.E.
,
Hinton
,
G.E.
and
Williams
,
R.J.
(
1986
), “
Learning representations by back-propagating errors
”,
Nature
, Vol. 
323
No. 
6088
, pp. 
533
-
536
, doi: .
Soltani
,
H.
,
Taleb
,
J.
and
Abbes
,
M.B.
(
2023
), “
The directional spillover effects and time-frequency nexus between stock markets, cryptocurrency, and investor sentiment during the COVID-19 pandemic
”,
European Journal of Management and Business Economics
, Vol.
ahead-of-print
No.
ahead-of-print
, doi: .
Suhartono
,
Rahayu
,
S.P.
,
Prastyo
,
D.D.
,
Wijayanti
,
D.G.P.
and
Juliyanto
(
2017
), “
Hybrid model for forecasting time series with trend, seasonal and salendar variation patterns
”,
Journal of Physics: Conference Series
, Vol. 
890
, 012160, doi: .
Tandon
,
S.
,
Tripathi
,
S.
,
Saraswat
,
P.
and
Dabas
,
C.
(
2019
), “
Bitcoin price forecasting using LSTM and 10-fold cross validation
”,
2019 International Conference on Signal Processing and Communication (ICSC)
, pp. 
323
-
328
, doi: .
Vidal-Tomás
,
D.
(
2022
), “
Which cryptocurrency data sources should scholars use?
”,
International Review of Financial Analysis
, Vol. 
81
, 102061, doi: .
Wang
,
P.
,
Li
,
X.
,
Shen
,
D.
and
Zhang
,
W.
(
2020
), “
How does economic policy uncertainty affect the bitcoin market?
”,
Research in International Business and Finance
, Vol. 
53
, 101234, doi: .
Wirawan
,
I.M.
,
Widiyaningtyas
,
T.
and
Hasan
,
M.M.
(
2019
), “
Short term prediction on bitcoin price using ARIMA method
”,
2019 International Seminar on Application for Technology of Information and Communication (iSemantic)
,
Semarang, Indonesia
,
IEEE
, pp. 
260
-
265
, doi: .
Yang
,
Z.
,
Li
,
L.
,
Xu
,
X.
,
Kailkhura
,
B.
,
Xie
,
T.
and
Li
,
B.
(
2022
), “
On the certified robustness for ensemble models and beyond
”,
arXiv
, doi: , available at: http://arxiv.org/abs/2107.10873
(accessed
 1 July 2023).
Yin
,
L.
,
Nie
,
J.
and
Han
,
L.
(
2021
), “
Understanding cryptocurrency volatility: the role of oil market shocks
”,
International Review of Economics and Finance
, Vol. 
72
, pp. 
233
-
253
, doi: .
Published in European Journal of Management and Business Economics. Published by Emerald Publishing Limited. This article is published under the Creative Commons Attribution (CC BY 4.0) licence. Anyone may reproduce, distribute, translate and create derivative works of this article (for both commercial and non-commercial purposes), subject to full attribution to the original publication and authors. The full terms of this licence may be seen at http://creativecommons.org/licences/by/4.0/legalcode

or Create an Account

Close Modal
Close Modal