This study aims to introduce an innovative multi-channel imaging technique aimed at mitigating deep learning overfitting and facilitating the automatic extraction of features from limited 1D data.
The proposed framework consists of five key component: dimensionality reduction, sequence image generation, image stitching, feature extraction and model training. It converts 1D multi-temporal data into multiple 2D images utilizing Markov transition field, Gramian angular field and recurrence plot. These single-channel images are stitched into a larger n-channel image, which is processed by a convolutional neural network for feature extraction and forecasted using a long short-term memory network.
The results demonstrate that the proposed multi-channel imaging technique outperforms all benchmark models. This conversion captures underlying patterns and enhances information transmission. Additionally, models with multi-time series configurations perform better than single-time series setups, highlighting that data are more crucial than advanced models in forecasting.
This pioneering study explores the role of non-economic variables in tourism forecasting. The proposed multi-channel time series imaging model not only applies to tourism but also offers potential for interdisciplinary applications.
1. Introduction
Tourism accelerates economic growth and promotes the development of other related industries (Gunter and Önder, 2015). Tourism resources’ high dependence and deep interactions result in an increased chance of exposure to risks in the face of a crisis (De et al., 2025). Tourism products and services are perishable, intangible, and indivisible compared with other fields (Goh and Law, 2002; Abellana et al., 2021). Accordingly, precise forecasting data are indispensable for industrial applications and represent a key focus for academic research (Hewapathirana, 2023). These data facilitate informed decision-making among tourism stakeholders, allowing for risk mitigation, optimal resource allocation, and economic enhancement (Liu et al., 2022; Zheng et al., 2024).
Researchers have developed a range of models to enhance accuracy (Song and Li, 2008). The proficiency of artificial intelligence (AI)-based models in elucidating non-linear datasets, coupled with their capacity to mitigate the limitations inherent in time series and econometric models, has catalyzed their widespread adoption across diverse scientific fields (Díaz and Mateu-Sbert, 2011). Deep learning, with its complex hidden layers and structures, can automatically build relevant features at different network levels. This mechanism is commonly used in solving unsupervised issues (Law et al., 2019) to extract more features from limited data (Hochreiter and Schmidhuber, 1997; Law et al., 2019). However, the over-fitting problem of deep learning cannot be ignored (Shi et al., 2017). Moreover, deep learning technologies can only create a 1D mapping of observations and predicted values with a lag order, resulting in the insufficient exploitation of some important features and a loss of important information (Bi et al., 2021). How to mine more informative features in limited data and solve the overfitting problem of deep learning is an effective means to enhance the accuracy of predictive models.
With the advancements of the computer vision technology, conversion of numeric time series data into images can efficiently mine the features reflecting the cyclic patterns and transient relationships in time series figures. However, the application of this technique in tourism demand forecasting is remains limited, and the data used for this technique are only a single variable, such as number of tourists (Bi et al., 2021). The robustness and generalization of the model can be guaranteed while addressing overfitting by integrating data from multiple sources and expanding the data sample (Goodfellow et al., 2016). This situation holds true for deep learning models using multi-layer neural networks, where the growth of network parameters requires additional available training data. Otherwise, the model will begin to capture noise, damaging the forecasting performance (Shi et al., 2017).
This study designs to construct a multi-channel imaging model to address the identified gaps. Appropriate multisource time series data are chosen to create a multivariate dataset, guided by the psychosocial framework for demand generation as conceptualized by Um and Crompton (1990). After composing multi-source data using the above variables, the proposed model (Figure 1) can expand data diversity and address deep learning over-fitting. Multi-channel imaging technology and image aggregation are used to solve the problem of visual feature extraction from limited structured data. The model has five parts: (1) dimensionality reduction; (2) convert multi-source 1D data into 2D images; (3) aggregate 2D images into a larger 2D image; (4) extract image features; and (5) train forecasting model. The initial step involves utilizing principal component analysis (PCA) to minimize the dimensionality of multi-source heterogeneous data. Subsequently, the dataset is transformed into image-compatible formats through encoding techniques, specifically the Markov transition field (MTF), Gramian angle field (GAF), and Recurrence plot (RP). Thereafter, these n single-dimensional channel images are stitched together into one large n-dimensional channel image and fed into the convolutional and pooling layers to extract features. Finally, the tensor created by the CNN is fed into a Long Short-term Memory Network model (LSTM) and the model is trained to predict travel demand.
2. Literature reviews
2.1 Tourism demand forecasting with multivariate time series
Tourism demand forecasting involves projecting future tourists by analyzing historical time series data and incorporating various pertinent factors and indicators. Forecasting studies can be divided into single or multiple time series depending on the data type. Models using multivariate time series can observe more data types compared with time series data, increasing degrees of freedom, declining multicollinearity, reducing variable bias that may be overlooked, and improving forecasting efficiency (Kuo et al., 2009; Song et al., 2019).
This study selects the features of news, climate, the Baidu index, and holidays through the social psychological framework model proposed by Um and Crompton (1990) to form the multimodal dataset. The following section provides an academic review of research pertaining to these variables in tourism demand forecasting. An essential challenge for tourism demand is the seasonality of destinations (Goh, 2012). The seasonality of tourism demand stems from natural factors in climate and institutional factors in holidays (Koenig-Lewis and Bischoff, 2005; Rudihartmann, 1986). Song et al. (2000) suggested that tourists’ choices are significantly shaped by social and psychological factors. Climate has been identified as a critical determinant in the selection of tourist destinations (Wall and Badke, 1994). The holidays’ enjoyment remains a central motivator in travel decision-making (Heung et al., 2001). Numerous studies combine environmental weather data to create a tourism demand forecasting dataset (Chen et al., 2017; Khatibi et al., 2020). Hu et al. (2021) distinguished holiday and non-holiday data in the time series by hierarchical pattern recognition and built different models based on historical data.
The Engel, Blackwell, and Miniard (EBM) model suggests that information search is a key early process that drives consumer decision making (Teo and Yeong, 2003). Online search behavior provides a valuable indication of prospective intentions (Höpken et al., 2021; Pan and Yang, 2017). The growing use of internet-based data offers a complementary resource for forecasting tourism demand (Wen et al., 2021). The combination of search engine indices and historical time series data has been featured in many articles (Li et al., 2021; Liu et al., 2021; De Luca and Rosciano, 2024). News has garnered academic interest as a measure of the stability of the broader social environment, offering a more subtle and consistent indicator compared with other data sources (Starosta et al., 2019). However, news research on tourism demand is limited, with only Park et al. (2021) examining how news coverage affects tourist numbers in Hong Kong from the US and China. Chen et al. (2024) argued that including news coverage can effectively improve tourism demand forecasting.
2.2 Data encoding as images
Researchers have started transforming time series data into visual formats, revealing patterns that may not be discernible in the raw 1D data. This approach is widely used in classification. Hatami et al. (2018) applied RP to shape 2D texture images and classified the UCR time series dataset using the CNN’s classifier. The performance was more competitive compared with the advanced classification algorithms. Martínez-Arellano et al. (2019) also used RP to classify tool condition monitoring data. The results illustrate that the images recreate the original time series without losing information, and deep learning can identify the intrinsic attributions of the primary sensory data.
The time series imaging method has also been recently applied in forecasting. Zhang and Guo (2020) converted data for weather conditions, residential building, and electricity consumption into GAF images, following dimensionality reduction via incremental kernel PCA (IKPCA). Hong et al. (2020) used GAF and convolutional LSTM to predict solar radiation time series data. After comparison with the benchmark model, they found that transforming 1D data into 2D image data enables the full advantage of computers in image processing. This transformation facilitates the comprehensive extraction of rich and diverse features inherent in the original dataset. Wang et al. (2022) forecasted the glucose concentration in the human blood by the GAF–CNN model. Li et al. (2020) converted time series into RP and used spatial bag-of-features to extract features from 2D graphs of M4 and tourism competition data. Li argued that automatic feature extraction reduces human intervention. RP will contain a more comprehensive view of the original data, resulting in higher predictive performance. In the tourism demand field, only one study used imaging techniques. Bi et al. (2021) combined the daily passenger flows into 2D images and utilized CNN to extract features and input them to the LSTM for prediction. Although converting 1D data into 2D images can improve the forecasting accuracy, multivariate time series data empirical research should be further developed and refined. The over-fitting problem in the forecasting models still exists. This study proposes an image stitch method, aggregating multiple 2D images into a large one. On this basis, the features of the multivariate time series can be extracted by CNN and fed into the forecasting model.
3. Methodology
This section briefly introduces the related methods including PCA, sequential time-series imaging, CNN-assisted stitching and feature extraction, and LSTM-driven model prediction.
3.1 PCA
PCA is a widely used unsupervised algorithm to find orthogonal directional axes between data with minimum error representation (Li et al., 2018). is the original information matrix with n sample data and m variables. PCA projects a high-dimensional feature dataset into a low-dimensional space with the following processes:
Step 1: Normalized by , where and are the mean and standard deviation.
Step 2: Alinear combination of k principal components by , where are the 1st, 2nd, …, and kth principal components of . is the eigenvector of the original variable located on the principal component and satisfies .
Step 3: Determine the number of principal components through and .
3.2 Data imaging generation
This section converts the multiple time series into 2D images based on GASF, GADF, MTF, and RP.
3.2.1 GAF
GAF utilizes polar coordinates for the representation of time series data. The core steps of GAF are divided into three main stages. Assume is a time series. The first step is to normalize . Step 2 entails transforming the normalized time series from Cartesian coordinates to polar coordinates using Equation (1). This process is akin to mapping these points into a unit circle, as illustrated in Figure 2. Step 3 represents the relationship between two points using either the sum of their angles (GASF, equation 2) or the difference between their angles (GADF equation 3). In GASF, brighter colors indicate a stronger relationship between the two points, while in GADF, brighter colors signify a notable change between them.
where is the timestamp, and N serves as a scaling constant determining the range of the polar coordinate system.
3.2.2 RP
RP (Figure 3) is an effective method for analyzing the similarity within time series data (Li et al., 2020). Its computation involves two main steps. First, the pairwise distances between data points () are calculated based on Equation (4). Second, a recurrence matrix is constructed to generate visualizationf the distance between two points is below a predefined threshold (), a point is marked at the corresponding position; otherwise, no mark is made.
where p is the trajectory’s dimension, and denotes the delay. i,j denote the time indices on the horizontal and vertical axes, respectively; and is a predetermined threshold value.
3.2.3 MTF
The core concept of MTF (Figure 4) is to discretize time series data into different intervals and compute the transition probabilities between these intervals. Specifically, the first step involves defining the intervals by dividing the time series () into intervals. The second step, based on Equation (5), involves constructing the state transition matrix () by counting the number of transitions between different intervals and calculating the corresponding conditional probabilities. Finally, MTF (equation 6) is generated using this state transition matrix. The resulting MTF is essentially an matrix, where each element represents the probability of transitioning from interval 2 to interval 3.
where is determined by the frequency, where is adjacent to the data in , is element of the Markov matrix M.
3.3 CNN
CNN is a good tool to process images. As illustrated in Figure 5, the CNN comprises three main steps for feature extraction. Given an image of size n × n, the process begins with convolution, which aims to extract local features from the image. Each convolutional kernel slides across the image and computes distinct features based on the specified formula (7). Where is the value of feature map of i,j. is the convolution operation of feature map X on convolution kernel K,
The pooling layer follows, serving to compress the feature values while preserving essential attributes. This study employs max pooling, which selects the maximum value within a defined region. Finally, the fully connected layer integrates the extracted features, transforming the input feature map into the output result.
3.4 LSTM neural network
LSTM, a modified form of RNN, adeptly addresses the challenges of long-term dependencies and gradient disappearance inherent in traditional RNNs. Originating from the seminal work of Hochreiter and Schmidhuber (1997), LSTM introduces forget gate, input gate and output gate to control the flow of information. The forget gate discards unimportant information. The input gate is responsible for selecting significant information to be stored. The output gate determines which information should be passed on to the next step or output result (Figure 6).
4. Empirical research
4.1 Data collection
Jiuzhaigou and Mount Siguniang have been selected as case studies to investigate the efficacy of the multi-channel imaging model when applied to multivariate datasets. The advent of COVID-19 has triggered a paradigm shift in the strategic planning and decision-making processes of individuals and societies globally (Gabriel-Campos et al., 2021). Separating the data before and after the outbreak is more accurate for building forecasting models due to the change in tourist motivation and behavior. The robustness of the model can also be further verified by building separate models for different periods. The data details are in Table 1.
Details of the dataset
| Sample | Period | Time | Data size | Mean value |
|---|---|---|---|---|
| Jiuzhaigou | Pre-COVID-19 | Dec. 25, 2015–Sept. 27, 2019 | 592 | 12536.147 |
| Post-COVID-19 | Mar. 31, 2020–Jan. 7, 2022 | 648 | 5715.907 | |
| Mount Siguniang | Pre-COVID-19 | Jan. 1, 2016–Nov. 30, 2018 | 1,065 | 1431.696 |
| Post-COVID-19 | Mar. 31, 2020–Dec. 31, 2021 | 641 | 1495.539 |
| Sample | Period | Time | Data size | Mean value |
|---|---|---|---|---|
| Jiuzhaigou | Pre-COVID-19 | Dec. 25, 2015–Sept. 27, 2019 | 592 | 12536.147 |
| Post-COVID-19 | Mar. 31, 2020–Jan. 7, 2022 | 648 | 5715.907 | |
| Mount Siguniang | Pre-COVID-19 | Jan. 1, 2016–Nov. 30, 2018 | 1,065 | 1431.696 |
| Post-COVID-19 | Mar. 31, 2020–Dec. 31, 2021 | 641 | 1495.539 |
Source(s): Authors’ own work
This study collects relevant variables based on the social psychological framework proposed by Um and Crompton (1990) (Figure 7). The model posits that external stimuli shape an individual’s awareness set. Social interactions and marketing communications serve as external inputs, which can be categorized into passive information catching and active information searching. News, functioning as a measure of the stability and turbulence of the social environment, significantly influences tourism demand. The Baidu search index reflects the active search behavior of travelers. Consequently, the study attributes both news and the Baidu index to the external input features of the model. Internal stimuli and situational constraints are crucial elements in converting the awareness set into an evoked set. Internal inputs stem from an underlying social psychological framework, such as motivation. Scholars have identified climate and air quality as significant motivational factors. Thus, climate (wind speed, minimum temperature, maximum temperature, average temperature, weather conditions) and air quality index (AQI) have been selected as the internal inputs for this study. Holidays, as situational constraints, play a pivotal role in the transition of tourists from the cognitive set to the evoked set. The sources for the relevant variables are as follows: the number of tourists (https://www.jiuzhai.com/; https://www.sgns.cn/), climate data (https://lishi.tianqi.com/), air quality index data (http://www.aqistudy.cn/historydata/), holiday data (https://wannianli.tianqi.com/fangjiaanpai/), search engine data (http://index.baidu.com/), and news data (WiseSearch). These sources contribute to the composition of a multi-source time series dataset.
For unstructured data (holidays and news), this study utilizes two methodologies to transform unstructured data into structured data that are amenable to deep learning algorithms. Dichotomous variable encoding serves as a straightforward and efficient method for distinguishing between different time periods, with “0” denoting workdays and “1” signifying holidays. In the context of news data, the study utilizes sentiment analysis to discern the emotional valence within news content, which is then integrated into the multivariate dataset.
4.2 Feature dimensionality reduction
The Baidu recommendation algorithm identifies 17 pivotal keywords from the Baidu index. Post-sentiment analysis, the news attributes comprise the sentiment score, probabilities of positive and negative sentiments, and the news item count. The climatic variables include the extreme and mean temperatures, meteorological conditions, AQI index, and wind velocity. Holiday periods are encoded as binary values. These data constitute a dataset of 29 distinct features. Although deep learning networks can automatically extract pertinent features, removing some irrelevant factors that are less related to tourist arrivals at first is necessary (Safi et al., 2020). This study uses the PCA to achieve this step. In terms of the pre-COVID-19 features, the first 10 principal components retain 85% of the original information, and adding any more principal components will only result in a 0.03 decrease in the growth rate (Table 2). Finally, the original 29 feature samples are downscaled into 10 dimensions. This study also chooses the 10 principal components for the post-COVID-19 data.
Results of the PCA
| Pre-COVID-19 | Post-COVID-19 | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| No | Cumulative variance contribution | Growth rate | No. | Cumulative variance contribution | Growth rate | No. | Cumulative variance contribution | Growth rate | No. | Cumulative variance contribution | Growth rate |
| 1 | 41.112 | 16 | 95.161 | 1.157 | 1 | 33.398 | 16 | 1.494 | 91.821 | ||
| 2 | 51.605 | 10.492 | 17 | 96.084 | 0.923 | 2 | 10.103 | 43.501 | 17 | 1.403 | 93.225 |
| 3 | 59.55 | 7.945 | 18 | 96.85 | 0.766 | 3 | 7.628 | 51.129 | 18 | 1.221 | 94.446 |
| 4 | 65.65 | 6.1 | 19 | 97.489 | 0.639 | 4 | 6.194 | 57.323 | 19 | 0.992 | 95.438 |
| 5 | 70.764 | 5.114 | 20 | 98.043 | 0.554 | 5 | 5.384 | 62.707 | 20 | 0.865 | 96.303 |
| 6 | 74.818 | 4.054 | 21 | 98.523 | 0.481 | 6 | 4.069 | 66.776 | 21 | 0.792 | 97.095 |
| 7 | 78.438 | 3.62 | 22 | 98.953 | 0.43 | 7 | 3.592 | 70.368 | 22 | 0.691 | 97.786 |
| 8 | 81.559 | 3.121 | 23 | 99.27 | 0.317 | 8 | 3.391 | 73.76 | 23 | 0.649 | 98.435 |
| 9 | 83.79 | 2.231 | 24 | 99.506 | 0.236 | 9 | 3.19 | 76.95 | 24 | 0.598 | 99.034 |
| 10 | 85.996 | 2.206 | 25 | 99.706 | 0.2 | 10 | 3.02 | 79.97 | 25 | 0.413 | 99.446 |
| 11 | 87.931 | 1.935 | 26 | 99.873 | 0.167 | 11 | 2.478 | 82.448 | 26 | 0.301 | 99.747 |
| 12 | 89.705 | 1.774 | 27 | 99.985 | 0.111 | 12 | 2.315 | 84.763 | 27 | 0.224 | 99.971 |
| 13 | 91.251 | 1.546 | 28 | 99.994 | 0.009 | 13 | 2.122 | 86.885 | 28 | 0.029 | 100 |
| 14 | 92.635 | 1.384 | 29 | 100 | 0.006 | 14 | 1.788 | 88.673 | 29 | 0 | 100 |
| 15 | 94.003 | 1.368 | 15 | 1.654 | 90.327 | ||||||
| Pre-COVID-19 | Post-COVID-19 | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| No | Cumulative variance contribution | Growth rate | No. | Cumulative variance contribution | Growth rate | No. | Cumulative variance contribution | Growth rate | No. | Cumulative variance contribution | Growth rate |
| 1 | 41.112 | 16 | 95.161 | 1.157 | 1 | 33.398 | 16 | 1.494 | 91.821 | ||
| 2 | 51.605 | 10.492 | 17 | 96.084 | 0.923 | 2 | 10.103 | 43.501 | 17 | 1.403 | 93.225 |
| 3 | 59.55 | 7.945 | 18 | 96.85 | 0.766 | 3 | 7.628 | 51.129 | 18 | 1.221 | 94.446 |
| 4 | 65.65 | 6.1 | 19 | 97.489 | 0.639 | 4 | 6.194 | 57.323 | 19 | 0.992 | 95.438 |
| 5 | 70.764 | 5.114 | 20 | 98.043 | 0.554 | 5 | 5.384 | 62.707 | 20 | 0.865 | 96.303 |
| 6 | 74.818 | 4.054 | 21 | 98.523 | 0.481 | 6 | 4.069 | 66.776 | 21 | 0.792 | 97.095 |
| 7 | 78.438 | 3.62 | 22 | 98.953 | 0.43 | 7 | 3.592 | 70.368 | 22 | 0.691 | 97.786 |
| 8 | 81.559 | 3.121 | 23 | 99.27 | 0.317 | 8 | 3.391 | 73.76 | 23 | 0.649 | 98.435 |
| 9 | 83.79 | 2.231 | 24 | 99.506 | 0.236 | 9 | 3.19 | 76.95 | 24 | 0.598 | 99.034 |
| 10 | 85.996 | 2.206 | 25 | 99.706 | 0.2 | 10 | 3.02 | 79.97 | 25 | 0.413 | 99.446 |
| 11 | 87.931 | 1.935 | 26 | 99.873 | 0.167 | 11 | 2.478 | 82.448 | 26 | 0.301 | 99.747 |
| 12 | 89.705 | 1.774 | 27 | 99.985 | 0.111 | 12 | 2.315 | 84.763 | 27 | 0.224 | 99.971 |
| 13 | 91.251 | 1.546 | 28 | 99.994 | 0.009 | 13 | 2.122 | 86.885 | 28 | 0.029 | 100 |
| 14 | 92.635 | 1.384 | 29 | 100 | 0.006 | 14 | 1.788 | 88.673 | 29 | 0 | 100 |
| 15 | 94.003 | 1.368 | 15 | 1.654 | 90.327 | ||||||
Note(s): The italic data is the number of selected principal components and corresponding variance contribution ratios and growth rate
Source(s): Authors’ own work
4.3 Performance measures
The proposed 2D multichannel imaging framework for integrating multi-source data is compared with traditional 2D imaging models designed for single time series data, as well as with other established benchmark models. Equations (8), (9), and (10) are used to evaluate the prediction performance, where and are the predicted and actual values.
4.4 Experimental results
This study utilizes a comprehensive set of benchmark models to ensure the robustness and generalizability of the findings. The first group includes the Autoregressive Integrated Moving Average (ARIMA) and Holt–Winters. These models use the single time series. The second group includes the OLS, KNN, and XG-BOOST, which use the multivariate time series. The auto_arima function is utilized to automatically ascertain the values for the ARIMA model’s p, q, and d parameters. The parameters for KNN, XG-BOOST, and LSTM are derived through grid search. The time series data are also simulated into a single-channel imaging model to evaluate the innovative aspect of this study, in line with previous studies.
Approximately 70% of the dataset is used for training, with the remaining 30% set aside for testing. A crucial parameter in AI models is the input sequence length (m), which represents the number of previous observations utilized for prediction. This study utilizes a grid search methodology to assess single-step and multi-step forecasting capabilities at various n values. Specifically, m is tested at values of 6, 9, 12, and 15, while the forecast horizons (h) is set to one step and three steps ahead. Table 3 outlines the optimal sequence length for each model, based on the respective forecast horizon.
Optimal n for each model with respect to different values of h
| Period | Pre-COVID-19 | Post-COVID-19 | |||
|---|---|---|---|---|---|
| Horizon | 1 | 3 | 1 | 3 | |
| Models | MGASF–CNN–LSTM | m = 15 | m = 15 | m = 15 | m = 15 |
| MGADF–CNN–LSTM | m = 12 | m = 15 | m = 15 | m = 15 | |
| MMTF–CNN–LSTM | m = 15 | m = 15 | m = 9 | m = 15 | |
| MRP–CNN–LSTM | m = 15 | m = 15 | m = 15 | m = 15 | |
| GASF–CNN–LSTM | m = 12 | m = 15 | m = 15 | m = 15 | |
| GADF–CNN–LSTM | m = 15 | m = 15 | m = 15 | m = 12 | |
| MTF–CNN–LSTM | m = 15 | m = 15 | m = 15 | m = 15 | |
| RP–CNN–LSTM | m = 15 | m = 15 | m = 15 | m = 15 | |
| ARIMA | m = 6 | m = 9 | m = 6 | m = 6 | |
| Holt–Winters | m = 6 | m = 6 | m = 6 | m = 12 | |
| OLS | m = 6 | m = 6 | m = 9 | m = 6 | |
| KNN | m = 15 | m = 15 | m = 15 | m = 15 | |
| XG-boost | m = 12 | m = 9 | m = 15 | m = 15 | |
| Period | Pre-COVID-19 | Post-COVID-19 | |||
|---|---|---|---|---|---|
| Horizon | 1 | 3 | 1 | 3 | |
| Models | MGASF–CNN–LSTM | m = 15 | m = 15 | m = 15 | m = 15 |
| MGADF–CNN–LSTM | m = 12 | m = 15 | m = 15 | m = 15 | |
| MMTF–CNN–LSTM | m = 15 | m = 15 | m = 9 | m = 15 | |
| MRP–CNN–LSTM | m = 15 | m = 15 | m = 15 | m = 15 | |
| GASF–CNN–LSTM | m = 12 | m = 15 | m = 15 | m = 15 | |
| GADF–CNN–LSTM | m = 15 | m = 15 | m = 15 | m = 12 | |
| MTF–CNN–LSTM | m = 15 | m = 15 | m = 15 | m = 15 | |
| RP–CNN–LSTM | m = 15 | m = 15 | m = 15 | m = 15 | |
| ARIMA | m = 6 | m = 9 | m = 6 | m = 6 | |
| Holt–Winters | m = 6 | m = 6 | m = 6 | m = 12 | |
| OLS | m = 6 | m = 6 | m = 9 | m = 6 | |
| KNN | m = 15 | m = 15 | m = 15 | m = 15 | |
| XG-boost | m = 12 | m = 9 | m = 15 | m = 15 | |
Source(s): Authors’ own work
The findings from various models utilizing pre- and post-COVID-19 data are detailed below, with analyses conducted for single-step-ahead forecasting (h = 1) and multi-step-ahead forecasting (h = 3).
- (1)
One-step-ahead forecasting
Table 4 shows the optimal m for each model in the single step ahead forecasting. The mean value of pre COVID-19 tourists is greater than those after COVID-19. However, the MAE, MAPE, and RMSE for models do not significantly differ. This finding suggests that all models are likely to fit the data pre-COVID-19 better. Moreover, some exciting results can be found.
Results of one-step-ahead forecasting for Jiuzhaigou
| Period | Pre-COVID-19 | Post-COVID-19 | ||||
|---|---|---|---|---|---|---|
| Measure | MAE | MAPE | RMSE | MAE | MAPE | RMSE |
| MGASF–CNN–LSTM | 569.605 | 0.059 | 2499.107 | 800.075 | 0.297 | 2723.166 |
| MGADF–CNN–LSTM | 572.295 | 0.059 | 2509.339 | 796.016 | 0.255 | 2556.829 |
| MMTF–CNN–LSTM | 569.212 | 0.057 | 2497.596 | 816.137 | 0.462 | 2804.063 |
| MRP–CNN–LSTM | 569.057 | 0.055 | 2495.883 | 772.264 | 0.203 | 2553.734 |
| GASF–CNN–LSTM | 5080.788 | 2.127 | 7458.790 | 5489.8581 | 2.727 | 7320.211 |
| GADF–CNN–LSTM | 4987.968 | 2.093 | 7457.710 | 5630.6738 | 2.910 | 7504.604 |
| MTF–CNN–LSTM | 5249.359 | 2.232 | 7504.604 | 5646.3492 | 3.025 | 7531.272 |
| RP–CNN–LSTM | 5167.184 | 2.195 | 7501.466 | 5599.9146 | 2.901 | 7457.710 |
| ARIMA | 6527.507 | 12.243 | 8208.974 | 6687.443 | 14.430 | 8465.265 |
| Holt–Winters | 6530.087 | 13.456 | 8236.313 | 6838.233 | 15.200 | 8550.719 |
| OLS | 1844.708 | 0.509 | 2869.429 | 2396.440 | 0.814 | 3543.890 |
| KNN | 1381.840 | 0.471 | 2816.723 | 1943.539 | 0.540 | 2922.243 |
| XG-BOOST | 1374.348 | 0.467 | 2812.333 | 1933.859 | 0.517 | 2871.385 |
| Period | Pre-COVID-19 | Post-COVID-19 | ||||
|---|---|---|---|---|---|---|
| Measure | MAE | MAPE | RMSE | MAE | MAPE | RMSE |
| MGASF–CNN–LSTM | 569.605 | 0.059 | 2499.107 | 800.075 | 0.297 | 2723.166 |
| MGADF–CNN–LSTM | 572.295 | 0.059 | 2509.339 | 796.016 | 0.255 | 2556.829 |
| MMTF–CNN–LSTM | 569.212 | 0.057 | 2497.596 | 816.137 | 0.462 | 2804.063 |
| MRP–CNN–LSTM | 569.057 | 0.055 | 2495.883 | 772.264 | 0.203 | 2553.734 |
| GASF–CNN–LSTM | 5080.788 | 2.127 | 7458.790 | 5489.8581 | 2.727 | 7320.211 |
| GADF–CNN–LSTM | 4987.968 | 2.093 | 7457.710 | 5630.6738 | 2.910 | 7504.604 |
| MTF–CNN–LSTM | 5249.359 | 2.232 | 7504.604 | 5646.3492 | 3.025 | 7531.272 |
| RP–CNN–LSTM | 5167.184 | 2.195 | 7501.466 | 5599.9146 | 2.901 | 7457.710 |
| ARIMA | 6527.507 | 12.243 | 8208.974 | 6687.443 | 14.430 | 8465.265 |
| Holt–Winters | 6530.087 | 13.456 | 8236.313 | 6838.233 | 15.200 | 8550.719 |
| OLS | 1844.708 | 0.509 | 2869.429 | 2396.440 | 0.814 | 3543.890 |
| KNN | 1381.840 | 0.471 | 2816.723 | 1943.539 | 0.540 | 2922.243 |
| XG-BOOST | 1374.348 | 0.467 | 2812.333 | 1933.859 | 0.517 | 2871.385 |
Source(s): Authors’ own work
First, Holt–Winters has the worst forecasting results among all the models, with MAEs of 6530.087 and 6838.233, MAPEs of 13.456 and 15.200, and RMSEs of 8236.313 and 8550.719. Meanwhile, MRP–CNN–LSTM has the smallest results with MAEs of 569.057 and 772.264, MAPEs of 0.055 and 0.203, and RMSEs of 2495.883 and 2553.734.
Second, when utilizing identical datasets, multi-channel imaging models outperform their counterparts by delivering superior results in multivariate time series analysis. Single-channel imaging models excel in forecasting univariate time series, underscoring the preeminence of time-series imaging coupled with deep learning predictive methodologies.
Moreover, OLS, KNN, and XG-boost results are considerably better than the time series models, even the single-channel imaging model. This notion means that multivariate data can better represent the nonlinear characteristics of tourism and be better fitted by the models compared with time series. 2D multi-channel imaging models have the best predictive performance. The performance differences between the multi-channel imaging technologies (GAF, MTF, and RP) are insignificant. Multi-channel image vision techniques can effectively improve prediction performance compared with single-channel image vision techniques and other models, regardless of the time interval.
- (2)
Multi-step-ahead forecasting
Table 5 has the same conclusions as the one-step-ahead forecasting. The multi-channel imaging models significantly outperform the baseline models in the multi-step ahead prediction. MGASF–CNN–LSTM and MRP–CNN–LSTM obtain the best results. The time-series models (ARIMA and Holt–Winters) always output the worst results. Time series imaging can effectively extract the data features whether using the multivariate or univariate time series.
Results of multi-step-ahead forecasting for Jiuzhaigou
| Period | Pre-COVID-19 | Post-COVID-19 | ||||
|---|---|---|---|---|---|---|
| Measure | MAE | MAPE | RMSE | MAE | MAPE | RMSE |
| MGASF–CNN–LSTM | 552.543 | 0.053 | 2119.660 | 809.310 | 0.341 | 2756.526 |
| MGADF–CNN–LSTM | 576.858 | 0.065 | 2516.128 | 799.318 | 0.257 | 2660.791 |
| MMTF–CNN–LSTM | 554.754 | 0.054 | 2229.129 | 811.480 | 0.461 | 2761.109 |
| MRP–CNN–LSTM | 567.366 | 0.054 | 2461.164 | 789.101 | 0.214 | 2555.037 |
| GASF–CNN–LSTM | 4977.555 | 1.981 | 7375.662 | 5450.734 | 2.718 | 7286.421 |
| GADF–CNN–LSTM | 4805.180 | 1.917 | 7286.421 | 5611.4971 | 2.909 | 7501.466 |
| MTF–CNN–LSTM | 4941.158 | 1.956 | 7320.211 | 5647.1357 | 3.084 | 7535.970 |
| RP–CNN–LSTM | 4965.838 | 1.958 | 7365.759 | 5529.1962 | 2.893 | 7375.662 |
| ARIMA | 6549.528 | 14.430 | 8325.730 | 6688.989 | 14.508 | 8480.441 |
| Holt–Winters | 6537.387 | 14.313 | 8305.170 | 6902.982 | 15.711 | 8599.571 |
| OLS | 2148.602 | 0.802 | 3459.231 | 2786.447 | 1.088 | 3776.166 |
| KNN | 1465.074 | 0.493 | 2832.017 | 1952.037 | 0.541 | 2954.167 |
| XG-BOOST | 1502.098 | 0.507 | 2861.728 | 2028.506 | 0.556 | 3268.129 |
| Period | Pre-COVID-19 | Post-COVID-19 | ||||
|---|---|---|---|---|---|---|
| Measure | MAE | MAPE | RMSE | MAE | MAPE | RMSE |
| MGASF–CNN–LSTM | 552.543 | 0.053 | 2119.660 | 809.310 | 0.341 | 2756.526 |
| MGADF–CNN–LSTM | 576.858 | 0.065 | 2516.128 | 799.318 | 0.257 | 2660.791 |
| MMTF–CNN–LSTM | 554.754 | 0.054 | 2229.129 | 811.480 | 0.461 | 2761.109 |
| MRP–CNN–LSTM | 567.366 | 0.054 | 2461.164 | 789.101 | 0.214 | 2555.037 |
| GASF–CNN–LSTM | 4977.555 | 1.981 | 7375.662 | 5450.734 | 2.718 | 7286.421 |
| GADF–CNN–LSTM | 4805.180 | 1.917 | 7286.421 | 5611.4971 | 2.909 | 7501.466 |
| MTF–CNN–LSTM | 4941.158 | 1.956 | 7320.211 | 5647.1357 | 3.084 | 7535.970 |
| RP–CNN–LSTM | 4965.838 | 1.958 | 7365.759 | 5529.1962 | 2.893 | 7375.662 |
| ARIMA | 6549.528 | 14.430 | 8325.730 | 6688.989 | 14.508 | 8480.441 |
| Holt–Winters | 6537.387 | 14.313 | 8305.170 | 6902.982 | 15.711 | 8599.571 |
| OLS | 2148.602 | 0.802 | 3459.231 | 2786.447 | 1.088 | 3776.166 |
| KNN | 1465.074 | 0.493 | 2832.017 | 1952.037 | 0.541 | 2954.167 |
| XG-BOOST | 1502.098 | 0.507 | 2861.728 | 2028.506 | 0.556 | 3268.129 |
Source(s): Authors’ own work
5. Robustness test
A dataset from Mount Siguniang, which encompasses various factors from the field of social psychology, is utilized to assess the robustness of the proposed model. This dataset includes the number of tourists, Baidu index, climate data, air quality information, holiday distribution, and news coverage, spanning from January 1, 2016 to December 31, 2021. The dataset is further segmented into pre-COVID-19 and post-COVID-19 periods. After the feature reduction, a concise summary of the data used for prediction is provided.
The best model is MRP–CNN–LSTM, as illustrated in Tables 6 and 7. The worst model is Holt–Winters. The findings are consistent with the empirical research. First, the forecasting performance of the proposed 2D multi-channel imaging models outperformed all baseline models, substantiating the hypothesis that the transformation of 1D data into 2D images can substantially enhance predictive outcomes. Second, when utilizing identical datasets, the multi-channel imaging models demonstrate superiority over OLS, KNN, and XG-BOOST algorithms. Meanwhile, single-channel imaging models yield less robust results than ARIMA and Holt–Winters method. Finally, the incorporation of multidimensional data markedly augments forecast precision, as evidenced by the superior performance of models (OLS, KNN, and XG-BOOST) that leverage multiple data sources over those (single-channel imaging technique, ARIMA, and Holt–Winters) reliant on time series data alone.
Results of one-step-ahead forecasting for Mount Siguniang
| Period | Pre-COVID-19 | Post-COVID-19 | ||||
|---|---|---|---|---|---|---|
| Measure | MAE | MAPE | RMSE | MAE | MAPE | RMSE |
| MGASF–CNN–LSTM | 210.976 | 0.242 | 816.533 | 337.048 | 1.323 | 1162.448 |
| MGADF–CNN–LSTM | 213.156 | 0.268 | 898.192 | 331.583 | 0.886 | 911.618 |
| MMTF–CNN–LSTM | 212.489 | 0.250 | 880.895 | 334.897 | 1.291 | 1056.314 |
| MRP–CNN–LSTM | 210.006 | 0.226 | 765.240 | 327.739 | 0.617 | 903.387 |
| GASF–CNN–LSTM | 1119.225 | 1.819 | 1953.703 | 1335.785 | 3.320 | 2722.238 |
| GADF–CNN–LSTM | 1247.142 | 2.054 | 2195.241 | 1632.381 | 3.927 | 3340.243 |
| MTF–CNN–LSTM | 1221.156 | 2.014 | 2014.321 | 1326.100 | 2.433 | 2357.950 |
| RP–CNN–LSTM | 1239.115 | 2.028 | 2172.926 | 1595.881 | 3.838 | 3264.474 |
| ARIMA | 2274.267 | 3.941 | 3250.282 | 2642.256 | 9.400 | 4374.119 |
| Holt–Winters | 2287.577 | 4.020 | 3254.243 | 2653.593 | 9.610 | 4410.260 |
| OLS | 622.386 | 1.618 | 1514.051 | 838.636 | 1.787 | 1527.959 |
| KNN | 461.517 | 1.333 | 1188.725 | 522.277 | 1.355 | 1494.061 |
| XG-BOOST | 478.059 | 1.339 | 1305.557 | 535.810 | 1.418 | 1504.567 |
| Period | Pre-COVID-19 | Post-COVID-19 | ||||
|---|---|---|---|---|---|---|
| Measure | MAE | MAPE | RMSE | MAE | MAPE | RMSE |
| MGASF–CNN–LSTM | 210.976 | 0.242 | 816.533 | 337.048 | 1.323 | 1162.448 |
| MGADF–CNN–LSTM | 213.156 | 0.268 | 898.192 | 331.583 | 0.886 | 911.618 |
| MMTF–CNN–LSTM | 212.489 | 0.250 | 880.895 | 334.897 | 1.291 | 1056.314 |
| MRP–CNN–LSTM | 210.006 | 0.226 | 765.240 | 327.739 | 0.617 | 903.387 |
| GASF–CNN–LSTM | 1119.225 | 1.819 | 1953.703 | 1335.785 | 3.320 | 2722.238 |
| GADF–CNN–LSTM | 1247.142 | 2.054 | 2195.241 | 1632.381 | 3.927 | 3340.243 |
| MTF–CNN–LSTM | 1221.156 | 2.014 | 2014.321 | 1326.100 | 2.433 | 2357.950 |
| RP–CNN–LSTM | 1239.115 | 2.028 | 2172.926 | 1595.881 | 3.838 | 3264.474 |
| ARIMA | 2274.267 | 3.941 | 3250.282 | 2642.256 | 9.400 | 4374.119 |
| Holt–Winters | 2287.577 | 4.020 | 3254.243 | 2653.593 | 9.610 | 4410.260 |
| OLS | 622.386 | 1.618 | 1514.051 | 838.636 | 1.787 | 1527.959 |
| KNN | 461.517 | 1.333 | 1188.725 | 522.277 | 1.355 | 1494.061 |
| XG-BOOST | 478.059 | 1.339 | 1305.557 | 535.810 | 1.418 | 1504.567 |
Source(s): Authors’ own work
Results of multi-step-ahead forecasting for Mount Siguniang
| Period | Pre-COVID-19 | Post-COVID-19 | ||||
|---|---|---|---|---|---|---|
| Measure | MAE | MAPE | RMSE | MAE | MAPE | RMSE |
| MGASF–CNN–LSTM | 212.742 | 0.252 | 888.094 | 332.737 | 1.039 | 916.476 |
| MGADF–CNN–LSTM | 212.915 | 0.258 | 888.652 | 333.725 | 1.222 | 944.808 |
| MMTF–CNN–LSTM | 214.062 | 0.278 | 902.145 | 333.772 | 1.227 | 949.239 |
| MRP–CNN–LSTM | 211.165 | 0.242 | 835.993 | 330.785 | 0.789 | 904.919 |
| GASF–CNN–LSTM | 1125.362 | 1.830 | 1956.290 | 1341.983 | 3.820 | 2735.917 |
| GADF–CNN–LSTM | 1273.273 | 3.384 | 2221.072 | 1654.494 | 4.963 | 3421.184 |
| MTF–CNN–LSTM | 1219.962 | 1.872 | 2007.890 | 1331.706 | 2.671 | 2368.031 |
| RP–CNN–LSTM | 1253.516 | 2.331 | 2205.644 | 1686.151 | 5.130 | 3436.304 |
| ARIMA | 2252.380 | 3.880 | 3203.655 | 2682.821 | 10.024 | 4423.429 |
| Holt–Winters | 2289.692 | 4.045 | 3293.116 | 2689.276 | 10.847 | 4436.791 |
| OLS | 699.960 | 1.667 | 1515.454 | 980.889 | 1.813 | 1838.232 |
| KNN | 561.789 | 1.492 | 1513.636 | 693.681 | 1.667 | 1514.997 |
| XG-BOOST | 512.356 | 1.350 | 1432.087 | 857.748 | 1.800 | 1539.139 |
| Period | Pre-COVID-19 | Post-COVID-19 | ||||
|---|---|---|---|---|---|---|
| Measure | MAE | MAPE | RMSE | MAE | MAPE | RMSE |
| MGASF–CNN–LSTM | 212.742 | 0.252 | 888.094 | 332.737 | 1.039 | 916.476 |
| MGADF–CNN–LSTM | 212.915 | 0.258 | 888.652 | 333.725 | 1.222 | 944.808 |
| MMTF–CNN–LSTM | 214.062 | 0.278 | 902.145 | 333.772 | 1.227 | 949.239 |
| MRP–CNN–LSTM | 211.165 | 0.242 | 835.993 | 330.785 | 0.789 | 904.919 |
| GASF–CNN–LSTM | 1125.362 | 1.830 | 1956.290 | 1341.983 | 3.820 | 2735.917 |
| GADF–CNN–LSTM | 1273.273 | 3.384 | 2221.072 | 1654.494 | 4.963 | 3421.184 |
| MTF–CNN–LSTM | 1219.962 | 1.872 | 2007.890 | 1331.706 | 2.671 | 2368.031 |
| RP–CNN–LSTM | 1253.516 | 2.331 | 2205.644 | 1686.151 | 5.130 | 3436.304 |
| ARIMA | 2252.380 | 3.880 | 3203.655 | 2682.821 | 10.024 | 4423.429 |
| Holt–Winters | 2289.692 | 4.045 | 3293.116 | 2689.276 | 10.847 | 4436.791 |
| OLS | 699.960 | 1.667 | 1515.454 | 980.889 | 1.813 | 1838.232 |
| KNN | 561.789 | 1.492 | 1513.636 | 693.681 | 1.667 | 1514.997 |
| XG-BOOST | 512.356 | 1.350 | 1432.087 | 857.748 | 1.800 | 1539.139 |
Source(s): Authors’ own work
6. Discussion and conclusion
This study addresses overfitting in deep learning and aims to advance the application of computer vision techniques in tourism forecasting. Accordingly, this study introduces a 2D multi-channel imaging model derived from multivariate time series data to achieve this goal. The proposed framework consists of five key components: dimensionality reduction, sequence image generation, image stitching, feature extraction, and model training. An empirical analysis and robustness evaluation are conducted using a multimodal dataset that includes specific metrics, such as tourist volume, climate conditions, holidays, search engine data, and news coverage for the Jiuzhaigou and Mount Siguniang regions.
Three main exciting findings emerged from this study. First, the time-series imaging models have superior prediction results when using the same data. Transforming 1D time series data into 2D representations allows the proposed models to capture more effectively underlying patterns, leading to improved predictive performance. 2D images have richer information transmission and connections than traditional 1D time series data. The degree of relationship between data point in 1D data decreases as the number of samples increases, resulting in a loss of valid information and an increase in noise. This problem can be solved by the 2D images. Each sample in an image can have an infinite number of neighboring data points, which breaks the limitation of the number of neighboring points and improves the correlation between the data. Thus, the multi-channel imaging model can accurately capture the subtle changes and trends and forecast the tourism demand.
Second, this study emphasizes the advantages of using the multiple time series in forecasting tasks. Multiple data can encompass a broader range of data types, detect highly complex nonlinear patterns, reduce the risk of overfitting, and improve forecasting performance. The proposed multi-channel time series imaging model can efficiently incorporate multivariate time series, integrate information from different data sources to analyze tourism demand, capture the influence of different factors, and improve forecasting accuracy compared with the single-channel imaging model.
Previous research has conclusively demonstrated the effects of the COVID-19 pandemic on tourism demand and consumer behavior. When data from the post-pandemic period are included, the correlations between the selected variables and tourism metrics weakened. Consequently, the model’s overall fit to the data diminished, and the improvement in prediction accuracy with the multi-channel imaging model is not as significant as it was prior to the pandemic.
The two main contributions of this study are as follows. First, this study integrates a multitude of features derived from a socio-psychological framework, including climatic data, media content, Baidu indices, and holiday patterns. This study delves into the influence of non-economic factors on tourism demand forecasting, thus broadening the analytical lens to encompass a more holistic understanding from a social scientific standpoint. The social sciences are fundamentally concerned with the intricate dynamics among various factors to decode human behavior. Therefore, the inclusion of a diverse array of data sources offers a sophisticated perspective on the multifaceted influences on tourism, culminating in enhanced precision in demand prediction.
Second, a multi-channel imaging model is proposed for the first time, extending the application of computer vision techniques. The model allows the inclusion of multivariate time series data. Dimensionality reduction techniques are used to maximize the retention of valid information from multiple data sources. The serial data imaging technique can effectively capture the hierarchical structure between data. Moreover, the convolution and pooling of CNNs reflect excellent ability in extracting image features. Previous research on the prediction and application of time imaging technology mainly focused on the single-channel imaging of 1D time series. The model proposed in this study is a multi-channel with a stitching process for images. This model can effectively solve the problem of over-fitting in deep learning and provide the basis for researchers interested in applying deep learning techniques for image processing performance in tourism demand prediction.
The findings of this research offer practical insights for managers overseeing tourism destinations and for governmental bodies involved in their administration. First, a model based on multi-channel imaging technology can provide tourism practitioners highly accurate and timely forecasts of tourism demand. In addition, relevant authorities can use multivariate data that do not rely on historical visitor numbers (e.g. weather data can be directly obtained from weather forecasts) to determine trends in tourism demand in advance.
This study has some limitations. As the research focuses on tourism forecasting, the dataset was constructed based on the Social Psychological Framework for Tourism Demand Generation and included specific factors. Future studies could expand the scope by incorporating more determinants, especially from a socio-psychological perspective. Additionally, applying the model to other fields, such as the hotel industry, could enhance its generalizability. Lastly, while the model performs well in forecasting, its post-COVID performance decline highlights its limitations in handling external disruptions. Future research could refine the model to improve its adaptability.
This project was partly supported by research grants funded by the University of Macau and Hainan University.
Declaration of interest statement:The author(s) declared no potential conflicts of interest with respect to the research, authorship and/or publication of this article.







