The purpose of this study is to make property price forecasts for the Chinese housing market that has grown rapidly in the last 10 years, which is an important concern for both government and investors.
This study examines Gaussian process regressions with different kernels and basis functions for monthly pre-owned housing price index estimates for ten major Chinese cities from March 2012 to May 2020. The authors do this by using Bayesian optimizations and cross-validation.
The ten price indices from June 2019 to May 2020 are accurately predicted out-of-sample by the established models, which have relative root mean square errors ranging from 0.0458% to 0.3035% and correlation coefficients ranging from 93.9160% to 99.9653%.
The results might be applied separately or in conjunction with other forecasts to develop hypotheses regarding the patterns in the pre-owned residential real estate price index and conduct further policy research.
1. Introduction
The Chinese real estate market has grown significantly during the past 10 years. Real estate price prediction problems have surely become one of the biggest issues facing governments and investors (Xu and Zhang, 2022d; Xu and Zhang, 2024a; Tekouabou et al., 2023; Xu and Zhang, 2022a; Lahmiri et al., 2023; Xu and Zhang, 2023m). It is critical to comprehend real estate price trends and changes because they have a direct impact on people’s decisions about where to live, how to invest in real estate and how regulatory authorities formulate and implement their regulations (Xu and Zhang, 2023a; Hjort et al., 2022; Xu and Zhang, 2023b; Yang et al., 2023; Xu and Zhang, 2022c). Predicting real estate prices is of course of interest to many customers and suppliers of estimates.
Predicting financial and economic time-series data with high accuracy and dependability is of interest to a large number of scholars and practitioners. Many of their expansions and modifications as well as certain important time series approaches, such the autoregressive (AR), vector autoregressive (VAR) and vector error correction (VEC) models, have been studied for a range of forecasting applications (Xu and Zhang, 2023c; Xu and Zhang, 2023d; Kim et al., 2007; Xu and Zhang, 2023e; Xu and Zhang, 2023f; Cabrera et al., 2011; Xu, 2020; Xu, 2019a; Webb et al., 2016; Xu, 2017a; Xu, 2017b; Wei and Cao, 2017; Xu, 2018a; Xu, 2017; Yang et al., 2018; Xu, 2018b; Xu and Zhang, 2021a; Liu and Wu, 2020; Xu and Zhang, 2024b; Xu and Zhang, 2023g; Milunovich, 2020; Xu and Zhang, 2022b; Xu and Zhang, 2021b; Glennon et al., 2018; Xu, 2019b; Xu, 2018c; Guo, 2020; Xu, 2018d; Xu, 2018e; Mei and Fang, 2017; Xu, 2015; Xu, 2019; Hepşen and Vatansever, 2011; Xu and Zhang, 2023h; Baroni et al., 2005). Recently, it has been discovered that many machine learning techniques, such as neural networks, boosting, bagging, regression trees (RT), random forests, nearest neighbors, deep learning, ensemble learning and Gaussian process regressions, are promising and effective solutions to a range of real estate price forecasting problems. The neural network model seems to be one of the most popular ways (Xu and Zhang, 2024c; Xu and Zhang, 2021c; Xu and Zhang, 2023i; Xu and Zhang, 2023j; Xu and Zhang, 2023k) for predicting real estate prices (Liu and Wu, 2020; Milunovich, 2020; Lam et al., 2008; Yasnitsky et al., 2021; Xu and Zhang, 2023 1), and these evaluations, while not comprehensive, are generally consistent with a range of empirical studies on the adoption of machine learning techniques for forecasting in finance and economics (Xu and Zhang, 2022f; Xu and Zhang, 2023n; Yang et al., 2008; Xu and Zhang, 2024d; Xu and Zhang, 2022g; Wang and Yang, 2010; Xu and Zhang, 2021d; Xu and Zhang, 2023o; Yang et al., 2010; Xu and Zhang, 2022h; Xu and Zhang, 2023p; Wegener et al., 2016; Xu and Zhang, 2023q; Xu and Zhang, 2022i; Xu and Zhang, 2023r; Xu and Zhang, 2022e). On the other hand, the predictions of time-series data from the real estate price index using the Gaussian process regression are not well studied.
A novel regression method is based on Neal’s work on Bayesian learning for neural networks (Neal, 2012). Modeling noisy data is appropriate because the method relies on priors over functions of Gaussian processes. It was demonstrated that a wide variety of neural network-based Bayesian regression models will converge to Gaussian processes at the limit of an infinite network (Neal, 2012). Regressions using the Gaussian process have demonstrated efficacy in simulating both noisy (Jin and Xu, 2024a; Jin and Xu, 2024b) and noise-free (Neal, 1997) data. Brahim-Belhouari and Vesin (2001) examined Gaussian processes with radial basis function neural networks for forecasting problems requiring stationary time-series data and discovered that Bayesian learning yields superior prediction results. According to research by Brahim-Belhouari and Bermak (2004), it is beneficial to look at different covariance functions, which is the method used in this study. In addition, prediction techniques based on Gaussian processes can be used to effectively solve forecasting problems for nonstationary time-series data. Brahim-Belhouari and Bermak’s (2004) research shows that Gaussian process regressions outperform radial basis function neural networks. Furthermore, the Gaussian process formulation’s value and benefit come from the exact matrix operations required to integrate the prior and noise models (Brahim-Belhouari and Bermak, 2004). Brahim-Belhouari and Bermak (2004) also recommended using the technique of multimodel forecasting using Gaussian process predictors, which is comparable to how we would use model averaging. Based on a number of recent research (Xu and Zhang, 2023s; Jin and Xu, 2024c), price time series for the resource and agricultural industries may be successfully projected using Gaussian process regressions.
Forecasting efforts for residential real estate prices have also been noticed using traditional econometric models and, more recently, machine learning approaches. In comparing parametric and semiparametric classical econometric models, for instance, Gençay and Yang (1996) and Gencay and Yang (1996) found that semiparametric models are better suited for predicting and assessing residential housing prices. Glennon et al. (2018) find that the accuracy of property appraisals based on house price indices is increased when many models are used. By devising a system to mitigate the effects of measurement errors associated with the repeat sales methodology, Clapp and Giaccotto (1992) demonstrate the superior efficacy of the assessed value strategy. Forecasts produced using formulas based on disaggregated data for the entire city have an edge over forecasts created using local average pricing formulas, claim Kaboudan and Sarkar (2007). Mei and Fang (2017) use trend analysis and multiple regression to build a dynamic state forecasting model of the average selling price. Levesque (1994) breaks down residential property prices using a case study of airport noise. Hepşen and Vatansever (2011) analyze the AR integrated moving average model’s performance. Baroni et al. (2005) use principal components analysis to generate an indicator of repeat sales that forecasts apartment prices. Guo (2020) uses regression models, both linear and nonlinear, to study future price stability. Paris (2008) looks into artificial neural networks for machine learning techniques in order to forecast changes in local and national price indices for the UK residential real estate market. Chi (2017) proposes estimating residential real estate prices via a spatial back-propagation neural network. Bee-Hua (2000) suggests merging neural networks and evolutionary algorithms to assess demand for residential building. When it comes to calculating the values of residential real estate, Štubňová et al. (2020) find that neural networks perform better than regression models. Seya and Shiroi (2022) examine the closest neighbor Gaussian process and deep neural network for residential unit rent price prediction and find that the former has more promise. Rafiei and Adeli (2018) recommend using an unsupervised deep Boltzmann machine (DBM) learning approach, a softmax layer to extract relevant features from the input data and a three-layer back-propagation neural network (or support vector machine) to convert the trained unsupervised DBM into a supervised regression network to estimate construction cost from economic variables and indices. Zhang et al. (2021) found that the radial basis function-based support vector regression (SVR) algorithm and the extra-trees regression approach are both useful for simulating the fine-scale spatiotemporal distribution of residential land values. Yoo et al. (2012) use cubist, random forest and traditional ordinary least squares to simulate the hedonic pricing of residential real estate sales. They find that the random forest model yields the best accurate results. Researchers Dimopoulos and Bakas (2019), Hong et al. (2020) and Dimopoulos et al. (2018) demonstrate how machine learning models may be applied to improve the precision of real estate mass evaluations. Ai et al. (2020) suggest that machine learning algorithms are useful for residential land appraisals as well. Picchetti (2017) shows how sample variation in hedonic geospatial residential property price evaluations may be overcome by applying the gradient tree boosting technique.
1.1 Research contributions and limitations
To keep with this topic, we concentrate on Gaussian process regressions for estimates of pre-owned residential real estate price indices for ten major Chinese cities from March 2012 to May 2020, a time when the real estate market expanded rapidly. Results here contribute to the literature in the following aspects. First, to the best of our knowledge, this is the first forecast research that looks at pre-owned residential real estate price indices in the Chinese market using Gaussian process regressions. In earlier studies, the Boston housing-related problems were often examined from the perspective of property valuation using the Gaussian process regression technique (Harrison and Rubinfeld, 1978; Rasmussen, 1997; Williams, 1997; Chu et al., 2004; Kim and Ghahramani, 2006; Titsias, 2009; Tokdar et al., 2010; Jylänki et al., 2011; Nielsen et al., 2012; Friedman and Nachman, 2013; Muñoz-González et al., 2014; Dearmon and Smith, 2016; Ranjan et al., 2016; Hartmann and Vanhatalo, 2019; Li et al., 2021; Miao et al., 2021; Murakami, 2021; Miao et al., 2022; Sebenius et al., 2022; Khosravi et al., 2022; Liu et al., 2022; Algikar and Mili, 2023). Second, it might not take much incentive given how well-known the Chinese residential real estate sector is. For investors and policymakers, forecasting residential real estate price indices should be both essential and possibly challenging, because having a solid understanding of pricing patterns may aid in making decisions. Numerous forecasting techniques exist, such as machine learning and econometric models. The nonlinear patterns displayed by the residential real estate price indices under examination, along with the literature’s acknowledged worth and potential for real estate price forecasting, led to the selection of the Gaussian process regression approach, which could effectively model nonlinear characteristics embedding the underlying price indices. Third, a variety of basis functions and kernels will be used in a Bayesian optimization method together with the cross-validation technique to train the Gaussian process regression. Particularly, Bayesian optimizations have great potential to improve the forecast models’ generalization performance and flexibility. Fourth, we believe that our results will contribute to a better understanding of using machine learning technology to anticipate residential real estate price in China, because the majority of previous studies have evaluated other forms of real estate assets based on a single location (Lam et al., 2008; Ma et al., 2015; Li et al., 2018; Wang et al., 2014; Ho et al., 2021). Because supply and demand are most active in these major cities and each may have unique price characteristics that merit examination, the current work’s coverage of these cities along with the data’s accessibility should represent an economically natural way to investigate the residential real estate forecast problem. Fifth, our findings may offer a more modern viewpoint on the effectiveness of machine learning techniques for real estate price index estimates for the Chinese market, as previous research in this field has generally focused on earlier time periods when researching other forms of real estate, such as Q12000–Q32010 (Wang et al., 2014), 7M2013–12M2013 (Ma et al., 2015), 6M1996–8M2014 (Ho et al., 2021), 12M2010–10M2017 (Liu and Liu, 2019) and 1M2005–11M2018 (Li et al., 2020). The level of uncertainty that currently exists in the residential real estate market might lead to more nonlinear pricing behavior (Xu and Zhang, 2024e; Xu and Zhang, 2022j). By developing Gaussian process regression models based on a more current time period, from March 2012 to May 2020, this work contributes to the body of literature on the applicability of the Gaussian process regression for real estate price index projections in dynamic situations. We also apply the developed models on pre-owned housing prices across the ten cities investigated during a more recent time period of February 2023–January 2024, and find at the models could still generate rather accurate forecast performance for November 2023, December 2023 and January 2024. Sixth, this study offers timely forecasting tools that policymakers and investors may find useful. Machine learning techniques could seem more complicated than econometric models to many forecast consumers who are not as experienced. Therefore, we create very simple yet accurate Gaussian process regressions to support technological predictions. Since econometric and machine learning models may overfit or underfit, we have considered a trade-off between stabilities and prediction accuracy in our model building processes. In particular, we perform out-of-sample forecasts from June 2019 to May 2020 and get relative root mean square errors (RRMSE) ranging from 0.0458% to 0.3035% and correlation coefficients ranging from 93.9160% to 99.9653% for the ten price indices. We also conduct benchmark analysis of the constructed models here against different traditional time-series econometric models and alternative machine learning models, and find that the Gaussian process regression models built here lead to statistically significant better performance in terms of forecast accuracy. Our results might be applied separately or in conjunction with other projections to develop hypotheses regarding the patterns in the residential real estate price index and conduct further policy research. It should be meaningful to note that the current technique adopted here relies on historical data as the source of prediction information rather than considering other economic aspects as additional potential predictors. As a consequence, one may integrate the prediction results from other sources with those generated by the models built here to overcome this potential limitation. In this situation, it could be necessary to build suitable model combinations.
2. Literature review
2.1 Real estate price forecasting using econometric methods
While machine learning approaches for residential real estate price index prediction are the focus of this study, econometric methodologies for broad real estate price forecasts have also been thoroughly investigated in the literature. Rental prices for the UK market are predicted using a variety of techniques, such as moving average, AR, integrated AR-moving average, simple least squares and seemingly unrelated regression equations (Silver and Goode, 1990; McGough and Tsolacos, 1995). In addition, the VAR regression approach indicates that the horizon affects how accurately retail rental prices are predicted for the UK market (Brooks and Tsolacos, 2000). Another study based on several linear least-squares regressions suggests that retail leasing costs in many British towns and cities may be predicted based on the size of the retail core and the demographics of the local population (Jackson, 2001). Time series analysis of retail real estate price returns has also been done. A generalized AR conditional heteroscedasticity in mean model indicates that macroeconomic factors are better predictors of price returns for office real estate in Australia than for retail or industrial real estate (West and Worthington, 2006). The VEC model suggests that retail sales and mortgage loans might be used to predict the price of Greek retail real estate. Numerous studies indicate that semiparametric models outperform parametric models in predicting residential real estate values (Gençay and Yang, 1996; Gencay and Yang, 1996). In addition, specific model combinations (Glennon et al., 2018), such as combining nonlinear and linear models (Guo, 2020), may improve forecasting accuracy. It is also a good idea to combine trend analysis with different linear regression techniques to build dynamic models that predict average selling prices (Mei and Fang, 2017). Principal component analysis has been demonstrated to enhance residential real estate price projections (Baroni et al., 2005); however, the AR integrated moving average approach also shows promise (Hepşen and Vatansever, 2011). In general, it is discovered that the VAR method and the directed acyclic graph technique are helpful for forecasting and studying the dynamics of Real Estate Investment Trusts in the USA (Kim et al., 2007). Another study emphasizes how crucial it is to combine machine learning approaches, like neural networks and other nonparametric regression techniques, with econometric approaches, like the AR approach, the functional coefficient approach, and the exponential generalized AR conditional heteroscedasticity approach, to forecast returns of global securitized real estate (Cabrera et al., 2011). Research has shown that machine learning models – most notably the neural network methodology – must be used in addition to econometric approaches for predicting real estate values (Kim et al., 2007; Cabrera et al., 2011; Liu and Wu, 2020). Several iterations of basic time-series econometric approaches are investigated in the context of real estate value forecasting research. For example, a generalized VAR technique that incorporates high-dimensional data processing approaches is used to analyze spillovers across residential home prices among 69 Chinese cities (Yang et al., 2018); a jump generalized AR conditional heteroscedasticity model is used (Webb et al., 2016); and the VEC technique incorporating the smooth transition weighting approach is used to improve the predictability of national housing price indices in the USA. Research has also suggested using a range of linear models to lessen the influence of nonlinearities on predictions of real estate prices (Webb et al., 2016). A more direct approach may be to take into account nonlinear models, which include machine learning methods.
2.2 Real estate price forecasting using machine learning techniques
Geographically, the USA (Peterson and Flanagan, 2009; Plakandaras et al., 2015), Australia (Milunovich, 2020), Europe (Ćetković et al., 2018), Russia (Yasnitsky et al., 2021), Nigeria (Abidoye and Chan, 2017; Abidoye and Chan, 2018), Uganda (Embaye et al., 2021), Malaysia, China (Wei and Cao, 2017; Liu and Wu, 2020; Lam et al., 2008; Ma et al., 2015; Li et al., 2018; Wang et al., 2014; Ho et al., 2021; Xu and Zhang, 2023t; Xu and Li, 2021), Italy (Morano et al., 2015), Spain (Rico-Juan and de La Paz, 2021), Malawi (Embaye et al., 2021), Tanzania (Embaye et al., 2021), South Korea (Kang et al., 2020) and Iran (Azadeh et al., 2014; Rafiei and Adeli, 2016) are the countries and/or regions where machine learning techniques for real estate price projections have been applied. Given their widespread application, machine learning techniques hold promise for forecasting real estate prices using time-series data with a variety of underlying variables (Xu and Zhang, 2023u; Xu and Zhang, 2023v; Xu and Zhang, 2023w). The Chinese market is the main subject of our present investigation. Most of the material that we have looked at so far is focused on a few particular areas. For the USA, Fairfax County of Virginia, California, Boston and Ames, and Wake County of North Carolina (Peterson and Flanagan, 2009) are examined. For Malaysia, Petaling of Kuala Lumpur and Mukim Pulai of Johor Bahru are investigated. For China, Hong Kong (Lam et al., 2008; Li et al., 2018; Ho et al., 2021), Tangshan (Gu et al., 2011), Beijing (Li et al., 2020; Yan and Zong, 2020), Chongqing (Wang et al., 2014), Dalian, Shenzhen (Liu and Liu, 2019), Shanghai (Ma et al., 2015), Xuzhou, Kunming, Handan, and Changchun (Liu and Wu, 2020) and Shenzhen, Beijing, Guangzhou, and Shanghai (Xu and Li, 2021) are explored. For Nigeria, Benin and Lagos (Abidoye and Chan, 2017; Abidoye and Chan, 2018) are assessed. For South Korea, Seoul is studied (Kang et al., 2020). For Iran, Tehran is researched (Rafiei and Adeli, 2016). For Italy, Bari (Morano et al., 2015), Taranto and Milan are taken into consideration. For Russia, Moscow is analyzed (Yasnitsky et al., 2021). For Spain, Alicante is the focus (Rico-Juan and de La Paz, 2021). To provide a more thorough understanding of real estate price forecasts using machine learning techniques and to enrich the relevant literature with expanded empirical evidence on the usefulness of such methods in this research field, which could benefit both investors and policymakers in examining their applications to a wider selection of real estate markets, the current study takes into consideration the ten major cities in China, which represent the largest real estate markets in China [1].
Prior research has used modeling and forecasting for a range of purposes in developing machine learning systems, depending on their specialized fields. One of the areas of concentration may be developing machine learning algorithms based on various real estate attributes for asset appraisals (Peterson and Flanagan, 2009; Morano et al., 2015; Abidoye and Chan, 2017; Abidoye and Chan, 2018; Kang et al., 2020; Ho et al., 2021; Xu and Li, 2021; Rafiei and Adeli, 2016), from many macroeconomic variables when valuing assets (Kang et al., 2020; Rafiei and Adeli, 2016), from time series of lagging real estate prices for technical forecasts (Liu and Wu, 2020; Ma et al., 2015; Wang et al., 2014; Li et al., 2020; Gu et al., 2011), from various real estate attributes for technical forecasts (Lam et al., 2008; Li et al., 2018; Yasnitsky et al., 2021; Liu and Liu, 2019; Yan and Zong, 2020; Embaye et al., 2021; Rico-Juan and de La Paz, 2021; Chen et al., 2017) and using several macroeconomic variables for technical forecasts (Wei and Cao, 2017; Milunovich, 2020; Lam et al., 2008; Azadeh et al., 2014; Ćetković et al., 2018; Li et al., 2018; Kang et al., 2020; Yasnitsky et al., 2021; Liu and Liu, 2019; Plakandaras et al., 2015; Rico-Juan and de La Paz, 2021). Aside from the usual macroeconomic factors such as gross domestic product, interest rates and unemployment rates, other popular real estate characteristics taken into account include vintages, lot sizes and locations. Our work is one of several that creates machine learning models for technical forecasts using delayed time-series data on real estate prices (Liu and Wu, 2020; Ma et al., 2015; Wang et al., 2014; Li et al., 2020; Gu et al., 2011). To develop the models for this study, a number of kernels and basis functions are investigated through the use of Bayesian optimization and cross-validation. In particular, Gaussian process regressions are utilized.
Previous research has focused on single machine learning models (Lam et al., 2008; Azadeh et al., 2014; Ma et al., 2015; Morano et al., 2015; Abidoye and Chan, 2017; Ćetković et al., 2018; Li et al., 2018; Kang et al., 2020; Yasnitsky et al., 2021; Wang et al., 2014; Li et al., 2020; Gu et al., 2011; Rafiei and Adeli, 2016; Chen et al., 2017), comparisons between different models (Li et al., 2017; Pai and Wang, 2020; Liu and Liu, 2019; Yan and Zong, 2020; Ho et al., 2021; Embaye et al., 2021; Xu and Li, 2021; Rico-Juan and de La Paz, 2021), machines learning models and conventional econometric models (Liu and Wu, 2020; Milunovich, 2020; Peterson and Flanagan, 2009; Li et al., 2017; Abidoye and Chan, 2018; Plakandaras et al., 2015) and combinations of models (Wei and Cao, 2017). More precisely, our findings indicate that the neural network approach is the most popular option for studies that concentrate on a single machine learning method (Lam et al., 2008; Azadeh et al., 2014; Ma et al., 2015; Morano et al., 2015; Abidoye and Chan, 2017; Ćetković et al., 2018; Li et al., 2018; Kang et al., 2020; Yasnitsky et al., 2021; Rafiei and Adeli, 2016), with SVRs (Wang et al., 2014; Gu et al., 2011; Chen et al., 2017) and ensemble learners (Li et al., 2020) coming in second and third, respectively. SVR models outperform back-propagation neural networks in terms of accuracy, according to a study evaluating many machine learning techniques. In another study, SVRs based on integrated hybrid genetic techniques achieve higher accuracy than fuzzy and back-propagation neural networks. However, it is discovered that gradient boosting techniques and random forest models both perform better than SVRs (Ho et al., 2021). The best results are produced by the RIPPER method (repeated incremental pruning to produce error reduction), which is compared to naive Bayesian algorithms, C4.5 and AdaBoost. A comparative study of fuzzy inference systems, neural networks and fuzzy least-squares demonstrates the potential advantages of fuzzy-based techniques. The effectiveness of different data processing techniques is demonstrated by a research that compares neural networks with single/double exponential smoothing, group data processing and back-propagation neural networks (Li et al., 2017). The comparison of neural networks’ convolution and long short-term memory versions, among many others, shows that the latter performs better than the former. Furthermore, it is shown that long short-term memory neural networks perform better than SVRs, back-propagation neural networks and evolution neural networks (Liu and Liu, 2019). Another study shows that SVRs are not as effective as long short-term memory dense neural networks. SVRs outperform linear regressions, boosting, decision trees and random forests in studies comparing machine learning techniques with conventional econometric methods, and XGBoost outperforms linear regressions, LASSO (least absolute shrinkage and selection operator), Ridge, random forests, bagging and boosting (Yan and Zong, 2020). However, random forests are shown to perform better than nearest neighbors, linear regressions, the XGBRegressor, CatBoost, AdaBoost, linear LASSO and linear Ridge (Rico-Juan and de La Paz, 2021). Furthermore, it is discovered that neural networks (Peterson and Flanagan, 2009; Abidoye and Chan, 2018), random forests, boosting, Ridge, LASSO and bagging (Embaye et al., 2021) all outperform linear regressions in terms of prediction accuracy. Random walk and Bayesian (vector) autoregression models are proven to perform worse than SVRs based on ensemble empirical mode decomposition (Plakandaras et al., 2015). These benefits are also evident when contrasting AR integrated moving average models with multilayer perceptron neural networks. Furthermore, studies show that when compared to deep learning neural networks, time-series econometric methodologies and a variety of other machine learning methods, there is no obvious benefit or drawback (Milunovich, 2020). Empirical results on different machine learning algorithms to estimate real estate values are, predictably, equivocal. Certain real estate price time series under consideration may have an impact on the machine learning techniques used by various researchers. Research similar to the ones listed here often point out benefits of machine learning techniques over conventional econometric ones. Our work is one of several that focuses on a single machine learning technique (Lam et al., 2008; Azadeh et al., 2014; Ma et al., 2015; Morano et al., 2015; Abidoye and Chan, 2017; Ćetković et al., 2018; Li et al., 2018; Kang et al., 2020; Yasnitsky et al., 2021; Wang et al., 2014; Li et al., 2020; Gu et al., 2011; Rafiei and Adeli, 2016; Chen et al., 2017) using Gaussian process regressions to anticipate real estate price time series. It has been suggested in a number of different research that integrating several models is a workable way to increase forecast accuracy. These include combining neural networks with multiple linear regressions, averaging dynamic models (Wei and Cao, 2017) and using neural networks in conjunction with case-based reasoning. These research use machine learning techniques to try to make up for any linear econometric underestimations of real estate pricing nonlinearities.
Various machine learning techniques have shown encouraging accuracy for real estate price projections in the literature. Based upon diverse evidence from empirical research, the measurement of the mean absolute percentage forecast errors could be below 1% (Liu and Wu, 2020; Ma et al., 2015; Pai and Wang, 2020; Liu and Liu, 2019; Ho et al., 2021), stretching from 1% to 2% (Liu and Wu, 2020; Ma et al., 2015; Pai and Wang, 2020; Rico-Juan and de La Paz, 2021), stretching from 2% to 3% (Pai and Wang, 2020; Liu and Liu, 2019; Plakandaras et al., 2015), stretching from 3% to 4% (Liu and Wu, 2020; Wang et al., 2014; Rafiei and Adeli, 2016), stretching from 4% to 5% (Kang et al., 2020; Wang et al., 2014), stretching from 5% to 6% (Liu and Wu, 2020; Kang et al., 2020; Li et al., 2020; Plakandaras et al., 2015), stretching from 6% to 7% (Liu and Wu, 2020; Kang et al., 2020), stretching from 7% to 8% (Pai and Wang, 2020; Kang et al., 2020) and stretching from 8% to 9% (Kang et al., 2020) [2]. More precisely, creating random forecasts (Ho et al., 2021), gradient boosting (Ho et al., 2021), neural networks (Ma et al., 2015), modified Holt’s exponential smoothing (Liu and Wu, 2020) and SVRs (Pai and Wang, 2020; Ho et al., 2021) results in the mean absolute percentage forecast error that is less than 1%. Creating neural networks (Ma et al., 2015), XGBoost, SVRs (Pai and Wang, 2020) and modified Holt’s exponential smoothing (Liu and Wu, 2020) results in the mean absolute percentage forecast error stretching from 1% to 2%. Creating SVRs (Liu and Liu, 2019; Plakandaras et al., 2015), classification and RT (Pai and Wang, 2020) and neural networks (Liu and Liu, 2019) results in the mean absolute percentage forecast error stretching from 2% to 3%. Creating neural networks, SVRs (Wang et al., 2014), deep restricted Boltzmann machines (Rafiei and Adeli, 2016) and modified Holt’s exponential smoothing (Liu and Wu, 2020) results in the mean absolute percentage forecast error stretching from 3% to 4%. Creating genetic algorithms (Kang et al., 2020), neural networks and SVRs (Wang et al., 2014) results in the mean absolute percentage forecast error stretching from 4% to 5%. Creating neural networks (Liu and Wu, 2020), SVRs (Li et al., 2020; Plakandaras et al., 2015) and genetic algorithms (Kang et al., 2020) results in the mean absolute percentage forecast error stretching from 5% to 6%. Creating neural networks (Liu and Wu, 2020) and genetic algorithms (Kang et al., 2020) results in the mean absolute percentage forecast error stretching from 6% to 7%. Creating genetic algorithms (Kang et al., 2020) results in the mean absolute percentage forecast error stretching from 7% to 8%. And creating neural networks (Pai and Wang, 2020) and genetic algorithms (Kang et al., 2020) results in the mean absolute percentage forecast error stretching from 8% to 9%. A model may perform well even with significantly larger mean absolute percentage forecast errors because different time-series data have different properties and some are harder to anticipate than others.
Last but not least, it is critical to remember that, like conventional econometric techniques, machine learning methods – including the Gaussian process regression models considered here – used for predictions may run into the overfitting or underfitting problem. Thus, we considered the trade-off between model prediction accuracy and forecast stabilities while developing the final model parameters and estimating pre-owned residential real estate price indices across the ten cities using Gaussian process regressions. The complexity of the model’s implementation may increase when comparing machine learning approaches to conventional econometric methodologies. Thus, it could be advisable to recommend simpler model topologies, especially for prediction users who lack experience. The Gaussian process regression models developed here have very basic structures as its foundation, can be implemented with ease and differ just little in running time from several classic econometrics models.
3. Data
The China Real Estate Index System (CREIS) provided the study’s data. It is an analytical tool designed to show how the main Chinese cities’ real estate markets are doing as well as their growth patterns. The National Real Estate Development Group Corporation, the Development Research Center of the State Council and the Real Estate Association developed the platform in 1994. Experts from the Banking Regulatory Commission, the Real Estate Association, the Development Research Center of the State Council, the Ministry of Land and Resources, the Ministry of Construction and other universities audited CREIS in 1995 and 2005. These days, CREIS releases a variety of real estate price indices each month, such as price indices for villas, retail real estate, office space, rental properties and residential real estate sales price indices for both newly built and existing homes, among many other categories. From a platform, this system has grown to encompass most of the Chinese real estate markets (Yang et al., 2013). In this work, we primarily use pre-owned residential real estate price indices to investigate forecasting problems.
The following ten major Chinese cities are included in the pre-owned residential real estate price indices gathered from CREIS: Chengdu, Wuhan, Nanjing, Hangzhou, Guangzhou, Shenzhen, Chongqing, Tianjing, Shanghai and Beijing. Data is gathered by phone, field and online surveys, according to CREIS. Every pre-owned residential property, excluding villas, in a given city that is sold in a given month is included in the samples used to create the price index. With an index value of 1,000, the base-period price index is based on the price of pre-owned residential real estate in Beijing in 12M2004. Pre-owned residential real estate price indices for different places and months are built by normalizing against this base-period price. The following is how CREIS calculates its pre-owned residential real estate price index for a given city: , where and reflect price indices associated with time t and time , respectively, expresses Project i’s total area of construction corresponding to time , and and signify average prices of Project i at time t and time , respectively. It is crucial to emphasize that the only data examined in the current inquiry are the pre-owned residential real estate price indices for each of the ten cities. No more CREIS platform data is available to us. The monthly data used in this research spans the period 3M2012–5M2020.
The ten price indices, their first differences, quantile–quantile plots and distributional plots using histograms and kernel estimations are all visualized in Figure 1. Table 1 also summarizes the ten price indices as well as their first differences. Based on the p-values of the Jarque–Bera test displayed in Table 1, none of the ten price indices exhibits a normal distribution at the 5% significance level. Non-normality is probably not startling for a wide range of financial and economic time-series data types. The price indices of Beijing, Shanghai and Shenzhen are left-skewed and those of other seven cities are right-skewed. For each of the ten cities, the platykurtic price index is determined.
Expressions of nonlinear characteristics at higher moments have previously been widely reported in the domains of finance and economics throughout a broad spectrum of time-series data (Yang et al., 2008; Wang and Yang, 2010; Yang et al., 2010). We use the Brock–Dechert–Scheinkman (BDS) (Brock et al., 1996) test to look for any possible nonlinear patterns in the ten time-series data sets of pre-owned residential housing prices. We apply the BDS test using 2–10 as the embedding dimensions and 0.5, 1.0, 1.5, 2.0, 2.5 and 3.0 multiplied by the standard deviation of a given price index time series as the ’s distance, which are used to evaluate the closeness of various data points. We find that all of the test’s resulting p-values are almost zero. These results imply that there are nonlinearities in each of the ten price indices. In light of these details, the goal of this study is to use Gaussian process regressions to forecast the ten non-normal and nonlinear pre-owned residential housing price indices.
4. Method
The forecasting method being studied in this work is the Gaussian process regression, a type of probabilistic kernel model that has been shown to be successful at predicting a range of nonlinear patterns in a number of scientific areas (Ou and Wang, 2011; Zhang and Xu, 2020a; Jin and Xu, 2024d; Zhang and Xu, 2020b; Han and Zhang, 2015). The training data having an uncertain distribution for the model’s representation are denoted by , the predictors in d-dimensions are shown by , and the target is represented by . Nine lagged price indices are used as predictors to forecast the pre-owned housing price indices for each city. For instance, nine price indices from the nine months prior will be used as predictors to anticipate the price index for the tenth month.
Let express a linear regression, where reflects the error term. In contrast, the target variable in Gaussian process regressions is defined using latent variables and explicit basis functions (Zhang and Xu, 2020c). It is possible to describe the basis function using b and the latent variables from the Gaussian process using in a way that satisfies the joint Gaussian distribution condition. The target’s smoothness will be represented by the latent variables’ covariance function, and the basis function’s objective is to project different predictors onto the feature space (Zhang and Xu, 2020d).
Two metrics that are commonly used to characterize a Gaussian process (GP) are the covariance and mean (Zhang and Xu, 2020e). We will use to represent the mean, and to express the covariance. Next, we will use to represent the Gaussian process regression, where and . We will parameterize via , a hyper-parameter. The following variables will often be estimated while training a Gaussian process regression using a particular technique: , and . In addition, we will specify the kernels (represented as k’s) and basis functions (represented as b’s) that will be used in the training of the model (Zhang and Xu, 2020f). Isotropic kernels and nonisotropic kernels – also known as automatic relevance determination kernels – are the two types of kernels that are taken into consideration in this study. Five different kernels are used to analyze both isotropic and nonisotropic kernels (Zhang and Xu, 2021a; Zhang and Xu, 2020g). The specifications for each of the kernels under consideration are provided by equations (1)–(10). The characteristic length scale of isotropic kernels is denoted by , the scale-mixture parameter is indicated by , the signal standard deviation is indicated by and . will be used to accomplish the positiveness of and (Zhang and Xu, 2021b; Xu et al., 2022). For nonisotropic kernels, each predictor will have a unique length scale, denoted by (). Accordingly, is reflected through (Zhang and Xu, 2021c; Zhang and Xu, 2021d):
This study considers four distinct basis functions, as described in equations (11)–(14), in a manner akin to the consideration of different kernels. In these equations: , , and .
The model parameters are estimated using tenfold cross-validation and Bayesian optimization, which is based on the expected improvement per second plus (EIPSP) approach (Zhang and Xu, 2021e; Zhang and Xu, 2021f). Let be the expression for a GP model. The Bayesian method chooses randomly picked data points of ’s inside variable boundaries to assess corresponding (Zhang and Xu, 2022a). In this case, data points were used to make preliminary assessments (Zhang and Xu, 2022a). When evaluation errors are encountered, the algorithm will continue to acquire data until it reaches successful evaluation cases (Zhang and Xu, 2021g). Then, as seen below, the algorithm’s first and second phases are repeated. The posterior distribution over will be generated starting with the updating of . Selecting a new data point (x) to ascertain the acquisition function’s () reduction objective will be the second stage. There will be a maximum of 100 iterations (Zhang and Xu, 2021h). The use of is to assess the goodness of x with respect to Q. Expected improvement acquisition functions assess expected quantities of improvements to the objective function, as opposed to values that would raise the objective function (Zhang and Xu, 2021i). We will let represent the data point at which the lowest posterior mean is attained, and indicate the corresponding numerical value of the lowest mean. may be used to indicate the expected improvement (EI). By applying the time-weighting method on the acquisition function, the Bayesian technique can provide greater benefits per unit of time, as the location can affect how long it takes to evaluate the target (Zhang and Xu, 2022b). It is feasible to maintain an additional Bayesian model of the time required to assess the goal as a function of x throughout optimization procedures. Considering this, we may write to describe the EI per second (EIPS) of the acquisition function, where denotes the posterior mean related to this additional timing GP model. The following modifications to the acquisition function’s behavior might be made to prevent it from overusing a particular region and from preventing a local minimum of the goal (Zhang and Xu, 2021j). Let represent the standard deviation of the posterior objective corresponding to x, and let represent the posterior standard deviation of the additive noises satisfying the formula . We will use to describe the exploration ratio. The acquisition function using the EIPSP algorithm checks if the next data point, x, satisfies the condition after each iteration. If this condition is satisfied, will be multiplied by the number of repetitions, and x will be considered overexploiting, and the kernel function will be adjusted accordingly (Bull, 2011). For data points between observations, the EIPSP approach modification essentially increases (Xu and Zhang, 2022k). The newly fitted kernel will then be used to create a new data point. In the event that the new data point is found to be overly exploitative, will be raised ten times in the next trials. To get a data point, x, that is not deemed overexploited, this approach will be used up to five times (Zhang and Xu, 2021k). The EIPSP algorithm will then accept the changed x as the subsequent exploration ratio. The algorithm finds a balance between examining fresh data points and previously examined neighboring data points to produce an overall answer that is more accurate (Zhang and Xu, 2021l).
We will perform Bayesian optimization operations over , basis functions, kernels and standardization status of predictors (Zhang and Xu, 2022c). The RRMSE, which enables comparisons of various prediction outcomes across several models or goals, will be used to assess forecast performance (Despotovic et al., 2016). The RRMSE may be written like this: , where n denotes the number of observations used for performance evaluations, expresses the target’s predicted numerical value and indicates the observed numerical value of the target variable. The mean absolute error (MAE) and root mean square error (RMSE), whose units match the target variable and whose magnitude is tied to the target variable, are two additional performance measures that are used to evaluate prediction accuracy. One way to represent the RMSE is: . And one way to express the MAE is: . Finally, we additionally take into account the correlation coefficient (CC) for performance evaluation, which is written as ,where and express averages.
5. Result
Data from pre-owned residential housing price indices for each city are used for model performance testing for projections one month in advance from 6M2019 to 5M2020, as well as for model training from 3M2012 to 5M2019. The results of EIPSP optimizations based on training data for all price indices are displayed in Figure 2. These findings suggest that (a) the isotropic exponential kernel [equation (1)], empty basis function [equation (11)] and standardized predictors are selected for the price index of Beijing, (b) the isotropic exponential kernel [equation (1)], constant basis function [equation (12)] and standardized predictors are selected for the price index of Shanghai, (c) the isotropic exponential kernel [equation (1)], empty basis function [equation (11)] and nonstandardized predictors are selected for the price index of Tianjing, (d) the nonisotropic Matern 3/2 kernel [equation (10)], constant basis function [equation (12)] and nonstandardized predictors are selected for the price index of Chongqing, (e) the isotropic Matern 5/2 kernel [equation (3)], empty basis function [equation (11)] and standardized predictors are selected for the price index of Shenzhen, (f) the isotropic Matern 5/2 kernel [equation (3)], linear basis function [equation (13)] and standardized predictors are selected for the price index of Guangzhou, (g) the isotropic exponential kernel [equation (1)], empty basis function [equation (11)] and standardized predictors are selected for the price index of Hangzhou, (h) the isotropic exponential kernel [equation (1)], constant basis function [equation (12)] and standardized predictors are selected for the price index of Nanjing, (i) the isotropic Matern 3/2 kernel [equation (5)], linear basis function [equation (13)] and nonstandardized predictors are selected for the price index of Wuhan, and (j) the nonisotropic exponential kernel [equation (6)], empty basis function [equation (11)] and standardized predictors are selected for the price index of Chengdu. The results of parameter estimates for the ten GPR models developed using the tenfold cross-validation for each city’s pre-owned residential housing price index are displayed in Table 2. The initials “CV1,” “CV2,” …, and “CV10” are used to signify these parameter estimation results, where ‘CV” refers to “cross-validation.”
The ten GPR models developed for each city’s pre-owned residential housing price index, Models “CV1,” “CV2,” …, and “CV10,” are shown in Table 2 and are used to forecast the price index’s numerical values for the testing period of 6M2019–5M2020. Therefore, there will be ten anticipated values for each price index for each month of the testing period. The average of the ten estimates is used to get the final price index estimate for that specific month. This method effectively eliminates any potential idiosyncratic forecasts made by a certain submodel, which might contribute to future projections that are stable and dependable. The advantages and desired characteristics of equal weighing have been covered in the literature (Costantini et al., 2017). Figure 3 compares the pre-owned residential housing price indices for each city, both predicted and actual. The percentage forecast errors are shown in Figure 4, which is consistent with Figure 3. It is evident that observed price indices and anticipated price indices generally track each other rather closely. For the findings in Figures 3 and 4, additional prediction performance data in terms of the RMSE, RRMSE, MAE and CC are summarized in Table 3. Specifically, the ten price indices have RRMSEs ranging from 0.0458% to 0.3035%. According to earlier studies (Despotovic et al., 2016), model prediction accuracy might be ranked as follows: excellent if RRMSE, good if RRMSE, fair if RRMSE, and poor if RRMSE. These criteria indicate that the GPR models developed here have an excellent level of prediction accuracy.
Figure 5 presents the results of an error autocorrelation study that was carried out to assess the adequacy of the constructed models. The study is carried out for up to 20 lags, with an emphasis on normalized autocorrelations. These findings validate the general validity of the models and rule out any blatant autocorrelations. It may be noteworthy, nonetheless, that including the AR conditional heteroscedasticity effect into a prediction model can improve its performance, even in the face of contradictory empirical evidence [e.g. 183, 184].
We also analyze the models’ performance for a more recent time period. Specifically, we apply the models shown in Table 2 on pre-owned residential housing prices across the ten cities for the time period of February 2023–January 2024. Recalling that the models use nine lagged series values as predictors, price forecasts are generated for November 2023, December 2023 and January 2024, which are shown as red dots in Figure 6. When comparing the forecasted and observed prices for each city during these three months, as depicted in Figure 6, it could be seen that the forecasts closely track observed prices. This further suggests the usefulness of the constructed models here.
6. Result analysis of contrast forecast models
So far, the investigation has mostly concentrated on regressions using the Gaussian process. The AR model, Arkansas-generalized autoregressive conditional heteroscedasticity (abbreviated as AR-GARCH) model, SVR model and RT model serve as four alternative models considered here for benchmark purposes. While comparing forecast accuracy of these different models, two measures are adopted: the RRMSE and the modified Diebold–Mariano (MDM) test (Harvey et al., 1997), latter of which evaluates variations related to mean squared errors (MSEs) of forecast results based upon two models. The basis for the MDM test is , where indicates the error term derived from model at time t and indicates the error term derived from model . would be one specific benchmark model (AR, Arkansas-GARCH, SVR or RT) and would be the GPR for the price index of a given city. The way to represent the MDM test statistic is through , where the length of the time period used for contrasting performance is represented by T, the horizon is denoted by h (in our situation, ), ’s average is , ’s variance is and ’s kth auto-covariance is for and . The equality of expected forecast performance generated by two distinct models is the null hypothesis of the MDM test. The MDM test would come after the t-distribution with degrees of freedom under the null hypothesis. The AR and AR-GARCH models use the same number of lags as the predictors in the GPR. The GARCH component has the form of GARCH(1,1). The predictors used by the GPR are also used by the SVR and RT models. The classification analysis and regression tree is the method that the RT model uses and the linear -insensitive version SVR is applied here. In the phase of out-of-sample testing, Figure 7 presents findings of benchmark analysis based upon RRMSEs. It is evident that the GPR has superior accuracy because, for each city’s price index, it yields the lowest RRMSE. The GPR is also compared with all four benchmark models in the MDM tests, and the resultant p-values are less than 0.01 for each city’s price index. This suggests that, when compared with the four benchmark models under investigation, the GPR generates accuracy that is statistically significantly better.
7. Implication
Forecasts of residential real estate price indices are a significant concern for governments and investors alike. For the purposes of risk management, strategic planning and portfolio allocation and modifications, investors require real estate price projections. For the purpose of assessing the market, developing, implementing and modifying policies – particularly to avoid overheating the market and stimulate the economy when needed – real estate price projections are essential to policymakers. To the best of the authors’ knowledge, econometric techniques – particularly time-series approaches where price indices are relevant – are frequently the basis for forecasting and valuation methodologies used by many investors, including those in the public sector. Professional opinions from specialists are also still used. This makes sense since econometric techniques and expert opinions are probably not too difficult to create, use and maintain. They have also been extensively used for many years by forecast users, and many of them may be able to provide a good degree of prediction accuracy. Although some decision-makers still see machine learning designs as unduly complex tools for making forecasts, it may be challenging for some policymakers and investors to take these models into consideration. Nevertheless, it is generally agreed that these models are worth exploring for their potential, especially in light of the growing accessibility of computational capabilities and the realistic basis for potential irregularities in price time-series data. In fact, a number of astute investors and decision-makers have recently shown a growing interest in machine learning techniques for predicting real estate prices. The work being done here carries on the legacy of examining how Gaussian process regressions may be used to tackle forecasting issues for indices of pre-owned residential housing prices. The models developed here could be straightforwardly applied on historical housing price indices of pre-owned residential properties of a city, with the reported model parameters, to generate price index forecasts into the future, thus offering necessary forecast information as part of various decision-making processes. It is worth attention that the present methodology does not take into account other economic factors as potential predictors but focuses on historical information for sources of predictive information. Thus, when using forecast results produced by the models constructed here, one might also combine them with other sources of forecasts. Under this circumstance, appropriate model combinations might need to be established. Overall, the presented methodology here for creating these forecast models for the ten largest Chinese cities, along with their demonstrated high prediction stability and accuracy, indicate that machine learning techniques are worthwhile exploring – possibly for a wider range of real estate types and locations.
8. Conclusion
The current study’s focus is on pre-owned residential housing price index estimations for ten significant Chinese cities. We create forecast models using monthly data from 3M2012 to 5M2020 using the Gaussian process regression technique. We pay special attention to four basis functions, ten kernels and two predictor standardization techniques while developing prediction models with Bayesian optimizations and cross-validation. We find that the developed models yield strong out-of-sample forecasts, with RRMSE ranging from 0.0458% to 0.3035% for the ten price indices during a one-year period from 6M2019 to 5M2020. Market participants and policymakers may find these forecast models useful in improving their comprehension of the pre-owned residential housing sector. Other Bayesian optimization strategies, in addition to the expected improvement per second plus method considered here, might be interesting to explore in future study. It is also possible to expand the forecasting procedure to include a wider range of real estate price indices and more cities/regions.
Notes
See ASKCI Consulting Co., Ltd., www.askci.com/
None of the earlier studies that are cited here and shown relatively substantial predicted errors have been classified by us.
Funding: There is no funding.
Conflict of interest: There is no conflict of interest.







