This paper aims to propose a method based on multi-source data and deep learning to early identify emerging technologies (ETs). On the one hand, this paper considers sufficient data sources related to the key characteristics of ETs, which can help enterprises and countries grasp technological innovation opportunities and market development trends in advance. On the other hand, this paper proposes a method based on deep learning to fit the complex and non-linear relationships more accurately between features, thereby achieving the early identification of ETs.
First, multi-source data sets are collected for emerging technology identification, such as papers, patents and industry reports, to cover multi-dimensional characteristics of ETs from different perspectives. Second, the corresponding features are designed and extracted from multi-source data, to expand and supplement current insufficient features. Finally, a time series forecasting model based on deep learning is designed to fuse multiple features, which can better fit complex and non-linear relationships between features.
The experiment on artificial intelligence (AI) demonstrates that integrating multi-source data can improve the performance of the model, with each individual data source contributing unique predictive value. Among them, patent data serves as the strongest individual predictor, while combining it with academic papers and industry reports yields the best performance by capturing diverse signals. In addition, long short-term memory networks (LSTM) achieve the strongest overall regression and ranking performance among the four compared models.
This work expands a single data source into multiple data sources, by fully considering knowledge impact, technology impact and market impact to comprehensively identify emerging technology. This study classifies multiple features into different categories based on the characteristics of ETs and design methods to extract corresponding features from multi-source data. This study designs a deep learning model to fuse multiple features for better fitting complex relationships among them. Compared with representative linear, tree-based and recurrent baseline models, LSTM achieves the strongest overall regression and ranking performance.
