Missing data represents a fundamental and pervasive challenge in modern data science, significantly impeding analytical capabilities and decision-making processes across an exceptionally broad spectrum of disciplines, including healthcare, bioinformatics, social science, e-commerce and industrial monitoring systems. Despite decades of research and the development of numerous imputation methodologies, existing literature remains fragmented across disciplinary boundaries, creating a critical need for a comprehensive, interdisciplinary synthesis that bridges statistical foundations with contemporary machine learning advances. This work systematically covers fundamental concepts – including missingness mechanisms, single vs. multiple imputation and varying imputation goals – and explores problem characteristics across different domains. The review extensively categorizes imputation methods, spanning classical techniques (e.g. regression and expectation-maximization algorithm) to modern approaches such as low-rank and high-rank matrix completion, deep learning models (autoencoders, generative adversarial networks, diffusion models and graph neural networks) and large language models. Special consideration is given to methods tailored for complex data types, including tensor data, time series, graph-structured data, categorical data and multimodal data, acknowledging their unique challenges and solution approaches. Beyond methodological considerations, they investigate the crucial integration of imputation with downstream machine learning tasks, including classification, clustering and anomaly detection, examining both sequential pipelines and joint optimization frameworks. The review also assesses theoretical guarantees for various methods, available benchmarking resources and comprehensive evaluation metrics. Finally, they identify critical challenges and future directions, emphasizing the complexities of model selection and hyperparameter optimization, the growing importance of privacy-preserving imputation through federated learning approaches and the ambitious pursuit of generalizable or universal imputation models that can adapt across domains and data types, thereby providing a roadmap for advancing this vital field of research.
Article navigation
6 May 2026
Research Article|
May 01 2026
An interdisciplinary and cross-task review on missing data imputation
Jicong Fan
School of Data Science,
The Chinese University of Hong Kong
, Shenzhen, China
Corresponding author Jicong Fan fanjicong@cuhk.edu.cn
Search for other works by this author on:
Corresponding author Jicong Fan fanjicong@cuhk.edu.cn
Received:
November 15 2025
Revision Received:
January 23 2026
Accepted:
February 02 2026
Online ISSN: 1932-8354
Print ISSN: 1932-8346
Funding
Funding Group:
- Award Group:
- Funder(s): National Natural Science Foundation of China under Grant
- Award Id(s): 62376236
- Funder(s):
- Award Group:
- Funder(s): General Program of Natural Science Foundation of Guangdong Province
- Award Id(s): 2024A1515011771
- Funder(s):
- Funding Statement(s): This work was supported by the National Natural Science Foundation of China under Grant No.62376236 and the General Program of Natural Science Foundation of Guangdong Province under Grant No. 2024A1515011771.
© 2026 Jicong Fan.
2026
Jicong Fan
Licensed re-use rights only
Foundations and Trends in Signal Processing (2026) 20 (3): 185–317.
Article history
Received:
November 15 2025
Revision Received:
January 23 2026
Accepted:
February 02 2026
Citation
Fan J (2026), "An interdisciplinary and cross-task review on missing data imputation". Foundations and Trends in Signal Processing, Vol. 20 No. 3 pp. 185–317, doi: https://doi.org/10.1108/FTSIG-11-2025-0139
Download citation file:
43
Views
Suggested Reading
Bayesian temporal factorization and improved transformer architecture for the prediction of aero-engine remaining useful life
Engineering Computations (July,2026)
The Q -matrix completion problem
Arab Journal of Mathematical Sciences (July,2020)
Three-way formal concept clustering technique for matrix completion in recommender system
International Journal of Pervasive Computing and Communications (January,2020)
An elevator failure mode and effects analysis method based on retrieval augmented generation
Journal of Intelligent Manufacturing and Special Equipment (March,2026)
Related Chapters
Missing-Data Imputation in Nonstationary Panel Data Models
Missing Data Methods: Time-Series Methods and Applications
Exploring the Potential of AI for Authentic Assessment in Education: Towards a New Model of Interaction
The Emerald Handbook of Active Learning For Authentic Assessment
AI and Knowledge Management: Navigating the Misinformation Maze
AI-Driven Knowledge Management Processes, Volume 1: Strategies for the Modern Business Landscape
Recommended for you
These recommendations are informed by your reading behaviors and indicated interests.
