Article navigation

Missing data represents a fundamental and pervasive challenge in modern data science, significantly impeding analytical capabilities and decision-making processes across an exceptionally broad spectrum of disciplines, including healthcare, bioinformatics, social science, e-commerce and industrial monitoring systems. Despite decades of research and the development of numerous imputation methodologies, existing literature remains fragmented across disciplinary boundaries, creating a critical need for a comprehensive, interdisciplinary synthesis that bridges statistical foundations with contemporary machine learning advances. This work systematically covers fundamental concepts – including missingness mechanisms, single vs. multiple imputation and varying imputation goals – and explores problem characteristics across different domains. The review extensively categorizes imputation methods, spanning classical techniques (e.g. regression and expectation-maximization algorithm) to modern approaches such as low-rank and high-rank matrix completion, deep learning models (autoencoders, generative adversarial networks, diffusion models and graph neural networks) and large language models. Special consideration is given to methods tailored for complex data types, including tensor data, time series, graph-structured data, categorical data and multimodal data, acknowledging their unique challenges and solution approaches. Beyond methodological considerations, they investigate the crucial integration of imputation with downstream machine learning tasks, including classification, clustering and anomaly detection, examining both sequential pipelines and joint optimization frameworks. The review also assesses theoretical guarantees for various methods, available benchmarking resources and comprehensive evaluation metrics. Finally, they identify critical challenges and future directions, emphasizing the complexities of model selection and hyperparameter optimization, the growing importance of privacy-preserving imputation through federated learning approaches and the ambitious pursuit of generalizable or universal imputation models that can adapt across domains and data types, thereby providing a roadmap for advancing this vital field of research.

Licensed re-use rights only
You do not currently have access to this content.
Don't already have an account? Register

Purchased this content as a guest? Enter your email address to restore access.

Pay-Per-View Access
$108.00
Rental

or Create an Account

Close Modal
Close Modal