This study aims to explore the development of automatic document classification, analyzes its related applications in library contexts, identifies key challenges and emerging opportunities and proposes future research directions.
The study examines historical and modern document classification methods. It synthesizes technical progress with insights from real-world cases. Furthermore, it investigates the implementation challenges libraries face and the opportunities offered by emerging technologies and proposes potential directions for future research.
Automatic cataloging and document classification have evolved significantly. AI-driven approaches have brought notable improvements in efficiency and consistency. Tools such as Annif, JEX and AutoMSC have demonstrated practical utility in real-world settings. However, several challenges remain, including scarcity of data and annotated resources, insufficient model interpretability, standard incompatibility and ethical concerns. Meanwhile, opportunities are emerging through human-AI collaboration, data-sharing initiatives, open-access models and advancements in multimodal and multilingual technologies.
This paper brings together recent developments and real-world examples to provide a comprehensive view of how automatic document classification and cataloging are evolving in libraries today. It provides insights for researchers, library professionals and technologists working to build more intelligent, equitable and inclusive library services.
