This study addresses a key technical challenge in automated subject indexing: improving prediction accuracy for large-scale library collections. We introduce and validate a hybrid AI framework that combines statistical and semantic embedding approaches to enhance the performance of automated subject heading assignment. The primary aim is to demonstrate how a minimal, targeted application of a powerful semantic embedding model can significantly boost the accuracy of a high-speed statistical classifier, particularly for descriptive, context-rich bibliographic titles.
A large-scale validation was performed using 2.4 million real-world bibliographic records from the National Library of Spain. We developed a hybrid framework combining a high-speed statistical engine (Omikuji) with an advanced semantic model (Qwen3-Embedding-8B). A key finding emerged during validation: an optimal “97/3” resource-efficient blend, where the statistical model handles 97% of the weight, and the semantic embedding model provides a minimal 3% corrective signal. The system's prediction accuracy was then evaluated on distinct test sets, including three stratified by title length (short, medium and long) to assess performance across varying levels of contextual richness.
The results demonstrate significant accuracy improvements. First, the “less is more” hypothesis is confirmed: a minimal 3% semantic contribution proves optimal, substantially improving baseline performance. Second, the enhancement effect correlates strongly with title length and contextual richness. While the hybrid model achieved a 7.3% accuracy gain (nDCG@5) on general records, performance improvements were more pronounced for longer titles: +24.9% for medium-length titles and +40.1% for long, descriptive titles, precisely the materials where statistical models typically struggle most.
This study provides empirical evidence for an effective method of integrating semantic embedding and statistical approaches in automated subject indexing. Its contribution lies in demonstrating that substantial accuracy gains can be achieved through minimal semantic enhancement rather than complete system replacement. By systematically evaluating performance across different title complexities, this research offers practical guidance for institutions seeking to improve their automated subject indexing systems, particularly for descriptive bibliographic records that challenge purely statistical approaches.
