This study evaluates the potential of large language models (LLMs) for subject indexing of Swedish LGBTQ + fiction using the dedicated QLIT thesaurus, based on Homosaurus. It examines both the quantitative performance and qualitative characteristics of AI-generated subject terms, with particular attention to thesaurus compliance and indexing policy.
Four general-purpose AI tools (ChatGPT, Claude, Gemini, and DeepSeek) were evaluated on a collection of 14 historical LGBTQ + literary works represented by 16 texts. Subject terms were generated from both metadata records and full texts and compared with existing indexing in the Queerlit database of Swedish LGBTQ + fiction. The metadata subset produced 518 AI-generated subject terms in total, while the full-text subset produced 590 terms. The analysis was based on a coding scheme, which was refined through several rounds of analysis until complete agreement was achieved.
Performance was substantially higher for metadata than for full texts, with average F1 scores of 44% and 12%, respectively. Across both datasets, fully incorrect terms constituted the largest category of generated outputs (37.5% for metadata and 56.1% for full texts), while fully correct terms accounted for only 30.0 and 8.3%, respectively. Qualitative analysis revealed recurring problems, including the assignment of broader concepts rather than, or in addition to, more specific concepts, failure to follow thesaurus definitions and indexing policies, confusion between general themes and LGBTQ + -specific themes, and the application of contemporary LGBTQ + concepts to historical texts. Poetry proved particularly challenging because themes were often implicit and open to interpretation. Although a small proportion of generated terms were judged potentially useful despite not appearing in existing indexing, no AI-generated terms were classified as fully correct additions beyond the current Queerlit metadata.
The study focuses on a relatively specific collection of Swedish historical LGBTQ + fiction and a specialized controlled vocabulary. The findings nevertheless demonstrate the importance of complementing quantitative measures with qualitative evaluation and suggest that assessment of automated subject indexing should consider thesaurus definitions, indexing policies, and interpretive aspects of literary aboutness in addition to agreement with existing metadata.
The results indicate that current general-purpose LLMs are unlikely to be suitable as semi-automated indexing tools for historical LGBTQ + fiction without substantial human oversight. While AI systems may assist in identifying candidate concepts and additional access points, effective indexing continues to depend on expert knowledge of controlled vocabularies and indexing policy.
This study contributes to research on AI-assisted subject indexing by combining quantitative and qualitative evaluation of LLM-generated subject terms in a specialized LGBTQ + knowledge organization context. It demonstrates the limitations of gold-standard-based evaluation for fiction indexing and highlights the role of human expertise in applying controlled vocabularies to historically and culturally complex literary materials.
Introduction
Although (semi-)automated subject indexing has long been explored as a means of increasing efficiency and scalability in libraries, the advent of large language models (LLMs) has introduced new opportunities. Their potential application to fiction indexing is especially important because fiction is typically indexed primarily by genre, place, and form, while users often search for works based on themes or topics (Saricks and Wyatt, 2019). By generating additional subject terms that capture themes, AI-based indexing may help address this longstanding gap between fiction indexing practices and users' information needs. However, while more and more researchers engage with using LLMs for subject indexing (Xu et al., 2026), previous research on the use of LLMs for subject indexing remains limited in scope and has generally produced weak or inconsistent results.
Moreover, most evaluations of automated subject indexing have relied on gold standard approaches that compare AI-generated terms with existing metadata using quantitative measures such as precision, recall, and F1 scores. While these measures provide valuable insights into system performance, they offer only limited understanding of the quality and usefulness of generated index terms. This reveals a significant research gap in the evaluation of automated subject indexing beyond agreement with existing metadata, particularly in interpretive domains such as fiction. Because subject representation in fiction potentially allows for multiple valid indexing choices, generated terms should also be assessed in terms of their degree of relevance, specificity, consistency, and usefulness for subject access. In this context, expert assessment can complement traditional performance measures by identifying both valuable terms absent from existing metadata and clearly incorrect or misleading terms (Golub et al., 2016).
Building on previous findings that LLMs perform poorly in reproducing professionally assigned index terms to Swedish LGBTQ + fiction (Golub and Ihrmark, 2026), because aggregate measures such as precision, recall, and F1 score provide only a partial picture of AI indexing quality, the present study complements quantitative evaluation with a qualitative analysis of AI-generated index terms. It therefore examines whether four AI tools assign terms at an appropriate level of specificity, whether they generate duplicates, near-synonyms, or overly broad concepts, and whether they produce clearly erroneous terms. It also investigates whether AI-generated outputs can identify relevant index terms that have not been assigned manually and could therefore contribute to improved subject access.
The remainder of this paper is structured as follows. First, the Related research section outlines previous research on automated subject indexing with focus on LLMs and evaluation thereof. The Methodology section then describes the data collection and evaluation procedures used in the study. This is followed by the Results section, which focuses on qualitative analysis of AI-generated index terms. Finally, the Conclusion discusses the implications of these findings for queer knowledge organization and the use of generative AI in libraries, while also suggesting directions for future research.
Related research
Research on the use of LLMs for subject indexing and classification is still limited, and findings have varied considerably across studies and indexing tasks. Holstrom (2025) tested ChatGPT and Copilot on Library of Congress Subject Headings (LCSH) and general classification numbers, finding that outputs included unauthorized terms and inaccurate headings. Chow et al. (2024) reported that ChatGPT produced correct LCSH headings for only 23.3% of theses and dissertations, often lacking specificity or exhaustivity. Dobreski and Hastings (2025) compared ChatGPT, Copilot, and Gemini, finding overall poor performance, especially in class number assignment, although ChatGPT showed slightly stronger results with LCSH. Włodarczyk (2025) evaluated GPT-4o for Polish LIS articles, achieving only moderate accuracy (F1 of 34%).
Harada et al. (2025) used ChatGPT to generate book classification numbers from the Nippon Decimal Classification (NDC) system at three hierarchical levels. Their results suggest that strong performance can be achieved with relatively small training datasets, while large increases in training data do not always produce equally large gains. For example, increasing the dataset from 100,000 to 1.2 million documents only slightly improved top level classification (one digit) accuracy (76.3%–78.5%). However, improvements were more noticeable at the second hierarchical level (two digit) with accuracy increasing from 64.9% to 70.7% and for the third level (three digit) where the accuracy increased from 48.4% to 58.8%.
Tanácsi and Micsik (2026) evaluated several automated approaches to classify publications in the Hungarian Science Bibliography using the OECD Frascati taxonomy. Although LLMs were useful for improving dataset quality through filtering and noise reduction, they did not achieve the highest classification accuracy. The best-performing method was a Support Vector Classifier with F1 of 83%, surpassing both LLM-based and more complex architectures (cf. below referenced work on Annif: D'Souza et al., 2025b; Suominen et al., 2025).
Luo and Liu (2025) employed a complex approach to LLM-based automated subject indexing using LCSH, employing three steps: first, they fine-tuned an LLM (Llama) using 20,000 Library of Congress bibliographic records as training data in order to generate initial terms for a given document; second, a text embedding model was used to retrieve a set of candidate terms from the controlled vocabulary that are most similar to the initial terms; third, a term re-ranking model was trained to score the candidate terms, of which the top-ranked terms were selected as the final recommended terms. They found this combined approach superior to other tested approaches and demonstrated it reduces hallucinations in automatically generated terms, achieving 76% recall for top 100 suggested terms. However, only the most general topic from each subject heading was evaluated (MARC field 650 and subfield “a” for the main topical term), while subfields that indicated specificity and often result in poorer automation performance were not taken into account (such as “x” for general subdivision, “z” for a geographic subdivision etc.).
SemEval2025 Task 5 titled “LLM-based Subject Tagging for the TIB Technical Library's Open-Access Catalog” aimed to assess how effectively LLMs can assign subject tags to scientific and technical library records from the German National Library of Science and Technology (TIB) (D'Souza et al., 2025a). The evaluation used a blind test set of English and German records and required systems to generate a ranked list of the top 50 most relevant subject terms for each document using German subject headings (Sachbegriff, from the German integrated authority file (Gemeinsame NormDatei – GND)). It combined quantitative assessment measuring performance with precision, recall and F1 scores for the top k terms, also analyzing English vs German documents, different document types (articles, books, reports, conferences, theses) and combination of the two. Unusual related to other automated indexing research is that beyond gold standard comparison, a subset of records underwent manual assessment by TIB subject specialists. This process was intended to mimic real-world library workflows and determine whether the generated terms were practically useful according to human indexers. The qualitative results were also summarized using precision, recall, and F1 measures.
D'Souza et al. (2025b) summarise the results of the Task 5 participants, involving 12 systems developed by 12 teams. LLMs used included Llama and Qwen for synthetic training data generation; an LLM ensemble comprising Llama, Mistral, OpenHermes, Teuken; a BERT model ensemble; and, RAG-based ranking with Qwen. Major findings indicate that hybrid systems combining traditional machine learning approaches with LLMs outperform pure prompting-based methods. The best quantitative results across all subjects were achieved by Suominen et al. (2025) with an average recall of 63%. Their approach combines traditional machine learning methods within the Annif toolkit with LLM-based techniques for translation, synthetic data generation, and prediction merging from monolingual models. These results suggest that integrating traditional machine learning methods with modern LLM techniques can effectively improve multilingual subject indexing accuracy and efficiency, as also reported by Tanácsi and Micsik (2026), although for a different data set and controlled vocabulary.
An interesting finding was that quantitative rankings did not always align with human judgments, resulting in different performance rankings for the 12 evaluated systems depending on whether quantitative or qualitative evaluation criteria were applied. The best qualitative results were achieved by Kluge and Kähler (2025) with an average recall of 57% when both “correct” and “irrelevant but technically correct” terms were counted as correct, or 51% when only “correct” terms were counted as correct ones. Like Luo and Liu (2025), they used a complex workflow involving several steps: first, they prompted selected LLMs; second, they conducted a series of post-processing steps that mapped the AI-generated keywords to the target vocabulary; third, they aggregated the resulting subject terms to an ensemble vote; and finally, ranked them as to their relevance to the record. For the top 5 recommended terms, the F1 measure was 32%. This is one of few papers in related literature that conducted an error analysis. It showed that the system worked best when subject terms were specific and explicitly mentioned or paraphrased in the text. More general concepts were often incorrectly predicted or replaced with more specific alternatives. The system particularly struggled with abstract relationships between text and subject terms, where relevant concepts were implied rather than directly stated.
Research focusing specifically on Swedish LGBTQ + fiction has demonstrated substantial limitations in current AI approaches (Golub, 2025; Golub and Ihrmark, 2026). In the 2025 study, ChatGPT was tasked with generating subject index terms from QLIT, an LGBTQ + dedicated thesaurus based on Homosaurus, for 20 metadata records and full-texts from the Queerlit database of Swedish LGBTQ + fiction, which were then compared against metadata assigned by information specialists. Using both full texts and metadata records as input, the evaluation showed consistently low precision and recall scores, with F1 values of only 7% for full texts and 8% for metadata records. The study also demonstrated serious difficulties in identifying works as LGBTQ + at all. ChatGPT failed to identify any of the 20 full texts as LGBTQ+, while only 5 out of 20 metadata records were identified as LGBTQ+, despite the fact that the metadata records already contained all the correct LGBTQ + index terms. The results further showed that automatically generated LGBTQ + subject terms often represented overly broad concepts related to gender or sexuality. In addition, the generated metadata contained inconsistencies, duplicates, and factual errors, including the simultaneous assignment of near-synonymous concepts and incorrect temporal descriptors.
The 2026 study expanded this investigation by evaluating four leading LLMs—ChatGPT, Gemini, Claude, and DeepSeek—on a larger corpus of 78 full-texts and their metadata. Results demonstrated that performance remained poor across all systems, with F1 scores below 15% for full texts and below 55% for metadata records, even though the metadata records themselves contained the index terms that the models were asked to generate.
With the exception of the SemEval2025 experiments (D'Souza et al., 2025a), evaluation in the above referenced studies was largely limited to comparing AI-generated subject terms against existing metadata records. As a result, relatively little attention has been paid to broader qualitative questions, such as whether automatically generated terms were assigned at an appropriate level of specificity, or whether terms that did not match existing metadata might nevertheless be useful. This is particularly important because manual indexing is shaped by institutional policies regarding exhaustivity and specificity, intended users, and the interpretation of a document's aboutness. In fiction indexing, the issue of aboutness further complicates evaluation, as different indexers may assign different yet equally valid subject terms (Saarti, 2002). Indexing consistency is also influenced by organizational practices: when senior cataloguing librarians provide guidance and quality control across multiple indexers, consistency is likely to be higher than in environments without such oversight. Consequently, evaluation of automatically generated terms should therefore incorporate multiple perspectives rather than rely solely on predefined subject assignments, including expert assessment of the quality and usefulness of suggested terms, workflow-based evaluation, and retrieval effectiveness (Golub et al., 2016). At the same time, evaluation should consider the frequency and severity of clearly incorrect or misleading terms, as aggregate performance measures may obscure errors that affect the practical usefulness of automated indexing.
Methodology
Background
The study focuses on qualitative analysis of AI-generated index terms as presented in Golub and Ihrmark (2026). The former study's dataset consisted of full text of 78 Swedish literary works from the Swedish Literature Bank (Litteraturbanken), an open-access digital collection of Swedish literary texts, and their associated metadata records from the Queerlit database of Swedish LGBTQ + fiction. The works spanned publications from 1848–1935. Subject indexing was based on Queerlit's dedicated QLIT thesaurus of LGBTQ + terms, which contained 913 subject terms as of 1 June 2025.
The study evaluated four large language models: ChatGPT-4o (OpenAI), Claude Sonnet-4 (Anthropic), Gemini 1.5-flash (Google), and DeepSeek Chat (DeepSeek). The prompting strategy used a structured prompt design, where the models were assigned the role of a subject indexer specializing in LGBTQ + literature. Prompts divided the task into smaller steps, involving identifying LGBTQ + themes and suggesting terms from the QLIT vocabulary. Separate prompts were used depending on whether the input consisted of full text or MARC metadata, while the complete controlled vocabulary was appended to the prompt to guide the models toward selecting standardized QLIT terms rather than generating free terms. The LLMs were supplied with the full set of QLIT preferred terms; however, scope notes accompanying the vocabulary were not included due to token limit. Likewise, the QLIT indexing policy was not explicitly provided to the models.
A custom Streamlit dashboard was used to process inputs, apply the structured prompts, and compare generated subject terms against human-assigned QLIT terms using precision, recall, and F1 measures. Using the pre-existing Queerlit metadata as the gold standard, the results indicated that overall model performance remained limited. When indexing full texts, all models produced relatively weak results, with F1 scores below 15%. Even when models were given access to the actual human-assigned metadata (the facit) rather than only the literary text, performance remained below 55%.
Aims and purpose
The purpose of this study is to evaluate the potential of LLMs for assigning QLIT subject terms to historical LGBTQ + fiction and to examine the quality, appropriateness, and usefulness of AI-generated subject terms beyond conventional quantitative performance measures.
The study aims to:
Evaluate the ability of four general-purpose AI tools to assign QLIT subject terms to metadata records and full texts of historical LGBTQ + Swedish fiction by examining qualitative characteristics of AI-generated subject terms, including their specificity, adherence to thesaurus definitions and indexing policies, and the types of errors produced.
Investigate whether AI-generated terms that are not part of the existing indexing may nevertheless represent useful subject access points.
Explore the challenges of evaluating AI-generated subject indexing for fiction, particularly in relation to aboutness.
Coding protocol
The coding framework used for evaluating generated subject terms was developed iteratively through a collaborative process involving all the paper's authors. Initially, two researchers in the field of knowledge organization independently examined a subset of about 100 AI-generated terms and developed preliminary coding categories based on their interpretations of relationships between AI-generated terms and the existing QLIT vocabulary.
After the initial independent coding phase, the researchers compared and discussed their coding decisions to identify disagreements and align their interpretations. Based on these discussions, the coding categories and definitions were refined and a shared coding scheme was established. The revised scheme was then independently applied to a new set of examples, after which the researchers again compared their coding, resolved remaining disagreements, and further refined the definitions where necessary. Following the third iteration, complete agreement was reached on all examples, and the coding scheme was finalized for use in the study.
The final coding framework resulting from this iterative process consisted of the following categories:
Fully correct, when there was an exact string match with an assigned QLIT term for the work in the Queerlit database.
Broader, where the generated term corresponded to a broader term (BT) to one of the Queerlit terms in the QLIT thesaurus.
Narrower, where the generated term matched a narrower term (NT) in the QLIT thesaurus.
Related, where the generated term corresponded to a related term (RT) to one of the Queerlit terms in the QLIT thesaurus.
Incorrect by definition of QLIT or subject indexing policy, for terms judged to be inconsistent with QLIT definitions or Queerlit indexing policies.
Not assigned by Queerlit but fully correct, for terms that were not assigned in Queerlit but should be assigned.
Not assigned by Queerlit but partially correct, for terms that were not assigned in Queerlit but could be considered relevant.
Fully incorrect, conceptually a complete miss.
Coding procedure
To ensure rigorous and consistent evaluation, the coding process was designed as a multi-stage procedure intended to reduce fatigue effects and improve agreement between evaluators. In the first stage, two researchers in knowledge organization independently reviewed the generated subject terms and assigned them to an initial set of categories: “fully correct”, “broader”, “narrower”, “related”, and “unsure if incorrect” for all remaining cases. These decisions were made independently to allow each researcher to develop their own interpretations.
Following the independent coding phase, the two researchers compared results and discussed differences in interpretation. After a shared understanding had been reached and disagreements resolved, the material was forwarded to a Queerlit subject librarian for expert review. The librarian was instructed to check both the already assigned categories (“fully correct”, “broader”, “narrower”, “related”) and assess the remaining “unsure if incorrect” cases as one of the four remaining categories: “incorrect by definition of QLIT or subject indexing policy”, “not assigned by Queerlit but fully correct”, “not assigned by Queerlit but partially correct”, or “fully incorrect”.
Queerlit indexing policy (Queerlit, 2026) is guided by several core principles intended to ensure consistency and specificity in subject assignment. Indexers are expected to use the most specific available term in the hierarchy rather than assigning both broader and narrower concepts. Subject terms should always be assigned according to their formal definitions in the thesaurus, and preferred controlled vocabulary terms should be used instead of free-text terms. Unlike non-fiction in the Swedish Union Catalogue, where the 20% rule serves as a guideline suggesting that a topic should occupy approximately one-fifth of a work before being assigned as a subject term, this principle is not applied in the same explicit manner to fiction. Thus, subject assignment for LGBTQ + fiction in Queerlit relies more strongly on the interpretive significance and thematic prominence of LGBTQ + content within a work. A theme does not necessarily need to occupy a large proportion of the text to justify indexing; rather, its narrative importance, relevance to characters, relationships, or broader meaning within the work may determine whether a subject term is assigned.
After the subject librarian completed the evaluation and assigned the coding categories, the two researchers without subject-specialist expertise reviewed the results and examined the coding decisions in detail. Following this review, all three evaluators met to discuss and resolve remaining ambiguities. These discussions also served as an opportunity to reflect on broader observations that emerged during the coding process, including challenges related to the interpretation of thesaurus relationships and the application of indexing policies.
Data collection
Due to the intensive time and effort required for the detailed coding and validation process described above, it was not feasible to apply the procedure to the entire dataset from the original study. Therefore, a smaller subset of 16 literary works containing at least one major QLIT term was selected for in-depth qualitative analysis. The selection was also supported by the quantitative results from the broader dataset. This subset achieved the strongest full-text performance, with the best-performing models (Claude and DeepSeek) reaching an F1 score of 14%, compared with only 3% for works containing exclusively peripheral LGBTQ + themes and 11% for works containing a mixture of major and peripheral themes (Golub and Ihrmark, 2026).
To provide additional context for the metadata used in this analysis, the Queerlit records contain bibliographic information, such as the work's title, publication year, and author, alongside subject metadata describing its thematic content. The subject metadata include major and peripheral LGBTQ + terms, representing central and secondary queer themes, respectively, as well as general subject terms from the Swedish Subject Headings (SAO). These categories are visually distinguished by colour and font size, with major LGBTQ + terms receiving the greatest prominence. The metadata records for all literary works included in the qualitative analysis are available at the URLs listed further below.
The sample included 14 works represented by 16 texts, as two works were published in two volumes; the volumes had identical metadata but different full-text content. The collection primarily represents historical Swedish literature from the transition between the nineteenth and twentieth centuries, spanning a publication period of 87 years, ranging from 1848 (Jungfrutornet by Emilie Flygare-Carlén) to 1935 (För trädets skull by Karin Boye). Of these, eight are books of poetry and six novels; one of the novels is of the children's/young adult genre (Tvillingsystrarna by Ellen Idström from 1893). The works are written by nine authors, with Edith Södergran having the largest number of books (n = 4), all belonging to the poetry genre. Karin Boye and Vilhelm Ekelund each contributed two poetry books, while Emilie Flygare-Carlén (n = 2) and Martin Koch (n = 2) were represented by fiction prose works. The remaining authors—Anna Branting, Ellen Idström, Aurora Ljungstedt, and Anna Sandberg—each contributed a single work, and all of them novels. The full list of the works reads as follows:
Boye, Karin. Härdarna (1927). Genre: Poetry. URL: Link to the website
Boye, Karin. För trädets skull (1935). Genre: Poetry. URL: Link to the website
Branting, Anna. Lena (1893). Genre: Novel. URL: Link to the website
Ekelund, Vilhelm. In Candidum (1905). Genre: Poetry. URL: Link to the website
Ekelund, Vilhelm. Hafvets stjärna (1906). Genre: Poetry. URL: Link to the website
Flygare-Carlén, Emilie. Jungfrutornet, Vol. 1 (1848). Genre: Novel. URL: Link to the website
Flygare-Carlén, Emilie. Jungfrutornet, Vol. 2 (1848). Genre: Novel. URL: Link to the website
Idström, Ellen. Tvillingsystrarna (1893). Genre: Youth literature/Novel. URL: Link to the website
Koch, Martin. Guds vackra värld I (1916). Genre: Novel. URL: Link to the website
Koch, Martin. Guds vackra värld II (1916). Genre: Novel. URL: Link to the website
Ljungstedt, Aurora. Moderna Typer (Samlade berättelser 6) (1874). Genre: Novel. URL: Link to the website
Sandberg, Anna. Vampyrer (1874). Genre: Novel. URL: Link to the website
Södergran, Edith. Septemberlyran (1918). Genre: Poetry. URL: Link to the website
Södergran, Edith. Rosenaltaret (1919). Genre: Poetry. URL: Link to the website
Södergran, Edith. Framtidens skugga (1920). Genre: Poetry. URL: Link to the website
Södergran, Edith. Landet som icke är (1925). Genre: Poetry. URL: Link to the website
Both full-text and metadata-based outputs from the 16 books with at least one major QLIT term were included in the assessment. This resulted in a total of 1,108 subject terms being evaluated by each researcher independently in their entirety and then revisited multiple times during collaborative discussions. Of these, 518 terms were from the metadata subset and 590 from the full-text subset.
AI-generated theme labels outside the QLIT vocabulary
Closer qualitative inspection of the outputs also revealed examples that raise questions regarding the strict use of existing gold standards for evaluation. Occasionally, the models added the prefix “Tema:” to labels that do not belong to the QLIT thesaurus itself but rather to the nine general themes displayed on the Queerlit website for browsing purposes (Link to the website), without any instruction to do so. The nine top-level browsing categories on the Queerlit website are Discrimination, Hate and Violence (Diskriminering, hat och våld), Identities and Practices (Identiteter och praktiker), Culture and Leisure (Kultur och fritid), Worldviews, Beliefs and Theories (Livsåskådning, tro och teorier), Medicine (Medicin), Relationships (Relationer), Movements and Rights (Rörelser och rättigheter), Sex, Intimacy and Embodiment (Sex, intimitet och kroppslighet), and Other (Övrigt).
In the metadata subset, Gemini preceded one generated index term with the label “Tema:” (“Theme:“). In the earlier quantitative evaluation, this output would be automatically counted as incorrect because the additional formatting prevented an exact string match with the QLIT term. In this particular case, the QLIT index term was actually incorrect, so the results remain unchanged. Although this occurred only once in the metadata set, it appeared more frequently in the full-text set.
In the full-text dataset, the most common example involved “Tema: Relationer (hbtqi)” (“Theme: Relationships (LGBTQI)”), which appeared six times in Gemini outputs, three times in DeepSeek outputs, and once in ChatGPT outputs. In all observed cases, the underlying QLIT term “Relationer (hbtqi)” (“Relationships (LGBTQI)”) had already been generated automatically, meaning that both the standard QLIT term and the erroneous form prefixed with “Tema:” (“Theme:”) appeared in the same output. A similar pattern was observed for “Tema: Identiteter och praktiker (hbtqi)” (“Theme: Identities and Practices (LGBTQI)”), which was generated three times by DeepSeek and twice by Gemini. Unlike the previous example, however, the prefixed form was not generated alongside the standard QLIT term “Identiteter och praktiker (hbtqi)” (“Identities and Practices (LGBTQI)”) in any of the cases. Finally, “Tema: Diskriminering, hat och våld (hbtqi)” (“Theme: Discrimination, Hate and Violence (LGBTQI)”) occurred twice in DeepSeek outputs. In one case, it appeared alongside the original QLIT term “Diskriminering, hat och våld (hbtqi)” (“Discrimination, Hate and Violence (LGBTQI)”), whereas in the other case, only the erroneous prefixed form was generated.
In addition to being from outside of QLIT and often conceptual duplicates, these cases also create practical difficulties for evaluations relying exclusively on existing gold standards and exact string matching. In a semi-automated indexing workflow such cases would be less problematic because information professionals would make the final decision regarding whether the generated term should be accepted, modified, or rejected. For the purposes of the present study, these outputs were evaluated as equivalent to the original QLIT terms, meaning that the “Tema:” prefix was ignored during analysis and the terms were evaluated as if they had been generated without the additional string.
Results
Metadata
The earlier study (Golub and Ihrmark, 2026) showed that AI tools perform with an average F1 of 44% on metadata. The highest precision is achieved by ChatGPT (39%), closely followed by Gemini (37%) which, just like Claude, also yields the highest recall (94%). Gemini attains the highest F1 score (50%), narrowly surpassing ChatGPT (49%).
Table 1 below shows the detailed qualitative assessment breakdown by the category code. The top row lists the AI tools and in parentheses the total number of generated terms by each tool. Claude and Deepseek created the highest number of terms for all the metadata records in the sample (n = 156), followed by Gemini (n = 110) and ChatGPT with least terms (96). The results show that across all models, the largest category on average was fully incorrect terms (37.5%), indicating that a considerable proportion of generated terms could not be considered valid at all. DeepSeek (43%) and Claude (41.7%) produced the highest proportions of fully incorrect terms, while Gemini generated the lowest proportion (30%). At the same time, the proportion of fully correct terms averaged 30% across all models, with ChatGPT (36.5%) and Gemini (35.5%) performing best in this category, followed by Claude (25%) and DeepSeek (23.1%).
Qualitative inspection showed that AI went for other text in the metadata record to often wrongly generate QLIT terms. For example, for E. Flygare-Carlén's Jungfrutornet (1848, Volumes 1–2) the description mentions that one of the recurring characters is the gender-changing “Mother Styrman”, which resulted in some systems assigning “Sex Changes” (Könsväxlingar). However, the QLIT term's definition states that it is intended for depictions in which individuals physically change sex through magical or otherwise unexplained means. The cross-dressing character in the novel does not undergo any such transformation. Another example of using the Queerlit description to assign an incorrect term is “Lesbian Studies” (Lesbiska studier), which was generated several times for different works where the description mentions how the author was cited in lesbian literature while none of the works of fiction are about lesbian studies. Yet another one is “Historical Terms (LGBTQI)” (Historiska termer (hbtqi)) which is intended for historical words used to describe LGBTQ + identities and experiences, not simply for works set in the past as this entire document collection; an example where this term was added was when the Genre facet depicted the work as historical fiction.
A substantial number of generated terms were not entirely incorrect but instead reflected semantic relationships with the assigned QLIT terms. Broader terms accounted for an average of 12.9% of outputs, with relatively little variation between models (11.5–13.6%). This shows that the models often did not follow the subject indexing principle of assigning the most specific term, even though they had been instructed to assume the role of a subject indexer. However, further inspection also showed that the broader term was frequently assigned alongside the more specific correct term rather than replacing it. This introduces conceptual redundancy of the assigned index terms.
From an information retrieval perspective, assigning both broader and narrower concepts introduces redundancy and reduces indexing precision by representing the same phenomenon at multiple levels of specificity. The results showed that some AI tools assigned the broader QLIT term “Women” (Kvinnor) to works already correctly indexed with “Lesbians” (Lesbiska). Although this may appear semantically reasonable, “Lesbians” is a more specific concept in the QLIT hierarchy; assigning the broader term “Women” is incorrect, as not all women are lesbians. Similarly, some systems assigned both “Romantic Friendship” (Svärmisk vänskap) and its broader term “Friendship” (Vänskap) to the same work. In Ellen Idström's Tvillingsystrarna (1893), all four systems generated both the correct term “Masculinizing Cross-Dressing” (Maskuliniserande crossdressing) and its broader term “Cross-Dressing” (Crossdressing), while some also assigned “Girls” (Flickor) despite the more specific term “Tomboys” (Pojkflickor) already being present.
In contrast, narrower terms were rare (1.2%), suggesting that the models were generally more likely to overgeneralize concepts than to select terms that were too specific. When narrower terms did occur, however, they reflected genuine conceptual misunderstandings rather than more precise interpretations of the work. For example, one AI tool assigned the QLIT term “Lesbians” (Lesbiska), a narrower term of the correct QLIT term “Women” (Kvinnor), to a work about women in general. This assignment incorrectly introduced a specific sexual orientation that was not supported by the content.
The category of related terms also remained relatively small (2.7% on average) but differed across systems. ChatGPT did not produce any terms classified as related, while DeepSeek generated the highest proportion (3.9%). From a subject-indexing perspective, semantically related concepts are not to be assigned if a more appropriate and precise controlled vocabulary term exists. Similar to the issue observed with broader terms, assigning related concepts instead of the most precise term may introduce ambiguity and conceptual overlap, creating retrieval failures.
Terms classified as incorrect according to QLIT definitions or indexing policy constituted a relatively stable category across all tools (10.4% on average), suggesting that some generated concepts appeared semantically plausible but conflicted with the rules underlying Queerlit indexing or with QLIT definitions. Adherence to indexing policy and thesaurus definitions is essential for ensuring consistency in subject assignment and maintaining the intended use of the controlled vocabulary. Furthermore, if thesaurus definitions are not followed to identify the intended meaning and appropriate use of a term, works will be misindexed.
An interesting result concerns terms categorized as not assigned by Queerlit but partially correct. Although these represented a relatively small category (5.4% on average), they appeared across all models, with DeepSeek (7.7%) and Claude (7.1%) producing the highest proportions. This suggests that the models occasionally generated subject terms judged to be potentially valid despite not appearing in the original Queerlit assignments. Such cases also illustrate the limitations of relying exclusively on pre-existing gold standards for evaluation. However, their potential value must be weighed against the much larger number of fully incorrect, overly broad, conceptually duplicate, or policy-inconsistent terms produced by the systems. In practice, this means that indexers would need to review a substantial volume of unsuitable suggestions to identify a comparatively small number of potentially useful ones, limiting the efficiency gains that AI-assisted indexing might otherwise offer.
The findings imply that AI tools have limited usefulness for subject indexing of the Queerlit collection based on metadata. Although the models occasionally generated valid additional concepts and achieved an average of 30% fully correct terms, the largest category across all systems consisted of fully incorrect terms (37.5%), alongside many broader, related, or policy-inconsistent terms. These results are particularly weak considering that all relevant subject terms were already present within the metadata records provided to the models, meaning that the task largely involved identifying and selecting existing terms rather than inferring them from full literary texts. Furthermore, there were no terms that any of the four tools generated and that were coded as “not assigned by Queerlit but fully correct”, suggesting that none of the models identified entirely new subject terms that were judged as valid additions beyond the existing indexing. Rather than providing new useful suggestions at a significant scale, the models most often appeared to use existing index terms to add also other semantically related, broader, or otherwise adjacent terms instead of selecting the most appropriate and most specific controlled vocabulary terms according to QLIT definitions and indexing policies.
Fulltext
The earlier study (Golub and Ihrmark, 2026) found that the four AI tools performed substantially worse on full texts than on metadata records, achieving an average F1 score of only 12%, precision of 8% and recall of 36%. This result is not unexpected, as the relevant QLIT subject terms are explicitly present in the metadata records whereas for full text they need to be entirely inferred. Claude and DeepSeek performed best overall, both achieving an F1 score of 14%, although Claude obtained a slightly higher recall (45%) than DeepSeek (42%). These models were more successful at identifying accurate subject terms, but at the cost of generating many additional terms that were not present in the gold standard, resulting in low precision (9% for both models). Gemini showed slightly worse performance, with an F1 score of 11%, precision of 7%, and recall of 30%. ChatGPT performed least successfully according to all three measures, achieving the lowest precision (6%), recall (27%), and F1 score (9%).
Table 2 below shows the detailed evaluation assessment breakdown for the subject terms generated from the full texts. The top row shows the AI tools and, in parentheses, the total number of generated terms. Claude and DeepSeek generated the highest number of terms (n = 178 each), followed by ChatGPT (n = 123) and Gemini (n = 111). The results show that, across all models, the largest category by far was fully incorrect terms (56.1% on average), indicating that more than half of all generated terms are entirely a conceptual miss. ChatGPT produced the highest proportion of fully incorrect terms (64.2%), followed by DeepSeek (56.7%), Gemini (54.1%), and Claude (49.4%). At the same time, the proportion of fully correct terms was very low, averaging only 8.3% across all models. Claude achieved the highest proportion of fully correct terms (10.1%), followed by DeepSeek (8.4%), Gemini (8.1%), and ChatGPT (6.5%). Compared with the metadata results, whereby the correct QLIT subject terms were explicitly present in the metadata provided to the models (yet performance was still only moderate), the performance on full texts was substantially worse, reflecting the considerable challenge of identifying and interpreting LGBTQ + themes from literary works where subject terms are rarely stated directly and instead must be inferred from narrative context.
At the same time, poetry also proved considerably more challenging than novels, both for subject indexing and for the evaluation of AI-generated subject terms, reflecting difficulties that have long been recognized in human indexing practice. Because poetic language often relies on symbolism and metaphor, it can be difficult not only for AI systems but also for human experts to determine which themes are genuinely present and which interpretations are justified. Novels, by contrast, were generally easier to evaluate because their themes tend to be expressed more explicitly. This challenge was particularly evident in the works of Edith Södergran, where recurring references to spirituality, Greek mythology, and nature made it difficult to determine whether concepts such as spirituality represented central themes or were merely metaphorical elements. Similar issues arose in Vilhelm Ekelund's poetry collection Hafvets stjärna, where AI systems identified references to boys (gossar), yet it remained unclear whether these references should be interpreted literally or metaphorically.
Several examples also highlighted the AI systems' difficulty in distinguishing between general themes and LGBTQ + specific themes. In Ellen Idström's Tvillingsystrarna, some AI-generated terms were judged thematically relevant, yet they did not represent LGBTQ + specific concepts. The AI tools had been instructed to assign only QLIT terms and were not given the option of selecting terms from SAO (Swedish Subject Headings), used in Queerlit for non-LGBTQ + themes. The AI systems sometimes assigned LGBTQ + concepts where a general subject heading would have been accurate instead; for example, “Grief (lgbtqi)” (Sorg (hbtqi)) was assigned instead of the general SAO concept “Grief” (Sorg).
Another important finding concerns terms categorized as incorrect according to QLIT definitions or subject-indexing policy, which made up the second largest category after fully incorrect terms, accounting for 13.9% of generated terms. The relatively stable proportions across models indicate that many outputs appeared relevant; however, as discussed above, consistent application of thesaurus terms and indexing guidelines is necessary to ensure accurate representation of works and reliable retrieval, as otherwise records may receive incorrect or inconsistent subject terms.
About 10.5% of generated terms were not entirely incorrect but were semantically related to the intended concepts: broader terms. Claude (12.9%) and Gemini (12.6%) produced the highest proportions of broader terms. As with metadata, this suggests that the models frequently recognized the general thematic area of a work but failed to select the most specific available concept, contrary to the Queerlit indexing principle of assigning the narrowest appropriate term. However, as with metadata, broader terms were often assigned alongside the more specific correct terms, creating redundancy and reducing the precision of the indexing output. The categories narrower terms (0.7%) and related terms (3.2%) remained relatively small. Related terms appeared somewhat more frequently, particularly for Claude (4.5%) and Gemini (4.5%).
The category “not assigned by Queerlit but partially correct” represented 7.4% of outputs on average, with DeepSeek producing the highest proportion (11.2%), followed by Claude (7.3%), ChatGPT (5.7%) and Gemini (5.4%). These terms were judged to be somewhat relevant but not fully accurate. Nevertheless, they may offer additional subject access points, particularly in fiction, where themes are often implicit and open to interpretation. However, their usefulness depends on human review, as loosely related concepts may also introduce noise and reduce retrieval precision.
As with the metadata subset, none of the models generated any terms categorized as “not assigned by Queerlit but fully correct” (0%), showing that AI primarily generates a mixture of incorrect, broader, partially correct, and related concepts, requiring considerable expert review. Taken together, the findings indicate that for this collection of historical LGBTQ + fiction, current general-purpose LLMs struggle not only with identifying the correct subject terms but also with applying the controlled vocabulary consistently and according to established indexing principles. At the same time, the evaluation revealed cases where AI tools identified non-LGBTQ + themes that should be represented via the Swedish subject headings (Svenska ämnesord – SAO, a vocabulary used in parallel with QLIT in Queerlit for general topics). For example, in Vampyrer by Algot Sandberg, the models suggested partially correct concepts such as jealousy and domestic violence, which had not been assigned earlier using SAO. This suggests that AI systems may have some potential for identifying additional subject access points for general themes.
Qualitative analysis further indicated that even when AI systems identified themes similar to those reflected in the existing indexing, they often failed to select the exact QLIT concept assigned by human indexers. For example, in Karin Boye's För trädets skull (1935), none of the four AI tools identified the assigned QLIT terms “Infatuation” (Förälskelse (hbtqi)) and “Unrequited Love” (Obesvarad kärlek (hbtqi)), instead generating related concepts such as “Desire” (Begär (hbtqi)) and “Intimacy”. According to the subject librarian, distinctions between concepts such as infatuation, desire, intimacy, longing, and unrequited love are often difficult even for human indexers, particularly in poetry where themes are expressed indirectly and remain open to interpretation. Similar observations were made for Vilhelm Ekelund's Hafvets stjärna, where themes of longing could be identified in the poems, yet not necessarily in a form that justified assigning “Unrequited Love” (Obesvarad kärlek (hbtqi)). These examples suggest that AI systems are often capable of identifying the broader thematic domain of a work but have greater difficulty selecting the specific controlled-vocabulary concept that best matches the intended indexing interpretation.
Another phenomenon occasionally observed was the tendency of AI systems to apply contemporary LGBTQ + concepts and identity categories to historical literary works. In several cases, the generated terms appeared plausible from a modern perspective but were difficult to justify either from the text itself or according to QLIT indexing policy. For example, in Ellen Idström's Tvillingsystrarna, terms such as “Queer” were regarded by the subject librarian as anachronistic in relation to the historical context of the work. Similarly, in Emilie Flygare-Carlén's Jungfrutornet, the presence of a cross-dressing character prompted the generation of terms relating to queer identities, transgender culture, and gender transgression, despite the fact that the narrative does not concern gender identity. These examples suggest that AI systems tend to interpret historical texts through contemporary conceptual frameworks, treating characteristics associated with modern LGBTQ + identities as sufficient evidence for assigning LGBTQ + -specific subject terms.
Interpretation versus textual evidence
A central finding was the difficulty of distinguishing between themes explicitly supported by a literary work and themes that emerge primarily through interpretation. This challenge was particularly pronounced in poetry, where themes are often indirect, symbolic, and open to multiple interpretations. The discussions between the non-subject experts and the Queerlit subject librarian revealed several cases in which AI-generated terms appeared plausible but relied on interpretations that extended beyond the textual evidence available in the work itself.
This finding aligns with longstanding debates in knowledge organization regarding the subject indexing of fiction. Unlike non-fiction texts, literary works often invite multiple interpretations, raising questions about the extent to which indexers should infer themes and meanings that are not explicitly stated. Some scholars have argued that interpretive indexing risks imposing the indexer's perspective on literary works and potentially constraining readers' own interpretations, while others contend that thematic indexing is necessary to facilitate discovery and access (Bell, 1992; Nielsen, 1997). One potentially useful conceptual parallel can be found in approaches to subject access for visual works of art, where distinctions may be made between different levels of subject analysis, such as description, identification, and interpretation. A comparable distinction could potentially help clarify whether an AI-generated term describes what is explicitly present in a literary text, identifies a recognizable subject or theme, or reflects a higher-level interpretation of the work. Developing or evaluating such a framework is beyond the scope of the present study, but it may provide a useful direction for future research on AI-assisted subject indexing of literary works.
Karin Boye's För trädets skull (1935) provided a particularly clear example. One non-subject expert considered terms such as “Women” (Kvinnor), “Lesbians” (Lesbiska), and “Love Between Women” (Kärlek mellan kvinnor) potentially justifiable because the Queerlit description notes that Boye has been canonized within lesbian literary history and that themes of love between women can often be read into her poetry. However, the subject librarian pointed out that explicit references to gender, sexuality, same-sex love, or women are absent from the poems themselves. Consequently, assigning such terms would require indexing an interpretation rather than what is directly expressed in the text.
The same tension between textual evidence and interpretation emerged in Anne C. Branting's Lena (1893). Although the novel contains depictions of intense affection between women, the subject librarian argued that assigning terms such as “Desire (lgbtqi)” (Begär (hbtqi)), “Lesbians” (Lesbiska), “Love Between Women” (Kärlek mellan kvinnor), or “Lesbian Relationships” (Lesbiska relationer) overstated what is actually represented in the work. While such terms may appear plausible, the librarian stressed that the novel primarily depicts romantic friendship rather than explicit lesbian desire or identity.
Similar interpretive challenges arose in Vilhelm Ekelund's Hafvets stjärna (1906), where “Desire (LGBTQI)” (Begär (hbtqi)) was considered a possible but highly interpretive assignment. The subject librarian acknowledged that another indexer might reasonably apply the term but would not personally assign it because the relevant themes remain implicit rather than explicit.
As mentioned above, AI systems also sometimes relied on cues present in metadata descriptions. For example, the metadata description for Karin Boye's Härdarna states that “in her poetry, the themes are never explicit.” This wording may have contributed to Claude generating terms such as “Secret Symbolism (lgbtqi)” (Hemlighetssymbolik (hbtqi)) and “Symbolism (lgbtqi)” (Symbolik (hbtqi)). Although the subject librarian regarded these assignments as interpretive rather than textually grounded, she nevertheless noted that she might have accepted some of them had they been proposed by a human indexer, illustrating the existence of grey areas between indexing and interpretation.
A similar pattern was observed in Anne C. Branting's Lena. The Queerlit description contains language referring to affection and desire between women, and Claude appeared to draw directly on this wording when generating “Desire (lgbtqi)” (Begär (hbtqi)). However, according to the subject librarian, the description uses language that may suggest themes not fully supported by the indexing. As a result, AI systems may generate terms based on the wording of the description rather than on the more cautious interpretation reflected in the assigned subject terms.
Conclusion
This study evaluated the ability of four general-purpose AI tools to assign QLIT subject terms to historical LGBTQ + fiction using both metadata records and full texts. The results indicate that general-purpose AI tools performed only moderately when assigning QLIT subject terms from metadata records and poorly when assigning terms from full texts. For metadata, the models achieved an average F1 score of 44%, with Gemini performing best (50%), followed closely by ChatGPT (49%). However, fully incorrect terms constituted the largest category across all systems (37.5%), exceeding the proportion of fully correct terms (30%). Performance declined substantially for full texts, where the average F1 score fell to 12%. More than half of all generated terms were fully incorrect (56.1%), while only 8.3% were fully correct. Across both metadata and full-text datasets, broader terms accounted for approximately one-tenth of all outputs (12.9 and 10.5%, respectively). Notably, none of the four tools generated any terms that were judged both fully correct and absent from the existing Queerlit indexing, suggesting limited ability to contribute entirely new and valid subject terms.
The qualitative evaluation revealed several recurring patterns that help explain the relatively weak quantitative performance. AI systems frequently relied on semantically related concepts rather than thesaurus definitions, indexing policies, or the principle of assigning only the most specific available term. They often generated broader, redundant, or policy-inconsistent concepts, and struggled to distinguish between general themes and LGBTQ + specific themes. The analysis also demonstrated that the systems tended to draw heavily on wording found in metadata descriptions and, in some cases, interpreted historical literary works through contemporary LGBTQ + conceptual frameworks, resulting in anachronistic assignments such as “Queer” or transgender-related concepts. Poetry proved particularly challenging because themes were often implicit, and open to multiple interpretations.
At the same time, a small proportion of generated terms (5.4% for metadata and 7.4% for full texts) were judged partially correct despite not appearing in the original indexing, suggesting some potential for AI-assisted identification of additional subject access points. Nevertheless, these potentially useful suggestions were greatly outnumbered by incorrect, overly broad, or policy-inconsistent outputs, indicating that effective use of AI for subject indexing would still require substantial expert review and careful application of controlled-vocabulary principles.
The study further demonstrated the limitations of relying exclusively on existing metadata as a gold standard for evaluating automated subject indexing. Many AI-generated terms that failed to match the assigned QLIT terms were not necessarily unrelated to the works but reflected alternative interpretations, broader thematic associations, or different understandings of literary aboutness. At the same time, qualitative evaluation revealed numerous cases in which generated terms appeared semantically reasonable yet conflicted with indexing policies, thesaurus definitions, or the intended scope of QLIT concepts. These findings confirm that evaluating AI-generated subject indexing requires more than measuring agreement with existing metadata. Effective evaluation must also consider thesaurus structure and rules, indexing policy, and the distinction between literary interpretation and evidence-based subject assignment.
From a queer knowledge organization perspective, the results highlight the complexity of representing LGBTQ + themes in historical literary works. Many indexing decisions depend on nuanced judgments regarding identity, cultural and historical context that cannot be reduced to simple keyword matching or semantic similarity. The study also illustrates how contemporary LGBTQ + concepts may be projected onto historical texts in ways that are not supported by the works themselves. The findings therefore reinforce the continuing importance of human expertise in queer indexing and suggest that effective knowledge organization requires not only recognition of relevant themes but also an understanding of context, the controlled vocabulary, and professional indexing practice.
For libraries considering the use of generative AI in subject indexing workflows, the results suggest a more cautious assessment of current capabilities. For this particular collection of historical LGBTQ + fiction, the AI systems generated relatively few potentially useful additional terms compared with the much larger number of incorrect, overly broad, or policy-inconsistent suggestions. Consequently, the effort required to review and filter generated terms may outweigh any efficiency gains achieved through automation. Nevertheless, the findings also indicate that AI systems may have value as tools for generating candidate concepts, identifying possible alternative access points, or supporting exploratory analysis of literary themes. The subject librarian noted that AI tools could potentially be useful in future indexing workflows, for example by processing reviews or other supplementary texts and suggesting candidate terms for consideration. However, such suggestions would still require expert evaluation, particularly because effective indexing depends not only on recognizing themes but also on distinguishing between concepts belonging to different vocabularies, such as QLIT and SAO, and understanding the specific indexing policies governing their use.
The study also points toward several directions for future research. First, future studies should investigate larger and more diverse collections of fiction in order to determine whether the patterns observed here generalize beyond historical LGBTQ + literature. In particular, the historical nature of the materials examined in this study may have presented additional challenges for AI systems, given changes over time in LGBTQ + terminology, identities, and conceptual frameworks. A direct comparison between historical and contemporary LGBTQ + fiction would therefore be valuable for determining whether AI systems perform more effectively on contemporary works and for identifying the extent to which historical and cultural context contributes to indexing difficulties. Such a comparison may, however, be challenging to conduct in practice because contemporary literary works are generally protected by copyright and are less likely to be freely available in full text for research purposes.
Second, comparison with broader controlled vocabularies such as SAO may help clarify whether AI systems perform more effectively when assigning general thematic concepts than when working with highly specialized LGBTQ + focused vocabularies. Third, research should explore ensemble approaches (similar to Kluge and Kähler, 2025) or hybrid approaches that combine generative AI with traditional machine-learning methods (Suominen et al., 2025; Tanácsi and Micsik, 2026), as recent studies suggest that such approaches are currently the most promising. Hybrid approaches may also help address practical limitations associated with prompt-based indexing, particularly token constraints that restrict the amount of thesaurus or indexing policy information that can be provided to a model when working with longer fulltext materials. Fourth, future evaluation should incorporate retrieval-based and workflow-based contexts in addition to conventional precision, recall, and F1 scores, allowing researchers to comprehensively assess whether AI-generated terms improve actual subject access for users (Golub et al., 2016). Future research could also examine whether approaches developed for subject access to other creative works, particularly visual art, could inform the indexing of fiction. Distinguishing between description, identification, and interpretation may provide a useful framework for evaluating when AI-generated terms are grounded in the work itself and when they instead represent interpretive readings.
Overall, the findings suggest that successful subject indexing of fiction depends not only on recognizing themes but also on understanding historical context, thesaurus definitions, indexing policy, and the distinction between textual evidence and literary interpretation. While generative AI shows some capacity to identify thematic content, it currently lacks the contextual and domain-specific understanding required for reliable indexing of historical LGBTQ + fiction. Human expertise therefore remains essential, both for ensuring accurate subject representation and for maintaining the integrity of queer knowledge organization systems.

