This paper aims to present a case study of integrating generative artificial intelligence (AI) technologies into archival metadata creation in an academic library, providing an example of improving descriptive metadata for digitized photograph collections to enhance discoverability and accessibility.
This study tested various vision-enabled and multimodal large language models for generating Metadata Object Description Schema-compliant metadata from digitized photographic prints. Models were evaluated on visual interpretation and metadata output quality. Following a pilot test that validated this approach, Anthropic’s Claude Sonnet 4 model was selected for production implementation, creating descriptive metadata for 2,263 digitized photos. A Python-based process helped integrate this into existing workflows for digital collections.
The process provided substantial improvements over existing minimal descriptive metadata. Generated subject terms mapped successfully to the Faceted Application of Subject Terminology vocabulary in 64% of cases, demonstrating compatibility with established standards. Project team discussions led to the implementation of user transparency notices about AI involvement in metadata creation. Post-project analysis revealed this approach to be a cost-effective method of metadata enhancement.
This case study provides actionable guidance for cultural heritage institutions evaluating AI integration into their workflows and can be adapted for different institutional contexts and collection sizes.
While AI applications in libraries are expanding, this study documents a practical approach to integrating generative AI into archival metadata creation. The workflows, quality controls, cost analysis and transparency measures offer a roadmap for those considering adopting these technologies at scale.
Introduction
Academic libraries have significantly expanded digitization efforts over the past two decades (Jaillant et al., 2022). While this has led to unprecedented levels of access to historical materials, a related challenge has emerged – the ability to generate comprehensive descriptive metadata has not moved at the same rate (Colavizza et al., 2021; Zavalin and Zavalina, 2025). One area where this disparity is particularly pronounced is in image-based collections. Whereas OCR technologies have long been used to increase discoverability for text-based collections, visual materials have traditionally required human analysis and interpretation to create meaningful textual descriptions that can be used to increase discoverability and accessibility. Mass digitization effectively opens large “hidden” collections, providing access to a variety of materials across the globe (CLIR, 2014). However, digitization without metadata has led to a different concern, where the digital collections themselves become inaccessible or undiscoverable, whether through opaque access in online repositories or even abandoned in local or shared drives (Hawkins, 2022; Kuźma and Mościcka, 2020; Lampert, 2017; Tebeau, 2021).
Recent advances in artificial intelligence, particularly vision-language models (VLMs) and large language models (LLMs), offer new possibilities for addressing this gap through automated metadata generation and enhancement (Ali et al., 2024; Jaillant and Caputo, 2022). Unlike earlier computer vision approaches that could identify and classify objects, contemporary VLMs can interpret visual scenes, extract text from images and generate meaningful descriptions about both content and context. In the field of digitized archival materials, these tools can enable automatic generation of descriptive metadata (e.g. title, description, subjects) based on image analysis for single or multiple photographs, as well as transcribing visible text such as handwritten notes or captions (Zhang et al., 2023). In this context, the ability of modern LLMs to create structured output becomes particularly valuable for metadata workflows, as they can provide results in formats that allow for integration into existing processes and systems. Working together, contemporary multimodal AI systems can produce metadata that can be used by cultural heritage institutions to facilitate large-scale descriptive cataloging (Schellnack-Kelly and Modiba, 2024; Spennemann, 2023). Early implementations have demonstrated both the promise and challenges of automated metadata generation for archival photographs, with success rates varying based on image quality, subject matter complexity and institutional requirements (Carter et al., 2022; Chen et al., 2024; Cocciolo, 2025; Thammastitkul and Petsuwan, 2024).
Research examining Metadata Object Description Schema (MODS) metadata quality in digital repositories has documented challenges with accuracy, consistency and completeness – the primary criteria used for metadata evaluation (Park and Tosaka, 2010). Park and Maszaros (2009) found inconsistent application of MODS elements across institutional repositories, attributing variations to local interpretations of metadata standards. Recent evaluations of AI-generated MODS metadata indicate that automated tools do not yet meet quality thresholds, with accuracy remaining particularly problematic (Zavalin and Zavalina, 2025). These findings underscore the need for practical evaluation approaches when implementing AI-assisted workflows in specific institutional contexts.
This article presents a case study on the integration of AI generated metadata into descriptive metadata records for a medium-sized digitized archival photograph collection from the Kenneth Spencer Research Library (KSRL) at the University of Kansas (KU). It will demonstrate a practical workflow, quality assessment methods and the means of integrating AI tools into existing workflows and practices while maintaining transparency for the user about limitations of the metadata. With regard to metadata quality assessment, this exploratory case study relied on a practical evaluation approach suited for this project. Bruce and Hillmann’s (2004) seven-dimensional model and Stvilia et al.’s (2007) comprehensive taxonomy provide robust metadata quality evaluation frameworks; however, considering the emerging nature of AI-generated metadata for archival materials and this project’s focus on the viability of this approach for our institution, we prioritized practical measures of accuracy, completeness and usability. These metrics will ultimately help inform decisions about implementation and workflow integration, though more comprehensive assessment using established frameworks would be valuable in future work.
Existing processes and limitations
To address the challenges inherent in creating digital objects without sufficient metadata, the KU Libraries adopted a general policy in 2017 to not undertake any large-scale digitization projects without a preexisting descriptive record (i.e. archival finding aid or library catalog record) detailed enough to create at least minimum metadata records for individual items (e.g. folder-level, item-level, record group, etc.). In this approach, description comes before digitization. This serves the dual purpose of ensuring that all digitized content can be minimally described and we can create links between each digitized item and its source record in the finding aid or catalog, which serves a secondary goal of helping to eliminate orphaned digital files.
For most large-scale digitization projects at KU, metadata records are created using Python scripts that extract the most detailed level of description available from the Encoded Archival Description (EAD) finding aid in ArchivesSpace, typically at the folder or item level. Depending on the level of processing that the physical collection has received, this can result in very detailed or rather general descriptions for the digital objects.
During digitization, staff record each item’s location information (e.g. box and folder numbers), which maps the item to its description in the finding aid. This location data, along with the unique identifier from the ArchivesSpace container record, creates persistent links between the digital object, its metadata record (in MODS format) and the original finding aid. When subject terms are present in the finding aid, they are mapped to the Faceted Application of Subject Terminology (FAST)-controlled vocabulary, and the terms and their URIs are included in the item’s metadata.
Imaging is done in the KSRL digitization lab, using a DT Versa copy stand using a Phase One iXH 150 MP camera with a 72 mm MKII lens. All photographs are digitized in accordance with the Federal Agencies Digital Guidelines Initiative (FADGI) guidelines to meet a minimum three-star rating, depending on the needs of the specific project, and saved as LZW compressed TIF files (FADGI, 2023). Following digitization, the TIF files are stored on a shared network drive until they are processed into our digital repository. For most projects, photographic prints are imaged front-side only, unless there are markings on the back (e.g. handwritten descriptions, photographer’s notes or other text/symbols regardless of content), in which case, both sides are captured. For this case study, “item” refers to the digital files created from either imaging approach.
The project: Watson Library photographic prints
To coincide with the 100th anniversary of the Watson Library, a project team (including members from Digital Initiatives, curators, administrators, processors and conservators) assembled to undertake a medium-sized digitization project of materials from the University Archives documenting the history of Watson Library. As part of this, approximately 3,500 archival items (photographs, documents and scrapbooks) relating to various aspects of the library’s history and operations over the years were identified and digitized. This study will focus on the 2,263 photographic prints included in the set of materials.
The items digitized for this project were reflected in the finding aid only by their respective Record group and subgroup entries (note: this level of finding aid description is often KSRL’s standard practice for university records of this type), so the most detailed descriptions available were broad categories, such as “Watson Library,” “Music Library,” “Reference,” etc. In all, two record groups and 28 subgroups were included in the digitized items. In most cases, photos were organized into folders by year or decade, with the appropriate date listed on the physical folder. During digitization, each of these values was captured in the filename of individual items. To illustrate how this works in practice, consider Plate 1.
The filename for Plate 1 – ksrl_ua_32_37_up_1993_004a – can be parsed as follows:
Item source
“ksrl”: “Kenneth Spencer Research Library”
“ua”: “University Archives”
Call number
“32”: Record group 32 – “University of Kansas Libraries records”
“37”: Record subgroup 37 – “Special Collections”
“up”: “University produced photographs”
Folder level
“1993”: date
Item level
“004”: sequential number within the folder
“a”: Optional. If present, an “a” indicates the front of a photograph, and a “b” indicates the back. If there is no “a” or “b,” only the front of the photograph was captured.
Given the limited detail provided in the finding aid, the descriptive metadata available for the item in Plate 1 is quite minimal, as shown below. This example includes a default institutional subject heading applied to all items in this project. (See Appendix 3 for the complete project XML template, including administrative and technical metadata.)
While this record is far better than having no metadata for this item, in that it provides some basic metadata for the image, there are multiple problems with this approach. First, the record provides only contextual information – nothing about the content that might be found by a keyword search. A user searching for terms like “students”, “class,” “books” or other keywords would never discover the above relevant image because the metadata contains no visual content descriptors. Table 2 provides an example of how AI can help capture these missing details to enhance the metadata record.
This example also illustrates issues of accessibility and discoverability. Accessibility in this context involves the libraries’ ability to provide equitable access for all users. With this absence of meaningful descriptive metadata, a patron using a screen reader would have no tangible information about the image itself. Discoverability refers to the pathways by which a user might come across a given item. In this example, the collection is likely to have several hundred items from the same record group or subgroup, resulting in dozens of pages of search results populated with nearly identical records that provide no basis for distinguishing between individual items (e.g. title = “Watson Library: Special Collections”). This creates a practical barrier to discovery, as users must manually examine each result to determine relevance and are likely to abandon their search before finding the specific materials they need (Abualsaud, 2020). The lack of unique or distinctive descriptive metadata essentially renders individual items invisible within the larger collection, undermining the purpose of making materials digitally accessible.
To address these challenges of access, the author researched ways that generative AI and computer vision models could be applied to automated image description. Building on foundational experience with computational text analysis and machine learning for metadata enhancement (Wolfe, 2019; Wolfe, 2023), the author had experimented capturing AI-generated metadata for individual items but had not yet applied these methods a larger-scale project. The combination of minimal existing metadata for the Watson Library materials and institutional support for AI experimentation made this collection an excellent candidate for applying these tools at scale while developing replicable workflows for similar collections. This case study describes this project in three phases: model evaluation and selection, pilot testing and production deployment.
Implementation
Phase 1: Model evaluation and selection
Beginning in early 2025, the author tested a wide range of AI models for this task, including of commercial Web-based products (e.g. Anthropic’s Claude models, GPT-4o) and open-source multimodal/VLMs (e.g. Qwen, Gemma) running on a local Mac with an M1 silicon chip. In all, nine models were tested using a five-image set selected to include a range of image composition and content (see Appendix 2). During this phase, all testing and review was done solely by the author. A detailed system prompt was created to define the query for the item and to identify the perspective the model should take in creating the metadata (see Appendix 1). To ensure consistency in testing, the same prompt was used for each model.
The models were given the task of analyzing individual photographs and generating values for the following metadata elements in accordance with the MODS metadata standard: title, description, subject(s), name(s) and date. The name and date values are optional and should be included only when clearly identifiable from the item itself, typically in writing on the back of the photo. In the earliest tests, additional metadata was requested, but an increase in the number of requested values directly correlated to a higher rate of errors and malformed outputs from the LLM. Thus, the scope was narrowed iteratively based on observed model performance. Models were asked to output results in a structured JSON format to allow for computational processing.
The results were assessed using criteria developed specifically for this testing phase, focusing on the model’s visual interpretation of the image – including thoroughness of visual content identification, accuracy of perception and text transcription – and metadata output quality – including natural language use, adherence to requested standards and proper JSON formatting. The review process was based on the author’s observation and personal judgment rather than controlled statistical testing methods, reflecting the exploratory nature of this project and the need to rapidly iterate on model parameters and output requirements. Model success was rated comparatively across the tested systems, evaluating both performance against these criteria and relative performance among models. See Appendix 2 for a full list of models tested and evaluation criteria used.
Common errors ranged from fundamental visual misinterpretations to technical formatting problems. For instance, models might describe a printing press as a microfilm reader, assume that all individuals in a photo are “library staff,” incorrectly transcribe visible text or fail to identify key visual components within the image. These content errors were often compounded by technical issues such as poorly formed JSON, extraneous explanatory/unstructured text and inclusion of unrequested metadata values. Hallucinations (i.e. textual descriptions of content not present in the actual image) were rare and were primarily identified in relation to extracted text – for example, transcribing the printed name “Professor Roberts” as “Professor Joseph Roberts.”
Although the models produced widely ranging levels of success across individual evaluation areas, the highest weighted criterion (and in many cases, hardest to achieve) was the model’s ability to return results in the requested structured format. Without consistently formatted results, even the most comprehensive metadata cannot be easily processed or integrated into existing computational workflows.
The Anthropic cloud-based model Claude-3-7-sonnet-20250219 (Sonnet 3) consistently produced the highest-ranking results in each of these areas. Although this phase was undeniably an uncontrolled and less-than-ideal testing environment, the project team agreed that the work completed to this point was sufficiently promising to allow moving forward with the next phase, which would involve a larger number of reviewers from across the library.
Phase 2: Pilot testing
For the pilot phase, a set of 100 images (50 front-only, 50 front-and-back) were randomly selected proportionally across all 28 subgroups. To improve on results obtained in the preliminary testing, the system prompt was revised to provide or clarify guidelines on image handling, metadata formatting and expected output (see Appendix 1 for details on the prompt development). A process was set up in a Jupyter notebook to pass individual queries, including the image(s) and text prompt, to Anthropic’s API using the Anthropic Python client library. This stage involved several iterations, as this larger query set revealed new issues with the responses. The most common issues were inconsistently formatted responses from the Sonnet 3 model, such as invalid JSON or extraneous text. Once the issues were addressed in the code, the output was processed, with the JSON parsed into a pandas dataframe and exported to a spreadsheet for documentation and review.
The metadata for each item was reviewed separately by three library staff members. At this stage, reviews simply included a Pass/Fail rating, along with a space for comments describing any noted issues. The Digital Initiatives Librarian reviewed these ratings, quantified them as minor (e.g. small errors in detail, omissions or interpretation that did not fundamentally alter the image’s meaning) or major (e.g. significant misinterpretations that could mislead users about the photograph’s content) and summarized the results. For this assessment, minor errors included imprecise but not misleading terminology (e.g. “desk” instead of “table”), minor transcription errors in visible text, or misidentification of secondary visual elements. Major errors included significant misidentification of primary subjects or objects, attribution of content not present in the image or incorrect interpretation of the photograph’s purpose, context and activities represented. Incompleteness was assessed by whether core descriptive elements (subjects, activities, locations, key objects) visible in the source material were included in the metadata. As shown in Table 1, minor issues were noted in 30%–37% of items, whereas major issues were only identified in 2%–6% of the items.
The metadata results and reviews were then discussed by the project team, with particular attention to the number and perceived significance of the noted issues. A notable concern involved the accuracy of AI-generated descriptions and the potential for misrepresenting historical materials, particularly given academic libraries’ institutional role as providers of trustworthy information. The project team acknowledged that inaccurate or misleading metadata could perpetuate misunderstandings about historical events, misidentify individuals in photographs, introduce unintentionally biased descriptions or provide incorrect contextual information that researchers might rely upon.
Despite these concerns and given the relatively minor nature of the errors noted to date, the team determined that the substantial increase in discoverability and access to these materials outweighed the inherent risks of AI-generated metadata. This decision was supported by two primary factors. First, existing metadata in the digital collections – created over decades by past library staff, students and volunteers – also contained errors, omissions and inaccuracies; Second, without descriptive metadata to provide avenues of meaningful discovery, many of these materials would remain effectively inaccessible to users. With these considerations, the team decided to move forward with the production phase of the project, with the addition of a note in the metadata alerting the user to the use of AI in the creation of the record, allowing users to evaluate the information appropriately.
To facilitate this transparency with our users, the Special Collections and University Archives developed a standard disclaimer for use as appropriate. The following default metadata values are included with each record:
A small amount of additional testing was conducted, making minor changes to the system prompt based on the reviews gathered during the pilot project. Some isolated testing was done using the LLMs identified in Phase 1 to confirm that the performance benefits of the paid Anthropic approach justified the higher cost and potential environmental impact compared to open-source alternatives. The results of this re-testing were consistent with the preliminary testing, so the overall model selected for this project remained unchanged. The most notable change implemented was that, by this time, Anthropic had released a newer version of the sonnet model: claude-sonnet-4-20250514 (Sonnet 4). Targeted testing showed this newer model to be more efficient and effective than the one previously used (see Table 2).
While both models captured the visual content well, Sonnet 4 demonstrated greater precision in identifying and applying details from the image and text. As shown in Table 2’s metadata descriptions of Plate 1, Sonnet 4 correctly identified this image as a “History of the Book” class (hand printed on the back of the photograph), whereas Sonnet 3 used the more generic “Rare book class.” This specificity extended to subject terms, with Sonnet 4 including “Book history” and “Medieval Manuscripts” compared to Sonnet 3’s “Rare books” and “Manuscripts.” However, Sonnet 3 showed greater contextual awareness in other instances, notably including “Special collections” as a subject term where Sonnet 4 used the more generic “College students.” Additionally, Sonnet 4 only included one of the two names (“Clement, Richard,” listed on the card as a subject of the photo under the “Who” label) and omitted the other, more clearly visible name (“McKinney, Dave,” credited as the photographer). This specific scenario was replicated multiple times, showing a consistency in approach of the model ignoring the name when specifically listed as photographer. Although it is hard to say with certainty, it is possible that the model intentionally omitted the name of the creator in favor of focusing on the contents of the image, a nuance that could be explored in future research and serves as a highlight of the types of subtle issues that metadata reviewers can consider when applying this method.
The potential for contextual decision-making became more apparent in another test case. When presented with a black-and-white smiling face icon (uploading as an image file, not as an in-text emoji) (see Plate 2), the Sonnet 3 model processed it and responded with the following JSON:
However, when given the same prompt, the Sonnet 4 model responded with the following text:
I can see that this image shows a simple smiley face emoji or icon - a white circle with two black oval eyes and a curved black smile on a dark gray background with small dots. However, this does not appear to be a historical photograph of an American academic library from the 20th century that would warrant MODS metadata creation according to your instructions.
This appears to be a modern digital screenshot or graphic element rather than archival photographic material that would be relevant to your specialization in American academic library history. Therefore, I cannot provide appropriate MODS metadata for this image as it falls outside the scope of historical library documentation.
Sonnet 4’s ability to recognize contextual appropriateness of the image demonstrates an unforeseen development in the model’s processing and introduces a level of decision-making that could be leveraged in other contexts. While technical compliance (i.e. Sonnet 3 output) can be preferable for batch processing, the interpretation of the image (i.e. Sonnet 4 output) could help prevent the generation of misleading metadata that might confuse users about historical context or image content. This contextual awareness could prove more valuable than technical compliance alone, potentially reducing the need for extensive review in large-scale implementations. This example suggests that model selection for projects of this nature might need to consider the model’s ability to make appropriate judgments in addition to its technical output. In follow-up testing after this project’s completion, changes to the system prompt resulted in the model correctly formatting error messages as JSON responses for improved logging and workflow continuity.
Phase 3: Production deployment
Following the pilot study and subsequent refinements to the model and system prompt, the project team decided to move forward with generating metadata for the entire set of digitized photographs. This comprised 2,263 items (868 captured front only, 1395 captured front-and-back).
The production workflow built on the approach developed during the pilot phase and was implemented using standard Python libraries (pandas, BeautifulSoup) in a Jupyter notebook. See Figure A1 in Appendix 4 for a visual representation of the workflow. Starting with the project directory, all TIF files were systematically identified and organized for processing. Relevant metadata (record group, subgroup, and date) was extracted for each item and included in the query as contextual information for the model.
To manage API costs, processing time and file size constraints while maintaining image quality, TIF files were converted in memory to JPEG format with a maximum side length of 2,400 pixels and then encoded for inclusion in the request. This provided sufficient resolution for content identification and text extraction while reducing the token count for each API request.
As in Phases 1 and 2, each item was processed individually. Responses were decoded and the JSON values parsed into a pandas dataframe. Newly generated metadata (title, description, subject and names) were combined with existing archival information and exported to a CSV file for documentation and for access during subsequent steps. A Python function was set up to collect any error responses for later review; however, there were none, indicating that the previous iterations and code changes had successfully helped avoid errors.
The model was asked to assign subject terms using the FAST vocabulary, with Library of Congress Subject Headings (LCSH) as a fallback if no FAST terms matched. Initial testing had shown only about 2/3 of assigned subject terms successfully mapped to FAST headings, so this fallback was an attempt to increase the number of authorized terms. However, this approach proved to be ineffective (see Project outcomes) and will not be used in future projects. Before inserting these into the MODS record, a separate Python function is run to validate the terms against the FAST API to identify and map authorized terms with their FAST identifiers (see Appendix 3 for examples).
Project outcomes
Metadata review
An initial library staff review of 113 randomly selected images’ metadata records (5% of the total) showed the results to be consistent with the pilot phase. Minor issues were noted in 21%–30% of items, and major issues were identified in 3%–8% of the items. As with the previous discussion, the project team remained convinced that, despite the noted imperfections, the AI-generated metadata provided significantly greater access than the alternative of leaving materials undescribed. To confirm these findings with a larger sample, an additional 113 were reviewed through the same process, again with consistent results. In addition to the 100 records from the pilot project, a total of 226 records (10%) were reviewed in the production stage. In the absence of established standards for metadata quality sampling, we adapted the National Archives and Records Administration (NARA) 10% benchmark for digitization quality control, which addresses image characteristics such as tone, brightness and contrast, rather than metadata. The consistent findings through all three stages of QA review provided our team with reasonable confidence in the overall quality of the metadata. To potentially reduce errors in generated metadata, future implementations could explore a multi-query approach, allowing the model to review each image twice to internally validate its analysis.
Prior to this process, these 2,263 photographs would have had only minimal metadata records available, with the record group and subgroup as a title (e.g. “Watson Library: Special Collections”), no description or named individuals and only generic subject terms (e.g. “University of Kansas. Libraries”). This set of materials was spread among 28 subgroups, resulting in only 28 distinct titles for 2,263 items. This AI-powered process generated new titles, descriptions and subject headings for all items submitted and identified named individuals in 505 items (22% of the total).
The resultant set of metadata records does contain some repeated or similar titles (e.g. 53 images titled “Watson Library exterior view,” with additional variations such as “Watson Library exterior with students on steps”), but they remain descriptive and useful. Despite these repetitions, there are 1,840 unique titles, displaying a significant amount of creativity on the part of the model (e.g. “Library worker operating microfilm reader in Engineering Library” or “Library patron viewing rare book exhibition display case”). There were no duplicated descriptions. Reviewers found the descriptions consistently provided greater specificity than titles, providing searchable access to visual content depicted in the images, such as specific architectural elements (e.g. “arched windows,” “Gothic style”), identifiable objects (e.g. “microfilm reader,” “card catalog”) and much more. The titles and descriptions generated by the model were often surprisingly perceptive, as illustrated by metadata generated for Plate 3. Without the metadata generated by this process, this image would have been included with the title “University of Kansas Libraries – Exhibit Program” and no description. The Sonnet 4 model was able to evaluate this image and provide the very perceptive title “Ukrainian cultural exhibit display case” and description “Display case featuring Ukrainian cultural artifacts and educational materials, including decorated Easter eggs (pysanky), traditional textiles, matryoshka dolls, and informational panels about Ukrainian folk traditions” (see Table 3).
A full computational review of the title fields revealed 223 cases (9.9%) in which a date had been added to end of the title, when it had not been requested. For these records, the title was either deleted or moved to a mods:dateCreated field, as appropriate.
Name and subject were both multi-value fields. The model identified names in 505 images (22% of the collection) and returned 2,536 names (504 unique). These names were written on the photographs, typically on the back. The majority of images did not have names recorded. In a random sampling QA of 25 images (5%), these had clearly been extracted from the text in the image. The accuracy of the transcription varied in line with expectation, with less legible writing resulting in inaccurate metadata. In this case, we will likely perform a targeted metadata cleanup project for the names. Other applications of this approach will need to use caution, depending on the sensitivity of the material or need for higher accuracy.
For the 2,263 images, the model assigned 10,558 subject terms (1,603 unique). A computational review of the subject terms found that nearly two-thirds (6,753 terms) were direct matches to the FAST vocabulary (see Table 4). Although the system prompt directed the model to use LCSH as a fallback vocabulary, post-processing validation through the Library of Congress authority API revealed that, of the terms with no FAST values, there were none that mapped to the LCSH terms. A random sampling of about 15% of the non-mapped terms was manually reviewed and found to be useful as keyword descriptors, so they were retained for increased searchability. For example, the top five non-mapped terms were “University buildings,” “Library staff,” “Library exhibitions,” “Book displays” and “Library interiors.”
Some testing was done using fuzzy matching (i.e. using part of the text to find a match) to map textually similar terms to authorized terms, but this uncontrolled approach led to unwanted matches due to the way in which the FAST API return results. For example, “Library exhibitions” (not an authorized term) did map to the authorized term “Library exhibits” (success), but the phrase “University library” (not an authorized term) mapped to the authorized term “Goa University. Library” (failure). This method was abandoned in favor of higher precision in the matched terms. Future explorations will look at using machine learning matching and embeddings to find closest authorized matches based on submitted keywords, such as that used by the LCSH Validation API (Tang, 2025).
Metadata integration
A collection-specific MODS XML template was created for this project, with default links to the finding aid, rights statements and other fields (see Appendix 3). Using the Python libraries pandas and BeautifulSoup, relevant metadata are added or updated to the template for individual MODS records.
To maintain links between digitized items and their archival context, individual items are mapped to the appropriate level in the finding aid (e.g. folder, item, subgroup, etc.). In addition to the call number, the ArchivesSpace UUID is extracted from the collection’s EAD and added as an identifier in the MODS record, using type= “archivesspace” as an attribute-value pair. This allows a separate Python-based process to generate a digital object in ArchivesSpace with a direct link from the finding aid to the item in our digital collections repository, increasing pathways for navigation and discoverability.
Although the metadata has been newly implemented in our repository, and the resultant digital objects have not undergone usability or user testing, the potential for increased discoverability and accessibility is apparent through visibly improved metadata.
Processing time and cost
To evaluate the overall feasibility and value of this approach and to help plan for future applications in similar projects, we analyzed both processing efficiency and associated costs. Actual processing time by the model was very quick, with an average of approximately 15 s per item. Gathering AI-generated metadata would represent an additional step in our workflow of processing digitized items for our repository; however, since the two primary steps (processing items for API requests and exporting output as structured data) are compatible with existing processes, this can be integrated with our overall workflow with minimal disruption.
An evaluation of the financial cost used for image processing reveals a relatively low amount when compared to traditional metadata creation, given the number of materials and the value returned. The Anthropic console provides a dashboard for reporting on API usage, with viewable metrics, including aggregated input/output tokens and cost. Some of the more specific costs (e.g. number of tokens for text input vs image input) are not reported and have been estimated based on separate testing.
The production phase comprised 2,263 items (868 captured front only, 1395 captured front-and-back), with a total of 3,658 images processed through the model. Based on data retrieved from the Anthropic console, the average per-image API request consisted of an estimated 505–560 tokens of text data (mostly from the system prompt), plus 3,560 tokens of image data, accounting for both single-image and two-image requests (see Table 5 for summary of costs). With a total input token cost of $27.77, this translates to $0.012 per request ($0.015 per request, including output tokens).
Batch processing, an approach that allows sending multiple queries in a single API call, is available with the Anthropic API. This would have had some impact on resource savings through reduced API calls and smaller input text token count. Batch processing was considered and briefly tested during preliminary testing. However, constraints such as local computer memory and large input token counts from the images rendered this impractical.
On review of the usage costs, using a batch processing could have been, in theory, accomplished in as few 75 API calls. The system prompt comprised the largest amount of text input tokens but would only need to be included once per batch request, rather than with each individual request. This efficiency would likely have reduced the text token input by about 90% across all requests. This would have had an estimated impact of a $3.24 reduction in total costs. Input token costs are determined by image dimensions rather than file size or visual complexity. While reducing image size would lower per-image input costs, it might compromise the model’s ability to accurately read text or identify fine details in the photographs. Future projects of this nature would be advised to do additional testing to identify the optimal balance between cost, metadata quality and process complexity for specific materials.
This cost analysis does not consider the staff time invested in model evaluation, prompt development, code creation and quality review. Although systematic time tracking was not implemented for these stages, staff estimates indicate approximately 70 h for testing and refinement, with quality review adding approximately 15–20 h. When distributed across 2,263 items, this represents approximately 2.3 min of staff time per item, the majority of which was foundational development work. Manual metadata creation would have been beyond the resources available for this project, leaving these materials minimally described and underutilized, whereas these records now include searchable titles, descriptions, subject terms and names. These workflows and procedures are now established and can be adapted with minimal additional development time, significantly reducing per-item costs for subsequent digital projects.
Conclusion
This integration of generative AI into an archival metadata workflow represents one practical solution to the challenge of creating descriptive metadata that keeps pace with digitization. The implementation described here demonstrates that AI-generated metadata, when coupled with appropriate transparency measures and quality controls, can significantly enhance collection accessibility without compromising professional standards.
This case study demonstrates how AI-generated metadata can address practical challenges in archival description. The project established a replicable workflow from model selection through production implementation, developed quality assessment methods combining automated vocabulary mapping with human review and integrated AI-generated content into existing systems with appropriate transparency measures. These elements provide a framework for institutions evaluating similar approaches to metadata enhancement.
The technical workflows and high-level cost analysis documented here provide an example for other institutions to evaluate and implement similar approaches. The emphasis on structured output, systematic quality review and transparent user disclosure offers a responsible model for AI adoption in cultural heritage settings. Most significantly, this work addresses fundamental questions of equity and access in digital collections. By transforming minimal and overly broad descriptions into detailed, searchable metadata, AI implementation directly supports both discoverability for researchers and accessibility compliance for users with disabilities. The transparent approach to AI disclosure helps to maintain user trust while enabling institutions to provide enhanced access to previously underutilized collections.
Several limitations should inform interpretation of these results. This implementation represents a single institution’s approach to one collection type: photographs of a known general topic area (i.e. library facilities and activities). Quality assessment relied on manual review of sampled records by staff familiar with the collection, which may not capture all potential issues. Although several models were tested, the production phase used a single commercial AI model, limiting broader conclusions about automated metadata generation capabilities. Additionally, the technical workflows require institutional resources and expertise that may not be universally available. Future research should examine how these approaches perform across different material types, institutional contexts and AI tools.
Despite the recognition of the benefits achieved through this approach, this study does not seek to discount other, less tangible costs and risks associated with the use of generative AI, such as ethical concerns around potentially introducing and sharing unintentional bias or misinformation, or to downplay the environmental impacts of large-scale AI processing. The efficiency gains shown in this case study should not diminish the recognized importance of human expertise in archival description. This case study provides evidence for the practical application of AI in metadata creation when implemented with appropriate quality controls and transparency measures. These technologies can address several challenges in this field, but the goal is to augment rather than replace human understanding and judgment in archival practice, and the application of them must be considered on a case-by-case basis.
As VLMs continue to improve, the methods established in this project are likely to scale to larger collections and different material types. Workflows are likely to change, and, given the rates of model development and improvement, metadata generated through these methods is likely to increase in quality. Beyond the specific technical improvements mentioned earlier, future implementations could explore the application of this approach to a variety of other, more complex, digital collections. Ongoing projects within the KU Libraries include the use of vision-enabled LLMs to create article-level indexing for an historical journal, develop page-level metadata and indices for a digitized newspaper and enhancing existing folder-level descriptions for an archival collection. These and other projects will provide valuable data on the adaptability and limitations of AI-generated metadata across different material types and institutional contexts.
The author would like to thank his library colleagues for their thoughtful feedback and encouragement.
References
Further reading
Appendix 1. Prompt development
In the context of this project, the API query consists of three parts: the system prompt, the text prompt and the image itself. The system prompt provides instructions that give the model guidance on how it will respond, such as its role or expected language. The text prompt is the specific question that is asked of the model. Images must be encoded as byte string to be included in the query.
The text prompt is relatively simple and is designed to present the model with a specific item. The subgroup and date are included when known, as in this example:
Describe this image from the University Archives’ photo collection. It is from the University of Kansas Libraries record group, subgroup Special Collections. The photo is dated 1993.
Development of the system prompt was iterative throughout the entirety of the project. Changes were made in response to model output and intended to elicit the desired results. The system prompt is necessarily more complex than the text prompt, as it needs to serve as a guide for how the model will interpret the query.
Initial system prompt
You are an historian, specializing in American academic libraries in the 19th and 20th century.
You are creating metadata from archival photos, following the MODS metadata standard.
Specific elements that should be created when possible are <title>, <date>, <description>, <subject>, <names>, etc. Response should be given in JSON format.
Scenario 1: Sometimes, there may be a single image. In this case, use your knowledge to evaluate the image and populate the desired elements, using appropriate length and formatting.
Scenario 2: Sometimes, there may be two images. In this case, this represents the front and back of a single photograph. Evaluate the image as in “Scenario 1", but also use any information extracted from the verso of the photograph to enhance the details.
You will also have a filepath for the image. When possible, extract the date from the file path to inform your answer. This should be considered an accurate date.
Final system prompt
You are an historian specializing in American academic libraries, primarily from 20th century. Your task is to analyze archival photos and generate metadata following the MODS (Metadata Object Description Schema) standard: Link to the cited website.
The specific elements to create are:
- ′<title>′ (Required).
- ′<description>′ (Required).
- ′<subject>′ (Required).
- ′<names>′ (Optional).
- ′<date>′ (Optional).
If a photographic element cannot be determined from the context, do not make assumptions and omit it.
Names should be formatted as Last, First. Do not assign roles or authority to your results - only include the values. Do not use [unidentified] for the name element.
When possible, subjects should adhere to the Faceted Application of Subject Terminology (FAST) vocabulary. Use Library of Congress Subject Headings (LCSH) as a fallback if no appropriate FAST term can be determined.
Descriptions should generally be limited to 50 words or fewer, unless warranted by the image composition.
**Image Quality Handling:**.
- If image is too dark, blurry, or damaged to make confident determinations, limit metadata to only what is clearly visible.
- Use bracketed qualifiers when appropriate: [possibly], [appears to be], [unclear], [unidentified].
- For completely illegible text, note as “[illegible handwriting]” or “[text too faded to read].”
**Scenario 1: Single Image**.
Analyze a single image to extract relevant metadata. Use appropriate length and formatting for each element. Base your response on the image content alone.
**Scenario 2: Two Images**.
These represent the front and back of a single photograph. Analyze both sides as in Scenario 1, but use additional information from the verso (back side) to enhance details. Only consider content related to the image, not inventory or processing notes. If numbers contain forward slashes:.
Two integers (e.g. “32/1”) are inventory numbers – ignore them.
Three integers (e.g. “1/12/99”) likely represent a date in M/D/Y format. Include this as an optional <date> element only if confidence is very high.
Return your findings as a JSON object for each scenario. Do not return any extraneous descriptions or text.
Summary of key modifications to the system prompt
Added specific notes on formatting names to facilitate consistent results.
Added a general word count limit to descriptions to help reduce overly verbose responses.
Included explicit directions on processing images of lower visual quality, designed to reduce inaccurate details in the output.
Moved the direction for JSON formatting to the end to aid in the retention of that command.
More clear directions on the prompt (e.g. removed vague wordings: “Sometimes, there may be….).
Inclusion and clarification of expected standards to be followed (MODS, FAST, LCSH).
Additional notes about system prompt development
Versions of the prompt were tested that includes a link to the FAST and LCSH vocabularies. This resulted in a significant increase in token cost and processing time with no notable improvement in increased vocabulary matching.
Following the completion of the project described here, additional testing in the area of prompt development and distributed computing for automated metadata generation continued, with an emphasis on increased complexity of workflow, digital objects and metadata created (Wolfe, 2025).
Appendix 2. Model evaluation
Criteria used for evaluation of model outputs
Visual processing quality
Thoroughness – Did the model identify and describe all relevant visual content in the image?
Perception – Did the model’s interpretation of the visual content of the image match that of the reviewer?
Text identification – If the result included text extracted from the image, was it transcribed correctly? Was any notable text omitted by the model?
Metadata output quality
Natural language – Are the metadata values created by the model conveyed in language that feels natural?
Adherence – How well were the generated metadata fields formed to meet MODS conventions and project specifications?
Technical formatting – Did the output adhere to the JSON standard, without any extraneous text?
Models tested and rated
Although there was not a strict rubric for making these evaluations, a basic three-tier rating system was used, as shown in Table A1. “High” indicates that the model performed very well with little room for improvement; “Medium” indicates the model completed the task as assigned but with errors, omissions or other formatting issues; “Low” indicates poor performance, resulting in output that was unsuitable for use.
Appendix 3. Example MODS records
Project XML template
MODS template created for this photo collection containing a combination of default and empty fields to be used as placeholders.
Example enhanced XML
MODS metadata describing Plate 1. Created from XML template and populated with AI-generated and item-specific metadata.

Appendix 4. Example MODS records
Graphic representation of the processing workflow, depicting image input, AI-generated metadata creation, validation, metadata mapping and repository integration.






