The library catalog is not only a useful tool for patrons to find and access library collections but also a valuable dataset for various types of quantitative analysis on our cultural assets. This research aims to develop the latter potential of library metadata by conducting a large-scale analysis of how the metadata fields encoded in the Machine-Readable Cataloging (MARC) standards.
We examined more than 6 million book records from 1980 to 2018 in the catalog of the Library of Congress. We particularly focused on how bibliographic records have changed over time and whether these changes correlate with the introduction of new cataloging policies.
Our results show that more than 200 unique fields and 1,300 unique subfields have been used in the MARC format, though the majority of them are used in fewer than 1% of all records. At the same time, bibliographic records have become increasingly complex, with more fields and subfields per record over the past 40 years. Additionally, there are clear, although sometimes asynchronous, parallel developments between MARC tags and cataloging standards.
This study represents the first large-scale quantitative analysis of the history of library metadata, revealing significant and interesting changes over time and highlighting challenges for the meaningful use of library metadata in the current data environment.
