Changes in information retrieval architecture are causing an increase in the duplication of metadata records in bibliographic databases. This paper aims to examine the problem by explaining how it is caused and the problems it creates for database users.
The paper describes the current state of duplication in bibliographic databases and presents a new way of measuring the duplication. The paper contrasts the nature of monograph, serial, and journal article metadata to show how duplication is different for each.
The new measure of duplication the paper presents could help define the concept of duplication and could aid efforts at eliminating duplicate records in online bibliographic databases.
As metadata becomes a commodity and is sold in aggregate packages it will cause increased duplication in bibliographic databases.
The paper is the first to describe the etiology of duplicate metadata records in web‐scale, bibliographic databases. Also, the paper introduces a new way to measure duplication in these databases.
