Crowdsourcing and collaboration in digital humanitiesIntroduction
What is now called crowdsourcing – the involvement of the general public in undertaking tasks with multiple granularities in the form of an open call – has a long history. It is important to note that the difference between the traditional mass collaboration projects and the modern phenomenon of crowdsourcing proposed by Howe (2006) is the use of the Internet and various online computer-mediated communication platforms to distribute tasks among large numbers of individuals and interest groups (Terras, 2016). To date, many crowdsourcing projects have been conducted in the business domain, which facilitates the idea of open innovation and leverage the wisdom of crowds to solve real-world problems (Zhao and Zhu, 2014).
In recent years, crowdsourcing has also been leveraged in the cultural heritage domain, such as Galleries, Libraries, Archives and Museums (abbreviated as “GLAMs”), to improve the collection, organization and evaluation of valuable resources (Holley, 2010; Oomen and Aroyo, 2011; Owens, 2013). Crowdsourcing initiatives of GLAMs strive to offer citizens the opportunity of being deeply involved in the production, utilization, communication, preservation and curation of digital collections in the cultural heritage domain (Lankes et al., 2007; Terras, 2016). For example, there have been some attempts to crowdsource complex tasks traditionally only handled by academics or domain experts to the general public (Ridge, 2013; Severson and Sauve, 2019; Zhao and Zhu, 2016). Terras (2016) stresses that many crowdsourcing projects carried out within GLAMs naturally fit under the digital humanities umbrella, and it is difficult to make a distinction between crowdsourcing in cultural heritage sectors and the area of digital humanities, as many projects from GLAMs are leveraging crowdsourcing not only to organize or manage cultural–historical information but to provide the methodologies and mediating tools for generating and co-creating novel information about our past and future. Therefore, in this editorial, we do not distinguish between cultural heritage and digital humanities when discussing the crowdsourcing projects in GLAMs.
In addition, some scholars argue that it is problematic to directly use the industry definition of crowdsourcing to the field of digital humanities (Owens, 2013; Ridge, 2016). Dunn and Hedges (2013) conceptualize crowdsourcing in the context of humanities research and identify four distinguishing factors pertaining to the asset type, process type, task type and output type. Ridge (2013) proposes a new definition of crowdsourcing in the cultural heritage domain and emphasizes the maintenance of granularities or to exploit the volunteer labor unethically as in the business field, but to establish a sustainable and effective interaction through cooperation and collaboration. We agree with Ridge's view (2013) that the concept and definition of crowdsourcing in the digital humanities field deserves further revisit and clarify, especially the new crowdsourcing projects in GLAMs that have emerged in recent years also provide more opportunities for this reflection.
There are several potential benefits in using crowdsourcing within digital humanities (Causer and Terras, 2014a; Holley, 2010; Ridge, 2016; Severson and Sauve, 2019; Terras, 2016). First, cultural heritage institutions can further activate collections and mobilize support through crowdsourcing in resource organization, platform development and service delivery. For example, in terms of resource organization, various resources related to themes can be collected more widely from the whole society through crowdsourcing activities. A typical case is the collection of genealogical information resources, in which crowdsourcing can help GLAMs collect auxiliary information related to specific genealogy, such as genealogical text, pictures, narrative audio, etc. Another example, cultural heritage institutions can leverage crowdsourcing approaches to tag, annotate, describe and visualize existing feature collections, thereby helping GLAMs to develop and utilize related resources. Second, for digital humanities research, the crowdsourcing approach can promote more interdisciplinary and cross-disciplinary cooperation and facilitate collaboration between institutions in the cultural heritage field and effectively engage with external communities, such as research institutions, universities and the general public. Last, from a long-term perspective, crowdsourcing can optimize the value co-creation mechanism of cultural heritage institutions to varying degrees from the resource, platform and service level, thereby enhancing the social impact of GLAMs and driving the sustainable development of research and practice in the field of digital humanities.
A brief overview of the crowdsourcing projects in digital humanities
The ongoing proliferation of online content generation and communication technologies creates many opportunities for public sectors, cultural heritage institutions and humanities scholars to engage the crowd in data collection, sharing, analysis, processing and reuse, sensemaking and value co-creation known as crowdsourcing activities in digital humanities. Although the literature on crowdsourcing and collaboration in digital humanities recognizes the opportunities of some topics – for example, conceptualization of crowdsourcing in GLAMs (Holley, 2010; Oomen and Aroyo, 2011; Ridge, 2016), task characteristics (Carletti et al., 2015; Terras, 2016), user motivation and engagement (Alam and Campbell, 2017; Severson and Sauve, 2019), technological appropriation (Granell and Martínez-Hinarejos, 2016; Iranowska, 2019) and quality control and assessment (Causer et al., 2018; McKinley, 2013; Parent and Eskenazi, 2010) – it remains relatively silent on the sociocultural acts and contextualization used to advance understandings. In this regard, there is a pressing need to explore and understand theoretical, methodological and practical issues of crowdsourcing and collaboration in digital humanities. We provide the following sketches on the existing crowdsourcing projects in digital humanities and try to identify some ongoing issues and challenges.
As shown in Table 1, we formulate a framework for examining the crowdsourcing projects in digital humanities. Particularly, we propose that the crowdsourcing projects in digital humanities can be deconstructed from three aspects, namely, tasks, platforms and participants.
Examples of crowdsourcing projects in digital humanities
| Project | Initiator(s) | Task(s) | Artifact(s) | Intermediary platform | Participants | References |
|---|---|---|---|---|---|---|
| Anno Tate | Tate Britain archive in London (Archives) | Transcription and classification | Artist's sketchbooks, letters and personal papers in the Tate British archives | Adapted | The general public | Iranowska (2019) |
| Australian Newspaper Digitization | National Library of Australia (Library) | Correction | Australian newspaper | Self-developed | Amateur and those with a general interest in Australian history | Alam and Campbell (2012); Terras (2016) |
| Old Weather | The citizen science group Zooniverse, the UK's national weather service the Met Office, the National Maritime Museum and Naval-History.net(Meteorological Department Museum and Gallery) | Correction and transcription | Ships' logs from voyages to the Arctic | Adapted | The general public | Blaser (2014); Terras (2016) |
| Steve Museum | Many American art museums (Museums) | Classification | Art museum collections | Self-developed | The general public | Chun et al. (2006); Trant (2006) |
| The National Archives Transcription Pilot | National Archives and Records Administration (Archives) | Transcription | Handwritten and typed historical documents (letters to civil war spies, presidential records, voting rights petitions and files on fugitive slave cases) | Self-developed | The general public | Noll (2013) |
| Transcribe Bentham | University College, London (University) | Transcription and contextualization | Bentham's manuscripts | Adapted | Students, researchers and the general public | Causer and Terras (2014b); Iranowska (2019) |
| What's on the Menu? | New York Public Library (Library) | Transcription, classification and co-curation | Historic New York restaurant menu | Self-developed | The general public | Lascarides and Vershbow (2014) |
| 1,001 stories of Denmark | The National Heritage Agency (Kulturarvsstyrelsen) and the Danish Communication Agency Advice | Contextualization | Comments, photos, recommendations and interesting stores about a place | Self-developed | The general public | Madsen (2014) |
| Digital (Cultural Institutions) |
| Project | Initiator(s) | Task(s) | Artifact(s) | Intermediary platform | Participants | References |
|---|---|---|---|---|---|---|
| Anno Tate | Tate Britain archive in London (Archives) | Transcription and classification | Artist's sketchbooks, letters and personal papers in the Tate British archives | Adapted | The general public | |
| Australian Newspaper Digitization | National Library of Australia (Library) | Correction | Australian newspaper | Self-developed | Amateur and those with a general interest in Australian history | |
| Old Weather | The citizen science group Zooniverse, the UK's national weather service the Met Office, the National Maritime Museum and | Correction and transcription | Ships' logs from voyages to the Arctic | Adapted | The general public | |
| Steve Museum | Many American art museums (Museums) | Classification | Art museum collections | Self-developed | The general public | |
| The National Archives Transcription Pilot | National Archives and Records Administration (Archives) | Transcription | Handwritten and typed historical documents (letters to civil war spies, presidential records, voting rights petitions and files on fugitive slave cases) | Self-developed | The general public | Noll (2013) |
| Transcribe Bentham | University College, London (University) | Transcription and contextualization | Bentham's manuscripts | Adapted | Students, researchers and the general public | |
| What's on the Menu? | New York Public Library (Library) | Transcription, classification and co-curation | Historic New York restaurant menu | Self-developed | The general public | |
| 1,001 stories of Denmark | The National Heritage Agency (Kulturarvsstyrelsen) and the Danish Communication Agency Advice | Contextualization | Comments, photos, recommendations and interesting stores about a place | Self-developed | The general public | |
| Digital (Cultural Institutions) |
As Oomen and Aroyo (2011) indicated that, some cultural institutions have devoted to engaging the wisdom of the crowd for various tasks, such as content generation and editing, problem-solving and troubleshooting or the organization of information resources and knowledge structures. Thus, a specific classification of crowdsourcing tasks in the cultural heritage domain is proposed according to their tangible outcomes and working practices. Six types of crowdsourcing tasks have been identified and defined, namely, correction and transcription, contextualization, complementing collection, classification, co-curation and crowdfunding (Oomen and Aroyo, 2011). It is important to note that task complexity and task granularity vary between different crowdsourcing initiatives in the field of digital humanities. For instance, social tagging and voting (labeled as classification task) are simpler than correction and transcription tasks, while idea generation (labeled as contextualization task) requires the most intellectual input and knowledge building. As shown in Table 1, the most prevalent crowdsourcing tasks conducted in digital humanities is in the area of manuscript transcription. Although advanced optical character recognition (OCR) technology has greatly facilitated the digitization of manuscripts, it still has limitations in generating high-quality transcripts of handwritten documents (Terras, 2016).
At the same time, there is a growing trend within the field of digital humanities to develop and use platforms to enhance technology-mediated social participation, collaboration and innovation (Carletti, 2016). From the perspective of technology appropriation, there are three main types of crowdsourcing platforms in the digital humanities field, namely, adopted, adapted or self-developed. Adoption means directly using mature crowdsourcing platforms to assign tasks and implement full-process management and curation. For example, Amazon Mechanical Turk (AMT) is often used directly by some crowdsourcing projects in digital humanities to carry out general activities without domain barriers. Adaptation means that the platform owners build their systems by leveraging some universal platforms to make a series of adjustments and extensions, such as using wikis or domain ontologies as means of establishing appropriate platforms that meet specific needs and requirements. Self-development means that the initiators of projects develop crowdsourcing platforms with strong specificity based on the requirements of the digital humanities projects and the characteristics of resources (e.g. feature collections on Chinese handwritten manuscript). Although such platforms are not universal, they can effectively facilitate the completion of specific crowdsourcing tasks. It is important to note that, for different crowdsourcing tasks and GLAM domains, the platform may have different forms and characteristics. In addition to supporting fundamental crowdsourcing tasks, the platform may also need to provide other affordances to encourage volunteer engagement and contribution. Therefore, the user experience design of the platform should play an important role in the success of crowdsourcing projects in the digital humanities field. We suggest that for the interactive design of the platform and interface, more consideration can be given to leveraging gamified elements to improve the usability and sociability of such platforms.
In terms of the participants involved in digital humanities projects, different crowdsourcing initiatives have various requirements. Volunteer engagement in the digital humanities can take many forms, such as rating and tagging to facilitate classification and discovery; commenting and reviewing to add contextual knowledge to cultural–historical artifacts; transcribing manuscripts to improve readability and understandability; collecting data to help information analysis and knowledge discovery; collaborating with the professionals and peers on action research to optimize policy-making (e.g. digital preservation) and develop best practice (e.g. National memory). It is important to note that some of the projects are highly specialized. For example, identifying some copies of ancient Chinese handwritten manuscripts requires both a certain literacy and historical knowledge. In this regard, not all crowdsourcing projects in the digital humanities are suitable for outsourcing to the general public as defined in the original crowdsourcing concept. Some specific crowdsourcing projects in the GLAM context require organizers to identify and recruit suitable candidates to participate; otherwise, it is difficult for such projects to attract and retain a sufficient amount of participants due to the narrow nature of the fields. Owen (2012) suggests that many GLAMs crowdsourcing initiatives endeavor to invite highly motivated and skilled volunteers rather than recruiting large and massive crowds labeled as amateurs. Moreover, the relationship between volunteers and cultural heritage institutions is worth further exploration. The initiators and sponsors should respect their volunteers and actively consider what the volunteers can gain from the crowdsourcing projects in digital humanities, rather than just exploiting them as cheap labors or freelancers.
Papers in this special issue
This special issue of Aslib Journal of Information Management is devoted to the subject of crowdsourcing and collaboration in digital humanities. The guest editors of this Special Issue were Yuxiang (Chris) Zhao, Xiao Hu and Kangning Wei. The deadline for submission was August 30, 2019.
With this special issue, we wished to bring together some of the concerns that seem to be of particular interest to academics today in investigating the theory and practice of crowdsourcing and social innovation in digital humanities. We received 22 manuscripts, and six papers were accepted for publication following a peer-review process performed by 24 external reviewers, for a final acceptance rate of 27%. Overall, the accepted papers explore different topics and dimensions for crowdsourcing in digital humanities and use different methods to analyze multiple levels of research units. The IT artifacts in our special issue collections range from constructs and models to methods and instantiations (e.g. implemented projects). Each contribution is briefly introduced in the following paragraphs. The issue starts with an article which was not submitted to this special issue but to AJIM's regular content but has been included within this issue. Minhyung Kang's research covers the subject of Dual paths to continuous online knowledge sharing: a repetitive behavior perspective.
Suissa, Elmalech and Zhitomirsky-Geffet report on the design of optimized crowdsourcing strategies for OCR postcorrection on selected historical documents. Digitization of historical archives is a challenging task in many cultural heritage projects. The traditional digitization method is to scan the documents into images and then convert images into text using OCR. However, OCR automatic identification of historical documents often has a lot of accuracy issues and needs postprocessing error correction. The authors investigate how crowdsourcing can be employed to correct OCR errors in historical text collections, and which crowdsourcing approach is the most effective in different scenarios and for various research objectives. A series of experiments with different micro-task structures and text lengths were conducted with 753 workers on a general crowdsourcing platform – AMT. The workers were asked to fix OCR errors in a selected historical text. To analyze the results, the authors propose new accuracy and efficiency measures. The results suggest that in terms of accuracy, the optimal text length is medium, and the optimal structure of the experiment is two phases with a scanned image. In terms of efficiency, the best results were achieved when using longer texts in the single-stage structure with no image. The work is one of the first attempts to systematically investigate the influence of various factors on crowdsourcing-based OCR postcorrection and propose an optimal strategy for this process. The findings yield some practical implications for researchers on building the principles and guidelines for automatic OCR postcorrection with a crowdsourcing approach.
Deng et al. explain the co-editing mechanism of wiki-based digital humanities projects (WDHPs) and focus on gathering contextual information in the cultural heritage domain. As defined by Oomen and Aroyo (2011), contextualization is one of the crowdsourcing initiatives in digital humanities and aims at adding contextual knowledge to objects, such as telling stories or writing articles/wiki pages with contextual data. The authors conduct an exploratory study by presenting the co-editing process, evaluating and improving co-editing efficiency. A representative entry's editorial records were reorganized to collect a data set of discussion statements (n = 608), based on which linked structures were built, and the PageRank algorithm was employed to analyze the co-editing process. Furthermore, skewness statistics were used to measure the consensus of co-editing and evolution over time. Linear or curve fitting was performed to analyze the correlation between consensus evolution and its influencing factors. The results suggest that co-editing of online content creation can be considered as a large-scale group discussion. Consensus can evaluate the efficiency of co-editing, which evolves with time and is influenced by the number of statements, breadth and depth of argumentation structure. The authors then examine the findings by selecting “Mogao Grottoes” as an example. In such a case, group discussions around 15 key issues dominate the content creating process. The consensus is on the rise with time and finally reaches a relatively high level. Besides, consensus evolution is more influenced by breadth than by depth of argumentation structure, which indicates that co-editing efficiency of “Mogao Grottoes” is fine, and more argumentation in an academic manner should be guided. The research sheds light on the design and implementation of WDHPs and contributes to the literature on the collaboration model in digital humanities.
To date, little attention has been paid to the impacts of the psychological and cognitive factors on crowdsourced manuscript transcription. In particular, how to recruit suitable volunteers for crowdsourcing projects in digital humanities is an important research topic. In this line, Zhang et al. explore the issues encountered in designing crowdsourcing strategies on manuscript transcription for both efficiency and effectiveness, while striving to encourage opportunities for identifying the appropriate volunteers for such digital humanities projects. The authors conduct a quasi-experiment using Transcribe-Sheng case, which is a well-known cultural heritage project in China, to investigate the influences of the psychological and cognitive factors on the performance of crowdsourced manuscript transcription. Social value orientation and domain knowledge were examined as two main constructs. The empirical study proposes some hypotheses and further examines the assumptions by employing ANOVA tests. Besides, interviews and thematic analyses were conducted to analyze the qualitative data to provide additional insights. The results confirm that in crowdsourced manuscript transcription, social value orientation has a significant effect on participants' cooperation level and transcription quality; domain knowledge has a significant effect on participants' transcription quality, but not on their cooperation level. The results also show the interaction effect of social value orientation and domain knowledge on cooperation levels and the quality of the transcription. The findings shed light on crowdsourcing transcription initiatives in the cultural heritage domain and can be used to facilitate volunteer recruitment in digital humanities.
In terms of the co-curation crowdsourcing task in digital humanities, Hong et al. propose a knowledge extraction framework to extract knowledge, including entities and relationships between them, from unstructured texts in digital humanities. The proposed cooperative crowdsourcing framework (CCF) utilizes both human–computer cooperation and a crowdsourcing approach to achieve high-quality and scalable knowledge extraction. CCF integrates active learning with a novel category-based crowdsourcing mechanism to facilitate domain experts labeling and verifying extracted knowledge. The authors select Tang poetry as the research case and show that CCF can effectively and efficiently extract knowledge from multisourced heterogeneous data in the collection of Tang poetry. Specifically, CCF achieves higher accuracy of knowledge extraction compared with state-of-the-art methods. The contribution of feedbacks to the training model can be maximized by the active learning mechanism. Furthermore, the authors believe that the proposed framework can scale up effective human–computer collaboration by considering the specialization of workers in different categories of tasks. This work has some theoretical contributions that can be generalized to other fields of digital humanities by introducing domain knowledge and experts. Practically, the extracted knowledge is machine-understandable and can support the humanities research of Tang poetry.
Another paper co-authored by Liang, Wang and Li also adopts the Transcribe-Sheng project as the case. The difference is that their research focuses on the task design and assignment of full-text generation on mass Chinese Historical Archives (CHAs) by crowdsourcing, with special attention paid to how to best divide full-text generation tasks into smaller ones assigned to crowdsourced volunteers, and to improve the digitization of mass CHAs and the data-oriented processing of the digital humanities. Task decomposition is identified as an important topic in crowdsourcing campaigns (Zhao and Zhu, 2016). The authors pay attention to task complexities of character recognition of mass CHAs and employ the theories of archival science, including diplomatic of Chinese archival documents and the historical approach of Chinese archival traditions as the theoretical basis and analytical framework in their research. The results show that tasks of full-text generation handled by volunteers in crowdsourcing transcription projects include transcription, punctuation, proofreading, metadata description, segmentation and attribute annotation. The study also provides a metadata element set for volunteers when creating or revising metadata descriptions. Along these lines, the study presents significant insights for application in outlining the principles, methods, activities and procedures of crowdsourced full-text generation for mass CHAs.
Crowdfunding is a relatively new area affiliated with the board concept of crowdsourcing, which could bring many opportunities and future benefits to digital humanities (Terras, 2016). So far, only a few research and impactful projects focus on crowdfunding exploration in digital humanities. Pratono et al. aim to understand how social enterprises adopt crowdfunding in digital humanities by investigating the mission drifting, risk-sharing and human resource practices. The authors present an exploratory case study by observing five different social ventures in Indonesia. This work involves observation of the social enterprises that concern on digital humanities projects and interviews with the key respondents who manage the crowdfunding for sponsoring the projects. This analysis adopts an interpretative approach by involving the respondents to explain the phenomena. The results suggest that the applications of crowdfunding platforms encourage social enterprises to reshape social missions with more responsive action for digital humanities. Crowdfunding allows social enterprises to share the risk with stakeholders that focus on fostering the social impact of digital humanities. In addition, crowdfunding stimulates social enterprises to hire professional workers with flexible work arrangements to attract specific donors and investors. The findings extend the traditional principles of social enterprises by introducing some concepts of crowdfunding in digital humanities. This study also explains the boundary conditions of digital humanities projects and how crowdfunding can support such projects in detail.
Conclusion
Although crowdsourcing, as a model of online collaboration and social innovation has received much attention from digital humanities scholars and cultural heritage institutions, there are still many relevant topics that deserve further exploration. The overarching question is how to use the crowdsourcing model to better promote the social impact and sustainability of digital humanities projects. To this end, research and practice need to be more closely integrated. We believe that scholars of library, information and archives studies should seize the opportunity to carry out more theoretical, empirical, critical and action research studies in this field. For instance, from the perspective of participatory design, how to iteratively develop and optimize user experience for crowdsourcing platforms in digital humanities? From the perspective of volunteers, more legal and ethical issues, such as the ownership of volunteer-generated data and privacy concerns should be further investigated (Terras, 2016). In addition, how to further improve participants' information literacy, media literacy and humanities literacy while launching such crowdsourcing campaigns in digital humanities is worthy of more consideration. From the managerial perspective, more emphasis should be placed on the cooperation and collaboration among the GLAMs to better facilitate the implementation of crowdsourcing projects in digital humanities. Specifically, future research on data infrastructure must pay more attention to data sharing, reuse and curation among cross-projects and cross-platforms.
The publishers would like to thank Professor Dirk Lewandowski, the Editor-in-Chief of the Aslib Journal of Information Management, for supporting and facilitating the special issue. The publishers would also like to express sincere appreciation to all of the reviewers who provided insightful and constructive review comments to the authors. Last but not least, the publishers are grateful to the authors who submitted their manuscripts for consideration and revised the papers based on reviewers' suggestions. Without everyone's efforts, the publishers could not complete this special issue. The publishers hope that the readers will enjoy reading these interesting research articles.
