Purpose

The purpose of this article is to increase our understanding of data papers as research narratives, with a focus on the functions that paradata – information about data creation and management processes and their underlying reasons – have, apart from describing data processes.

Design/methodology/approach

Seven papers from archaeological data journals were selected based on the number of citations they have received specifically for the use of their associated data. The paradata in the seven papers were analysed through close readings of them as narratives, and prominent functions were identified and examined.

Findings

Three expressive paradata functions were found in the data paper narratives, contributing to the papers’ arguments for the usefulness of the datasets, to the tone of the data papers and to the papers’ construction of credibility.

Originality/value

We are aware of no previous studies of paradata as part of data paper narratives or of any studies of data papers employing close reading as an analytical tool.

A data paper is part of the tools, methods, and infrastructures developed for sharing research data in the context of open science (Schöpfel et al., 2019; Li et al., 2020). It has been defined as “a scholarly publication of a searchable metadata document describing a particular online accessible dataset, or a group of datasets, published in accordance with the standard academic practices” (Chavan and Penev, 2011, p. 3). Metadata in data papers include information about data creation and management processes (McGillivray et al., 2022; Li and Jiao, 2022; Jiao et al., 2024). Such data descriptions, or paradata, facilitate research data reuse by enabling data (re)users to gauge the relevance, reliability, and provenance of repository datasets (Huvila et al., 2024, 2025). The concept of paradata is close relative to the notion of provenance (meta)data, however, often with stronger emphasis on being a description of certain processes whereas provenance is typically conceptualised in terms of referring to the provenance of a thing (Sköld, 2025). A narrative structured around the decisions and processes related to data creation and management emerges in data papers (Li et al., 2020), with paradata as key components of that narrative. However, what functions paradata acquire as part of this narrative and how these functions contribute to the communication of the data paper remain unexplored. Our understanding of data sharing via data journals can be enhanced by exploring what paradata can express beyond data description, within the narrative framework of the data paper. Despite the extent of earlier research on data papers and paradata, there is a gap in the literature examining what paradata communicate as part of narratives, beyond descriptive information on data processes.

This article aims to increase our understanding of data papers as research narratives by examining paradata functions in data papers, apart from describing data processes. Archaeology provides a useful context for investigating the question as it is a field without rigorously standardised workflows and varied information-making practices (e.g. Morgan and Wright, 2018; Khazraee, 2019), and one where paradata are crucial for data reuse (Huggett, 2012). These characteristics are, to varying extents, shared across all disciplines within the social sciences and humanities (SSH) domain. The study is based on a close reading of the narratives (Brummett, 2019; Bal, 2004a, b) of the seven archaeological data papers that have received the highest number of citations due to the use of their associated data. The findings suggest that except for describing data, paradata contribute to establishing the credibility of the data producers, emphasising the usefulness of the dataset, and presenting a tonal quality of the text.

Although Parsons and Fox (2013) conclude that no perfect metaphor exists for making data available for other researchers to use in their research, data sharing has become a widely used concept. Pasquetto et al. (2017, p. 2) define data sharing as “the act of releasing data in a form that can be used by other individuals,” but observe that sharing provides little insight into the usability of data. They further note that the distinction between data use and data reuse is that reused data are deployed in a later project by someone other than the original data producer (p. 3).

An increase in the number of research datasets made available for others does not necessarily lead to greater data reuse (Faniel and Jacobsen, 2010). Wang et al. (2021) identify three main paths for data reuse: via data repositories, directly from the original data creator, or through publications. However, these paths are not viewed equally by data reusers. Trust is a key factor in researchers’ willingness to reuse data shared by others, encompassing trust in both the original data producers and the data repository (Faniel and Jacobsen, 2010; Faniel et al., 2019; Yoon and Lee, 2019). The specific contextual information required varies between reusers or groups of reusers (Faniel et al., 2019), but the better the data documentation, the higher the reuser’s satisfaction (Faniel et al., 2016). Data producers are often unaware of the needs of reusers and view the additional documentation required by reusers as unnecessary extra work (Nusser et al., 2021).

Studies of data sharing in archaeological research practice shows that it is a heterogeneous process carried out in many ways using various mediums with few standardised methods (Marwick and Pilaar Birch, 2018). The literature indicates that “informal” means of sharing archaeological research used in settings where colleagues and collaborators are the main stakeholders often include data sharing via e-mail, non-repository databases, notes and drafts, and in meetings and other conversations (see, e.g. Sköld and Andersson, 2025). More “formal” approaches to sharing research data includes the main outputs of scholarly results-reporting (papers, books, reports) including data papers published in data journals, and the sharing of data using data repositories (see, e.g. Börjesson et al., 2022). Despite this variability, common characteristics of data sharing across archaeology and its sub-disciplines have been identified in the literature. Data sharing is understood to be beneficial, as it facilitates research collaboration, promotes more efficient and rigorous practices, opens new avenues for scholarly discovery, and improves the ability to trace the path from research design to results (Sköld, 2025; cf. Borgman, 2012). Data have also been suggested as essential for verifying the reliability of findings; the destructive methods of archaeology make data sharing crucial for preventing irreplaceable knowledge loss; and archaeological data hold public interest, warranting their availability for reuse (Anichini and Gattiglia, 2015). Faniel et al. (2019) found that archaeologists seek rich contextual artifact data to evaluate the reuse potential of data, particularly context information available to excavators in situ, such as stratigraphy and associated artifacts.

A recurring viewpoint in previous studies is that the dissemination of archaeological data is insufficient both in terms of frequency and modes of dissemination (Kansa and Kansa, 2013; Marwick and Pilaar Birch, 2018). Marwick and Pilaar Birch’s (2018) study of how archaeologists share data revealed that data sharing remained largely ad hoc and disaggregated, with structured datasets rarely being shared. Their study further indicated that dataset citations were “almost nonexistent,” with few incentives for data sharing (p. 139). Other commonly mentioned hurdles include the difficulty for research funders and journals to enforce data sharing policies (Marwick and Pilaar Birch, 2018), a lack of incentives, resources, and support within prevailing research funding models (Kansa and Kansa, 2022), and limited data sharing know-how among archaeologists (Kansa and Kansa, 2013).

Data journals are a relatively recent form of research publication, not least within the humanities (García-García et al., 2015), including archaeology. Their purpose is to publish detailed descriptions of datasets intended for use by other researchers (Chavan and Penev, 2011; McGillivray et al., 2022). Data papers have been repeatedly acknowledged for their value to the research ecosystem, particularly because they encourage data sharing (e.g. Chavan and Penev, 2011; Chavan et al., 2013; Candela et al., 2015; Thelwall, 2020; Walters, 2020; Fu et al., 2023). Although the academic value of data papers remains in doubt in some areas of academia (Huang and Jeng, 2022), and although the citation of data papers as sources of data is still fairly rare (Jiao et al., 2024), recent studies suggest that data papers positively impact the scientific and scholarly system (Thelwall, 2020; Jiao and Darch, 2020; McGillivray et al., 2022; Fu et al., 2023). McGillivray et al. (2022) and Fu et al. (2023) conclude that data papers increase both the visibility – this is observed also by Thelwall (2020) – and the citation count of the associated research papers. Jiao and Darch (2020) find that data papers facilitate data sharing and data reuse, but only to a limited extent. Additionally, there has been a steady increase in using data papers as a means of citing datasets (Jiao et al., 2024), and it has been suggested that data journals could be a way to overcome some of the barriers to data sharing (Candela et al., 2015; Jiao et al., 2024).

Studies also underscore the significance of reporting paradata related to data creation and related methodology in data journals. García-García et al. (2015) conclude that data papers enhance the visibility of methodological aspects in research, while Li et al. (2020) found that data collection events constituted the largest category in their analysis of data events described in data papers. Although the vast majority of researchers in the sciences and social sciences expect data collection methods to be included in data peer review (Kratz and Strasser, 2015), and future peer reviewers will need to understand data collection, management, and publishing (Chavan and Penev, 2011), Li et al. (2020) caution that the methodological details recorded in data papers should not be assumed to fully reflect the actual research process.

Documentation of contextual information on data and especially data generation has been shown to be demanding and is often missing in established data documentation schemes (Birkin, 2020; Faniel et al., 2019). Paradata are therefore documented unevenly and are often missing in formal documentation (Huvila, 2022). However, research indicates that information that qualifies as paradata can be found across data documentation. Studies of archaeological documentation have found paradata in, for example, investigation reports, narrative descriptions, diagrams, and naming of specific methods (Huvila et al., 2021), as well as in citations to methods descriptions and other literature (Huvila et al., 2021, 2022b). Paradata can also be extracted from the data by analysing traces of instrumentation used, naming conventions, and data structures (Börjesson et al., 2022; Huvila et al., 2023a). They are present in standards, though unevenly and often without clear guidelines on implementation (Chao, 2014; Börjesson et al., 2020).

To document and share intellectual decisions underlying data generation, traditionally informally documented in illustrations and personal notebooks, is very challenging (Sleath et al., 2024; Canfield et al., 2011, p. 232). Research on paradata in archaeological reports revealed incomplete coverage that focused on practices central to data generation and management rather than on practices significant for reuse (Huvila et al., 2022a). While data creators emphasise technical and administrative details, reusers seek narrative descriptions of data generation and prior use, as well as information on what data were collected or included in the final dataset (Huvila et al., 2025; Gregory et al., 2023; Bishop and Kuula-Luumi, 2017). In the case of data papers – similar to research reports in their recording of data context, creation, and processing – no corresponding research has been conducted to date. Consequently, it remains unclear what types of paradata data papers contain, and what is their extent.

Two data journals publish data papers specifically for archaeology research. As of 2023, the Journal of Open Archaeology Data (JOAD) has published 75 data papers since 2012, and internet archaeology (IA) has published 7 since 2013. Only a few papers have been cited more than once or twice in publications that used the associated data in some fashion.

A lack of citations for a data paper does not necessarily indicate a lack of reuse of the associated dataset, and there is no reliable method to measure data reuse (Pronk, 2019). Reuse can be acknowledged by, for instance, citing the dataset directly or mentioning dataset and creator/-s within the body of the text (Jiao et al., 2024). The point of data papers is to make data available for reuse as well as provide citable objects that can contribute to giving academic credit for the production of data that are useful to others. To focus our analysis on data that have demonstrated usefulness beyond its original research context, we selected data papers with at least five citations, using citation frequency as a proxy for their usefulness and impact, to identify the types of data most commonly cited.

Bibliometric analysis was used to determine which of the JOAD and Internet Archaeology data papers have had their data cited most. Using Publish or Perish 8, the total number of citations for JOAD and Internet Archaeology papers in Google Scholar were identified. Google Scholar proved to have better coverage than alternative bibliometric databases for these journals. The search was run in January 2024 on all articles published by JOAD from the journal’s inception in 2012 through 2023, and Internet Archaeology since its introduction of data papers in 2013 through 2023. Of the 82 data papers, 27 articles were cited at least 5 times (JOAD 25 and IA 2). Each citation of the data papers was manually reviewed to identify cases when the data papers were cited for their data. The seven cases with at least 5 citations were selected for closer analysis (Table 1).

Table 1

Papers with 5+ data citations in JOAD and Internet Archaeology, 2012–2023

YearVol: no.Author(s)CitationsData citations (self*)Abbrev.**Article title
20165:2Manning, K., Colledge, S., Crema, E., Shennan, S., Timpson, A.5024 (9)JOAD 5:2The cultural evolution of Neolithic Europe. EUROEVOL dataset 1: Sites, phases and radiocarbon data
20186:2Palmisano, A., Bevan, A., Shennan, S.269 (5)JOAD 6:2Regional demographic trends and settlement patterns in central Italy: archaeological sites and radiocarbon dates
20165:1Colledge, S.137 (2)JOAD 5:1The cultural evolution of Neolithic Europe. EUROEVOL dataset 3: Archaeobotanical data
20165:3Manning, K.106 (0)JOAD 5:3The cultural evolution of Neolithic Europe. EUROEVOL dataset 2: zooarchaeological data
20197:5Bates, J185 (2)JOAD 7:5The published archaeobotanical data from the Indus civilisation, South Asia, c. 3200–1500 BC
20219:1Martínez-Grau, H., Morell-Rovira, B., Antonlín, F.125 (1)JOAD 9:1Radiocarbon dates associated to Neolithic con-texts (c. 5900–2000 Cal BC) from the North-western Mediterranean Arch to the High Rhine Area
20197:4Pardo-Gordó, S., Puchol, O. G., Aubán, J. B., Castillo, AD.105 (0)JOAD 7:4Timing the Mesolithic-Neolithic Transition in the Iberian Peninsula: the radiocarbon dataset

Note(s): *) Self-citations are defined as citations in a publication that has at least one author in common with the data paper

**) For simplicity, these abbreviations will be used in the discussion below

Source(s): Authors’ own creation/work

Only a fraction of citations to data papers indicate data reuse; for instance, Jiao and Darch (2020) found this was true for about half, while our study found barely a quarter. Non-reuse citations may credit previous studies, employ the same approach, or provide background (Jiao and Darch, 2020; Jiao et al., 2024). One common reason among the papers that cited JOAD data papers was to cite them as examples of available data for a particular geographic area or of a particular type. The results contain examples of self-citations, defined as data papers cited by one or more of the paper’s authors in order to use the data in another article. As this is one of the reasons frequently mentioned for publishing a data paper – to provide a citable object for a dataset and thus provide academic credit for producing data that can be used by others (Chavan and Penev, 2011; Kratz and Strasser, 2014; Stuart, 2017; Huang and Jeng, 2022) – we have included self-citations in determining which data papers have been most cited for data reuse.

Three general points should be borne in mind when evaluating the citation of data papers, not least in archaeology, and which have a bearing on our results. First, older papers have an advantage in terms of cumulative number of citations due to longer visibility – it takes time for a paper to be disseminated, read, and cited. Second, although data reuse could be discussed in terms of actual practice or efficiency gain (Faniel and Jacobsen, 2010; Wallis et al., 2013; Faniel et al., 2019; Pronk, 2019), it should also be considered in relation to a dataset’s potential long-term value. The dataset of Swedish funerary practices in the 1990s (Gustafsson, 2013) may be interesting shortly after its creation, and again several decades later, for a comparative study. In the open data discussion, there is a tendency to expect data reuse within a few years of the original study (Pronk, 2019). While this expectation may be reasonable in some disciplines, it does not hold true for others.

Finally, not all datasets presented in data papers are equally useful or interesting to other users (Wallis et al., 2013). Pasquetto et al. (2017) posited that little is known about how scientists decide between collecting new data and using existing data, and in a survey of researchers, Shen (2016) found that only 6% of researchers across disciplines engaged in data reuse. According to Huggett (2016), archaeologists are not using archived digital data to any greater extent and Sobotkova (2018) suggested that the reason why archaeologists avoid reusing others’ data is because it is time-consuming and lacks the prestige associated with fieldwork. Although sharing data through a data paper has the merit of providing another publication for the data producer (Huang and Jeng, 2022), the majority of associated datasets are not in great demand.

A key limitation of our study is the restrictive selection criteria for data papers. By including only those papers from archaeological data journals that have been cited at least five times, we limited our sample to seven cases. This approach was designed to focus on data papers that demonstrated some degree of reuse, ensuring that our analysis centred on influential examples. Our findings cannot be assumed to represent all data papers, or even all archaeological data papers. Less frequently cited papers may contain valuable insights into paradata expression but remain unexplored in this study.

Once the seven data papers were identified, the paradata they contained were subjected to close readings (Brummett, 2019). All information that described or explained the processes of dataset creation, including reasons that had shaped those processes, was analysed for language, structures, and meanings. Particular attention was paid to semantic nuances, narrative devices, and tacit implications. Striking and recurring patterns in the analyses were then identified, in particular those that occurred in multiple cases.

The paradata close reading was grounded in narratology. As pointed out by Bal (2004a, b), narratology can be usefully applied to non-fiction narratives, examples of which include those describing data creation processes. Data papers result from both intentional and unintentional processes of selection (inclusion and exclusion) and presentation, and are thus not neutral accounts of events – they are narratives. Like descriptions of data processes in general (Knorr-Cetina, 1981, pp. 115–118; Li et al., 2020; Huvila et al., 2021), data papers were not assumed to neither be complete nor completely accurate accounts of events and reasons. However, as with other narratives, careful analysis can uncover how different types of meaning are expressed through various components – their expressive functions. In literary narratives, components can express for instance stance (e.g. irony), genre (e.g. fantasy), or narrator reliability. In a research narrative, they can express meanings such as authority, theoretical position, or community of practice.

The paradata analysis is text-focused, treating the data papers as narrativised accounts of processes of data creation and integration. Instead of making assumptions about the individuals who wrote the text or its actual readers, our analysis focuses on the textual representation of paradata, in order to avoid the risk of being misled by conjectures. Having no direct knowledge of actual intentions, assumptions about them could introduce unwarranted interpretations and compromise the analysis. We employ a simplified communicative model of the text, considering, on the one hand, the “narrator,” which, according to Bal (1997), is the text’s “linguistic subject, a function and not a person, which expresses itself in the language that constitutes the text” (p. 16). On the other hand, there is the “reader,” the constructed (implied) recipient of the text’s message. In other words, we conflate the textual construct implied author with the narrator function, and the construct implied reader with the narratee function (see Chatman, 1978, p. 151). As mentioned, we make no pronouncements about the actions, considerations, or intentions of the actual authors or readers of the data papers; our analysis is limited to what is present in the texts. For studies of actual actors’ intentions with paradata, see Huvila et al. (2024, 2025).

When enquiring into archaeological data papers, it is relevant to take into consideration archaeological data practices and their characteristics. Unlike many other disciplines especially in social sciences and humanities, archaeology is heavily focused on data creation and management. Archaeological fieldwork generates a great deal of data and in most parts of the world the legislation mandates their adequate documentation and preservation for future use. There is great variation in practices and regulations, however. Most data are preserved in archives and repositories rather than published. Moreover, even if archiving of digital data is advancing, not all repositories accept digital datasets (Huvila, 2018; Richards et al., 2021). This means that some data are only available in printed or PDF reports and therefore difficult to reuse (Sobotkova, 2018). Considering the volume of archaeological data generation, so far very few data papers have been published even if such papers have been featured as a means to advance the reuse of datasets that are often underutilised (Farina et al., 2025 cf. Boi et al., 2015). As a previous pilot study on archaeological data papers indicated, their structure and contents vary considerably.

The seven most-cited data papers have all been published in JOAD, and they are all descriptions of integrated datasets. The data described in JOAD and Internet Archaeology have either been produced as primary data, for instance from a particular excavation, or they integrate datasets with different origins. The most-cited data papers have all been produced by gathering data from different sources, transforming them, and combining them into a single, unified dataset (Kintigh et al., 2018, p. 31), often a database. For example, JOAD 9:1 integrates radiocarbon dates from 745 sites and JOAD 7:5 integrates archaeobotanical data from 63 sites, in both cases using numerous sources. The paradata provided therefore reflect the integration process rather than the original data production (for challenges related to harmonising and integrating archaeological data, see McKeague et al., 2020). In the following discussion, we distinguish between creation paradata, related to the processes involved in original data creation, and integration paradata, associated with the integration of several data sources into a larger database or dataset.

The seven data papers are almost completely restricted to integration paradata, referring secondary users of the data to the constituent datasets. The Methods sections of the seven data papers focus on the compilation and integration of previously produced (original) datasets, with little or no paradata pertaining to the original data creation processes. Most data papers explain that they include references to the constituent datasets, indicating that secondary data users in need of creation (as opposed to integration) paradata can find them in the constituent datasets. Two papers (JOAD 5:1, 5:3) explicitly state that the underlying data have been archived in a particular place, suggesting that this would be a convenient way to find creation paradata. Other papers are less specific about access to the constituent data and creation paradata, referring to the information provided in the integrated dataset. No paper includes any significant amount of creation paradata.

The lack of creation paradata describing the original data creation processes in the studied data papers aligns with observations made in earlier research on data publishing. It has been observed that data papers are expected to contain paradata about the main actors and procedures involved in creating and analysing the dataset at the centre of the data paper (Callaghan et al., 2012; Jiao et al., 2024). These expectations are often underpinned by the notion that data papers should facilitate dataset reuse and trust-building (Chao, 2015). However, studies show that data papers alone are often not sufficient as stand-alone providers of creation and methods paradata to support data reuse (Jiao and Darch, 2020). Paradata in data publishing emerge as a fragmented and networked phenomenon, where information describing various aspects of data creation processes exists across datasets, data papers, research publications, and unpublished research documentation (Sköld and Andersson, 2025). One explanation for this may be the greatly varying disciplinary traditions and the flexible, general policies of data publishing outlets regarding what paradata data authors are requested to document when publishing research data (Jiao and Darch, 2020). From this perspective, the lack of data creation paradata in the studied dataset is, to some extent, consistent with earlier observations. It is nevertheless notable that all data papers in the corpus contained little to creation paradata – despite several studies identifying it as the most commonly occurring type in data papers (Li et al., 2020; Li and Jiao, 2022).

A pilot study by Huvila et al. (2023b) suggests that although authors of data papers implement various strategies to document the research process, these papers may not effectively serve as paradata documentation. Given that data papers detail the methods used to produce data – a fundamental aspect of paradata – it raises the question: if they are not effective as paradata documentation, what are they good for?

A close reading of the paradata in the seven data papers revealed how, apart from describing the associated data, paradata express three other communicative dimensions of the paper. The central function of paradata is to describe the processes related to research data creation and management. However, a close analysis and interpretation of how paradata in the seven data papers are presented and what paradata are included or left out revealed that paradata also have three other, interrelated and expressive, functions: to establish the credibility of the authors, to emphasise the usefulness of the dataset, and to contribute to the tone of the text.

The expressive functions of paradata are interrelated concepts, addressing various aspects of what paradata communicate much like how the Vitruvian triad addresses aspects of a building. In De Architectura (c. 27 BCE), Vitruvius introduced a triad of fundamental principles for constructing buildings: firmitas (structural integrity), utilitas (fitness for purpose), and venustas (aesthetic qualities). These principles apply to the architectural design of buildings, but also provide tools for thinking more broadly about combinations of other interrelated concepts. From the descriptive paradata in the data papers emerge three other functions: the credibility of the data papers and datasets, the scientific usefulness of the datasets, and the textual tone of the papers. These aspects can be used analogously to the Vitruvian triad, to discuss what paradata descriptions communicate. The three concepts are not mutually exclusive; a single paradata element expresses all three functions – like an architectural element functions within, and can be discussed in terms of, all of Vitruvius’ categories. Paradata that emphasise the usefulness of a dataset may simultaneously contribute to the tone of the text and affect the (constructed) credibility of the author. Consequently, a comprehensive analysis of the expressive functions of paradata requires the consideration of all three aspects.

The usefulness function uses paradata to construct or express arguments for the utility or scientific value of the dataset described in the data paper. Earlier information science research has occasionally referred to the (notion of) usefulness of information although typically eclipsed by the related concept of relevance (Huvila et al., 2019). One purpose of data papers is to describe datasets that can be shared and reused. Paradata are frequently used as part of an argument in favour of reusing the data that are being described. Rather than in terms of its intrinsic relevance (cf. Huvila et al., 2019; Borlund, 2003), data usefulness is expressed through paradata related to, for instance, the user-friendliness of database structures and interfaces, dataset versatility and quality, and potential for different kinds of analyses. Usefulness of data does not necessarily equal its relevance or usability. A relevant dataset can be non-useful – it does not help to reach specific goals – similarly to how it can be unusable – it cannot be used for a particular purpose (cf. Huvila et al., 2019; Buchanan and Salako, 2009).

The usefulness function emerges to different degrees in the data papers, with JOAD 9:1 presenting the clearest case of focusing on establishing usefulness. According to the JOAD template, the subsection Steps should include “[t]he series of procedures followed to produce the dataset,” including “any source data used, as well as software and instrumentation involved” (p. 2). Instead, JOAD 9:1 uses the subsection to detail the “[s]ix main fields of information related to each radiocarbon date” (p. 3). These fields are described in a way that outlines the dataset structure – what categories of information are available for each radiocarbon date – rather than any data creation process. This focus on the structure of the data, rather than on the processes used to collect or integrate them, suggests that the dataset’s usefulness is prioritised in the paper over the hows and whys of its creation. Neither creation nor integration paradata are included to any greater extent; instead, the detailed description of how the data are structured in the dataset foregrounds ways in which it can be reused for future research. In JOAD 9:1, data usefulness is also foregrounded in other ways, for example by stressing that the database is built as a “user-friendly tool” with “high variability of information fields” (p. 2), although it is possible to read the connection between the dataset and the “broader [project name] database” (p. 3) as suggesting that the project database is more useful than the presented dataset.

One way of using paradata to elevate the usefulness of a dataset is to foreground its versatility. When the dataset is presented in JOAD 7:5, it is stressed that although grey literature [1] is not included in the dataset for the sake of access and verifiability, the data structure is designed to facilitate future incorporation of data by reusers. In this way, the paper demonstrates the usefulness of the dataset despite the restrictions placed on the data collection. A versatile data structure is used in order to minimise the trade-off between data verifiability and usefulness. Versatility of a different kind is exemplified by the three datasets presented in JOAD 5:1, 5:2, and 5:3. A central aspect of these three datasets is that they are linked to each other, something which is made very clear in the data papers. This link between the datasets not only means that paradata overlap and combine between the papers, it also means that it is possible to carry out data analyses that combine radiocarbon, zooarchaeological, and archaeobotanical data for the same geographical area and timespan. These links serve to elevate the scientific value of reuse by providing a broad, analytical versatility. (Their usefulness is indicated by the fact that the three papers are three of the four most-cited data papers in this study.)

The expressive function of tone uses paradata in a way that sets the tone or attitude of the text in relation to the dataset. Textual tone has been approached quantitatively, by analysing the frequency and distribution of, for example, positive, neutral, and negative words in texts, often using standardised wordlists (Bassyouny et al., 2020; Li et al., 2022) and specialised computer software (Fisher et al., 2020). This approach does not take context or connotations into consideration, however, and is not a substitute for “close reading” (Fisher et al., 2020, p. 85). Instead, we draw on literary scholarship, in which the tone of a text is a quality which pertains to the mediating or narratorial voice a person imagines they hear when they read (Stockwell, 2014, p. 362; Roof, 2020, p. x). It has a ubiquitous character, arising out of a combined effect of “diction, syntax, connotation, and cultural associations we might link to specific words, naming practices, and cultural patterns” (Roof, 2020, p. 7).

All utterances in a text add to the tonal quality, and in a data paper, paradata has a central and dominant position. They are an integral part of the paper’s data descriptions, lending weight to the tonal aspect of how paradata are communicated as well as the specific paradata conveyed. In the seven data papers, paradata contribute to or amplify tones that come across as, for instance, anxious or assertive, didactic or defensive. As demonstrated by Roof, the elusiveness of tone and the complexity of its composition means that an exhaustive tonal analysis needs to be extensive and detailed, far more so than the scope of this article allows. What we have included should be considered brief examples only.

The following two sentences illustrate how paradata in JOAD 7:5 are presented with, and contribute to, an anxious or meticulous tone:

This dataset was obtained directly from source publications. PhD theses and unpublished reports were not included as these grey literature [sic] have not always been digitised or made available in the same way that journal articles are. (p. 2)

The first sentence communicates the most fundamental piece of paradata for a dataset: its provenance or origin. It does so in a syntactically straightforward manner and the information is not even new: dataset provenance is described in much more detail on the previous page (“The dataset described here was created through the systematic collation of primary archaeobotanical results published up to October 2017”). The neutral sentence stress lies on the new piece of information, the adverb “directly” – the data were obtained directly from publications – in the middle of the sentence. Nothing has come between the dataset and the source; what is in the dataset is, as it were, pure, uncontaminated. The tone confers authority from the publications to the dataset (thus contributing to its credibility), and the publications, in turn, have the authority of the academic publishing system. But the adverb plays a double role: it draws on the authority of the publications but it also shifts responsibility to them. Data were obtained directly from publications: whatever problems they may have is the fault of those publications.

The second sentence also deflects problems with data away from the dataset. It explains that “publications” do not include PhD theses and unpublished reports; they lack the authority of the quality checks involved in a publishing system. There is also additional information underpinning the decision to exclude theses and reports. This unit of paradata exemplifies how the tone can be affected by the inclusion of information – the explanation is not necessary and thus gains weight by its inclusion. The reason for excluding theses and reports is that there is no guarantee that these can be accessed: they are not always digitised or – otherwise? – made available. This access is important; the reuser must be able to verify the data in the dataset. This suggests that neither credence nor correctness is assumed. Like the previous sentence, it shifts responsibility, this time to the reuser, by clarifying that verification is possible and, therefore, the responsible thing to do.

Our second example of tone comes from JOAD 7:4. The two sentences corresponding to the previous example reads as much more authoritative:

The database has been built based on rigorous research criteria. The data core was obtained directly from both published papers and grey literature, as well as information provided by colleagues that work in the Iberia Peninsula. (p. 2)

The first sentence offers underpinning information that contributes very little information about the data-related processes apart from their quality. It does not attempt to persuade rationally, it is a statement of the nature of the data process. Not only are there research criteria underlying the database, these are “rigorous” – thought-through, adhered to – and require no explanation; their scientific quality should be taken as read. Just like the sentence compresses two clauses (about the building of the database and about what the database is based on), it compresses the concepts of the database and of scientific rigour. There is no space between the two concepts; any room for doubt is syntactically eliminated. The second sentence expresses provenance through the same adverbial construction as the first example: data have been obtained directly. Here, there are other options for neutral stress than the adverb, however. Directly emphasises the scientific rigour of the first sentence, but it does not shift authority in the same way. With three data sources – publications, grey literature, and colleagues – there is no quality system to provide authority, and there is no need for it. It has already been declared that rigorous research criteria provided the basis for the database; credibility is assumed, authority resides in the authors of the paper. The three types of sources could even be taken to cover all possible kinds of sources, including directly from colleagues, suggesting privileged access to data that have not even been made available as grey literature. These sentences appear to convey that the database is a first-rate data product, containing all data there is on this subject.

The previous examples, of anxious and authoritative tones, illustrate how elusive tone can be; the third example will demonstrate a tone that is more readily identifiable. Part of JOAD 5:1 adopts a distinctly didactic tone as it explains the potential constraints of the dataset:

In comparison to charring, for example, waterlogging results in less taphonomic bias and thus a far greater diversity of taxa is preserved which is more likely to comprise the full spectrum of species originally used; under all other preservation conditions the large seeded crops and wild edible fruits/nuts are resistant to decay but fragile taxa rarely survive, hence the dataset is biased in favour of the more robust plant parts […]. (p. 3)

This long sentence not only points out a potential bias, it explains this bias by means of comparisons and examples. A series of implications are stringed together: the result of waterlogging as opposed to charring, what thus follows from this result, which in turn has (likely) consequences for what can be known; the effect on certain kinds of taxa but not on fragile ones, and, finally, what this means (hence) for the dataset. Apart from the final part about dataset bias, the series of implications are not specific to this particular dataset; they are broadly applicable archaeobotanical facts, supported by a reference to an article. This reference both corroborates the facts provided and offers a way for the reader to learn more.

The way these facts are narrated constructs the reader as having some archaeological knowledge and willingness to learn more. Through the paradata, the (implied) reader is presented as someone who is familiar with archaeology – they are expected to know what, for example, “taphonomic bias” is – but who needs to be informed about or reminded of how preservation conditions give rise to bias in the preserved archaeobotanical material. This sentence exemplifies a didactic tone that conveys facts and causal links cogently in order to ensure that a particular quality of the dataset is comprehensible.

Already the cursory readings of these examples demonstrate paradata’s effect on tone. In the first example (JOAD 7:5), the tone is characterised by avoiding responsibility and drawing on external authority by carefully providing detailed paradata. This tone comes across as anxious to avoid problems, keen to not contribute to mistakes. The second example (JOAD 7:4), has a tone which is decidedly more self-assured or authoritative. The tone is based on the assumption that a declaration of scientific rigour is sufficient to be trusted; there is no need to defer authority or blame. And in the third example (JOAD 5:1), the tone is noticeably didactic, explaining general facts and causal links in order to elucidate a particular bias in the dataset.

The final expressive function of paradata is to enhance the text’s credibility. Credibility is one of the key concepts in information science research (Rieh, 2010), with a range of meanings, entangled with other concepts, not least trust and trustworthiness (Rubin, 2022, p. 62). The predominant focus for credibility research appears to be either investigating criteria or explicating strategies for credibility assessment (Huvila, 2020), and thus discusses credibility in terms of perception and assessment. However, in this article we use credibility to refer to a textual construct. It is established by the narrator through a combination of strategies meant to give credence to the text’s account, and is related to what Francke (2008) refers to as the text “signalling […] cognitive authority potential” (p. 318; she employs Wilson’s (1983) concept). Huvila et al. (2021, p. 1118) point out that building trust is a major reason for archaeological documentation of information making, and draw parallels to various other accounts. They note that to build trust, transparency might not be the best or most economic approach; other means and various levels of details can be used to appear credible enough. Depending on analytical perspective, the method descriptions in the data papers signal cognitive authority potential, build trust, or construct credibility through a range of strategies that use integration paradata in various ways. Some examples include drawing on an authority external to the text (such as the cognitive authority of someone else, cf. Francke, 2008, pp. 320–330), using conventions of something with greater credibility (Francke, 2008, pp. 335–343), or demonstrating scientific or scholarly expertise. In the data papers, establishing credibility emerges as a significant expressive function of paradata, whether by invoking scientific rigour, establishing authority, or accentuating the validity and verifiability of the datasets.

When we looked at the respective tones of JOAD 7:4 and 7:5, above, we demonstrated how they either confidently assumed that the paper would be considered credible or anxiously used various strategies to demonstrate credibility. Of the seven papers, these two provide the extremes of a scale of credibility strategies, between the assumption of credibility from the start and the assumption that credibility must be established in various ways. But all papers employ various strategies to enhance their credibility, and some of these strategies involve paradata.

The most common strategy in which paradata are used to enhance credibility is by situating the dataset in a context which lends credence to it. Descriptions of prestigious funders (e.g. JOAD 9:1, 6:2, 7:4) or major projects (JOAD 5:1, 5:2, 5:3), or the presentation of the data collection as the successor or “heir” to previous (important) projects (e.g. JOAD 6:2, 7:5), are strategies in which external authority is invoked in order to enhance a dataset’s (or data paper’s) credibility. Such paradata are mainly located in the Context subsection, near the top of the papers, but can occasionally be found towards the end, in the subsection Creation Dates. By book-ending the data description with such credibility-enhancing paradata, the reader would both enter and leave the description being informed of the credibility of the data.

Another way to enhance credibility is to demonstrate or foreground scientific expertise. This is a strategy particularly visible in JOAD 7:5, in which both data collection and integration processes and their underlying reasons are clearly laid out in the text. There is a comprehensive table of taxa that may cause problems with regard to data quality, and notes on how these problems have been dealt with. The EUROEVOL datasets in JOAD 5:1 through 5:3 are presented with similar clarity, carefully explaining data processes and their underpinnings. In JOAD 5:2, the display of expertise and meticulous care combines with a defensive tone in the Constraints subsection:

Although we have undertaken painstaking steps in data cleaning and checking, a considerable portion of this dataset was received without source publications or original reports. As such we initially inherited a number of errors, and it may still be possible for some to remain. We would encourage all users to double check for duplications, and if possible to inform us of any possible errors in order for us to correct the online repository. (p. 3)

This paragraph contrasts the solid, scholarly work carried out to construct the dataset with the source of possible errors. Painstaking steps have been undertaken to rid the “considerable portion” of the dataset of inherited errors. The awareness of imperfection adds credence to the claims of expertise, although the thrice repeated “possible” stresses the hypothetical nature of any errors. The final sentence is arguably not paradata but part of the construction of credibility through demonstrated expertise.

A third strategy to establish credibility involves foregrounding the dataset’s validity and verifiability. JOAD 7:5 uses verifiability as a reason for why only published data sources have been used to construct the dataset, and validity for why data have been converted from values into presence/absence information. It is common to include information about data sources in the datasets, but not all papers use paradata to underscore the possibility to validate the datasets. In JOAD 9:1, it is clearly explained that there is a Reference field in the database, which contains the references for the sources of radiocarbon information and whether it was obtained from a published database (p. 3). In JOAD 6:2 and 7:4, this kind of paradata is included in the description of the data objects. The presence of references for each data point implies that any reuser could verify the data if needed thereby suggesting validity and enhancing credibility.

Credibility is not only enhanced by paradata, it can also be undermined by them. While a strong focus on detail on the whole lends credence, inconsistencies in the details may detract from the overall credibility of the data paper. For example, in JOAD 9:1, inconsistencies and irregularities in the paradata – such as inconsistent terminology and unclear descriptions – convey an impression of sloppiness, suggesting a lack of care, which may reflect poorly on the quality of both the data and the text. Inconsistent and incorrect paradata could thus undermine the credibility that the article aims to establish.

The notion of data quality is a prominent feature that contributes much to all three expressive functions, and serves almost as much as an argument in favour of data usefulness as it boosts credibility. The JOAD template encourages reflections on data quality through the Quality Control subsection, in which a paper should list “the methods used for quality control in the production of the data” (p. 2). This subsection varies in length across the seven data papers, suggesting that they differ in emphasis on quality control and view on data quality. Being a highly discipline- and context-specific, multidimensional concept, data quality requires adaptability rather than standardised approaches for interpreting its expressions (Stvilia et al., 2007; Stvilia and Lee, 2024). In the brief Quality Control subsection in JOAD 7:4, quality is implied to be data consistency and as much data-related information as possible: data were checked for (unspecified) inconsistencies and geospatial information was added. In JOAD 9:1, the subsection is more detailed. The paradata describe how the dataset has been subjected to a range of controls that are carefully explicated. Data reliability is closely tied to quality, and for each radiocarbon date, a quality coefficient has been added, derived from the relation between standard deviation and before-present date. The description of how the derivation of the quality is achieved underscores how quality is an objective or scientific variable, “quantif[ied]” and providing “quantitative information” (p. 4). This added coefficient works on a level beyond the acceptance/rejection criteria and offers the (re)user the chance to judge the reliability of the dates in the dataset. The implication is that the dataset includes only information of sufficient quality, information which is also ranked by quality. Credibility is constructed by demonstrating the extent to which a reuser can trust the quality of the data.

Even minor discrepancies can suggest variations in how quality is judged, something which comes across clearly in the EUROEVOL papers (JOAD 5:1, 5:2, 5:3). The three datasets in the papers are focused on the western, temperate parts of Europe, but for somewhat different reasons. In JOAD 5:2 and JOAD 5:3, it is because “the available data are best” there (p. 1; p. 1); in JOAD 5:1, it is because “more site data are published there” (p. 1). Where the former papers suggest a qualitative reason, the latter indicates a quantitative one. This notion that quantity equals quality recurs in the Quality Control subsection, which opens with an almost identical sentence in each of the three papers:

We have adopted a fully inclusive approach to the data collection [including all archaeobotanical/faunal/radiocarbon data] irrespective of

  1. Whether or not they pre-dated the adoption of the current standard methods of sampling, recovery and recording. (JOAD 5:1, p. 3; archaeobotanical data)

  2. The date of publication or original analyst. (JOAD 5:3, p. 3; zooarchaeological data)

  3. The date at which the sample was processed. (JOAD 5:2, p. 3; radiocarbon data)

In each case, the sentence implies that data quantity has been prioritised over data quality. Disregarding minor discrepancies in sentence structure, these three sentences also highlight what the key features are for the respective datasets (and types of data). The quality of archaeobotanical data concerns the methods of sampling, recovery, and recording; for faunal data, quality is mainly tied to the requirements of publications or to the person who analysed the material; and for radiocarbon data, quality depends on how they have been processed. The sentence versions indicate differences in emphasis and highlight how analytical methods have changed over time.

Olaisen (1990) suggested that information quality could be shaped by cognitive authority factors (“how information is perceived”) and technical user-friendliness factors (“what the user is offered”) (p. 96). In expressing data usefulness, including quality, the data paper paradata tend to argue that the associated datasets should be perceived as worthy of the trust afforded a cognitive authority but also that they are user-friendly. In short, that the datasets are useful.

Paradata derive meaning from context. As has been previously pointed out by, e.g. Sköld and Andersson (2025), their significance varies depending on the disciplinary context in which they are produced and reused, whether they are created intentionally or unintentionally, and what purpose they are intended to serve. Additionally, paradata take on different meanings depending on the context in which they are communicated – whether they have been entered into a standardised physical or digital form, extracted from a site photo or data table, or incorporated into a data paper narrative.

As per findings regarding the aim of this study to increase our understanding of data papers as research narratives by examining paradata functions in data papers, first, we found that data papers constitute a citable format for shared data, yet in archaeology, they appear to be cited infrequently. Of the 82 data papers published in two journals featuring archaeological data papers, only seven papers were cited for the use of their associated data at least five times over a twelve-year period. While this number will likely increase over time – both because some datasets may attract interest of reusers only in the long term and because the (cumulative) likelihood of discovery and reuse grows with time – it nevertheless suggests that the academic merits gained from sharing data through data papers currently outweighs the actual interest in data reuse. Making data papers more popular among reusers would seem to require putting more emphasis on publishing papers describing datasets with genuine and broad enough reuse potential rather than data scholars happen to have to share.

Second, in data papers, the descriptions of data creation and management processes, along with their underlying rationale and decisions – the paradata – are often incomplete, for various reasons and in various ways. These descriptions form a narrative, with a textually constructed sender (narrator) and receiver (reader). This narrative context shapes the function of the paradata, framing them linguistically and determining which paradata are included and which are omitted. Far from being neutral, paradata express meaning. These expressive functions emerge through the interplay of descriptive and narrative elements and are inherently interconnected. Like the Vitruvian triad, which describes the interrelated functions of a building in terms of fitness for purpose (utilitas), structural integrity (firmitas), and aesthetic qualities (venustas), the triad of expressive paradata functions in a data paper are interrelated aspects that depend on and affect each other.

Third, in a data paper, a paradata element argues regarding the usefulness of the data, that is, how fit data are for the purpose of scientific analysis, it lays a foundation for such research in terms of credibility, and it conveys an impression to the reader through its tone. And like a building the paradata narrative is ultimately a whole, whose argument for usefulness, foundation for credibility, and tonal impression need to be considered in relation to each other and the larger context. Earlier research has so far tended to focus on enumerating factors affecting the (perceived) credibility and usefulness of information, including that usefulness affects credibility and vice versa (e.g. Van House, 2002; Du et al., 2013). Regarding tone, we posit that it merits closer consideration in future information science research relating to credibility and usefulness of information in general, not only relating to data papers. Huvila (2020) argued previously that regarding credibility criteria, it is important to enquire into how they are used to justify and what. The notion of tone clearly provides practicable means to approach the how question both in the context of data papers but also regarding information, documents and information systems in general.

On a practical note, we are inclined to suggest that authors of paradata in data papers would benefit from considering how to find an appropriate tone to argue for the usefulness of data for reuse within which it has opportunities to be credible. The available paradata, the practices described, audience and anticipated reuse affect the options. If there are few caveats, a neutral or authoritative tone can be helpful. If there are major uncertainties, a cautious tone can help to convey an appropriate sense of potential usefulness and credibility whereas if the practices are complex, arguing for usefulness and credibility might require a pedagogical explanation.

While the small sample size constrains the generalisability of our findings, the consistency of the three expressive paradata functions across all seven papers suggests their significance. Rather than claiming universal applicability, our study highlights a pattern that warrants further investigation. Additionally, by employing a close reading of paradata as part of the research narrative, we have demonstrated a novel analytical approach that reveals new dimensions of paradata itself and suggests a new route for paradata exploration.

Ultimately, data papers are not isolated research outputs offering complete, neutral, and objective descriptions with all the paradata necessary for any form of reuse. Instead, they are embedded within the broader narrative of academic research, where researchers craft trustworthy personas, research findings serve both as marketable, citable products and contributions to human knowledge, and data are strategically hoarded, discarded, or shared. Within this narrative, the paradata in data papers are far from neutral; they carry meaning and function to bolster their creators’ academic reputations.

The authors would like to extend their deepest gratitude to Dr Helena Francke for invaluable contributions to writing up this study and her mentorship for Stefan Ekman in the wonderful world of the information studies discipline. This work has received funding from the European Research Council (ERC) under the European Union’s Horizon 2020 research and innovation programme grant agreement No 818210 as a part of the project CApturing Paradata for documenTing data creation and Use for the REsearch of the future (CAPTURE).

1.

In her PhD thesis on grey literature in archaeology, Donnelly (2016) uses a definition from the 12th International Conference on Grey Literature (2010), adding the word archaeological: “[Archaeological] Grey literature stands for manifold document types produced on all levels of government, academics, business and industry in print and electronic formats that are protected by intellectual property rights, of sufficient quality to be collected and preserved by library holdings or institutional repositories, but not controlled by commercial publishers, i.e. where publishing is not the primary activity of the producing body” (2016, p. 12; quoting Schöpfel, 2011, p. 17).

Anichini
,
F.
and
Gattiglia
,
G.
(
2015
), “
Verso la rivoluzione. Dall'Open Access all'Open Data: la pubblicazione aperta in archeologia
”,
European Journal of Postclassical Archaeologies
, Vol. 
5
,
May
, pp. 
299
-
326
.
Bal
,
M.
(
1997
),
Narratology: Introduction to the Theory of Narrative
, (2nd ed.) ,
University of Toronto Press
,
Toronto
.
Bal
,
M.
 
(Ed.)
(
2004a
),
Introduction, Narrative Theory: Critical Concepts in Literary and Cultural Studies
, Vol. 
I
,
Routledge
,
London
, pp. 
1
-
8
.
Bal
,
M.
 
(Ed.)
(
2004b
),
Introduction, Narrative Theory: Critical Concepts in Literary and Cultural Studies
, Vol. 
IV
,
Routledge
,
London
, pp. 
1
-
7
.
Bassyouny
,
H.
,
Abdelfattah
,
T.
and
Tao
,
L.
(
2020
), “
Beyond narrative disclosure tone: the upper echelons theory perspective
”,
International Review of Financial Analysis
, Vol. 
70
, 101499, doi: .
Birkin
,
J.
(
2020
), “
Institutional metadata and the problem of context
”,
Digital Culture and Society
, Vol. 
6
No. 
2
, pp. 
17
-
34
, doi: .
Bishop
,
L.
and
Kuula-Luumi
,
A.
(
2017
), “
Revisiting qualitative data reuse: a decade on
”,
Sage Open
, Vol. 
7
No. 
1
, doi: .
Boi
,
V.
,
Marras
,
A.M.
and
Santagati
,
C.
(
2015
), “
Open Access and archaeology in Italy: an overview and a proposal
”,
Archäologische Informationen
, Vol. 
38
, pp. 
137
-
147
, doi: .
Borgman
,
C.L.
(
2012
), “
The conundrum of sharing research data
”,
JASIST
, Vol. 
63
No. 
6
, pp. 
1059
-
1078
, doi: .
Börjesson
,
L.
,
Sköld
,
O.
and
Huvila
,
I.
(
2020
), “
The politics of paradata in documentation standards and recommendations for digital archaeological visualisations
”,
Digital Culture and Society
, Vol. 
6
No. 
2
, pp. 
191
-
220
, doi: .
Börjesson
,
L.
,
Sköld
,
O.
,
Friberg
,
Z.
,
Löwenborg
,
D.
,
Pálsson
,
G.
and
Huvila
,
I.
(
2022
), “
Re-purposing excavation database content as paradata: an explorative analysis of paradata identification challenges and opportunities
”,
KULA: Knowledge Creation, Dissemination, and Preservation Studies
, Vol. 
6
No. 
3
, pp. 
1
-
18
,
Article 3
doi: .
Borlund
,
P.
(
2003
), “
The concept of relevance in IR
”,
Journal of the American Society for Information Science and Technology
, Vol. 
54
No. 
10
, pp. 
913
-
925
, doi: .
Brummett
,
B.
(
2019
),
Techniques of Close Reading
, (2nd ed.) ,
SAGE
,
Thousand Oaks, CA
.
Buchanan
,
S.
and
Salako
,
A.
(
2009
), “
Evaluating the usability and usefulness of a digital library
”,
Library Review
, Vol. 
58
No. 
9
, pp. 
638
-
651
, doi: .
Callaghan
,
S.
,
Donegan
,
S.
,
Pepler
,
S.
,
Thorley
,
M.
,
Cunningham
,
N.
,
Kirsch
,
Ault
,
L.
,
Bell
,
P.
,
Bowie
,
R.
,
Leadbetter
,
A.
,
Lowry
,
R.
,
Moncoiffé
,
G.
,
Harrison
,
K.
,
Smith-Haddon
,
B.
,
Wearherby
,
A.
and
Wright
,
D.
(
2012
), “
Making data a first class scientific output: data citation and publication by NERC's environmental data Centres
”,
International Journal of Digital Curation
, Vol. 
7
No. 
1
, pp. 
107
-
113
, doi: .
Candela
,
L.
,
Castelli
,
D.
,
Manghi
,
P.
and
Tani
,
A.
(
2015
), “
Data journals: a survey
”,
JASIST
, Vol. 
66
No. 
9
, pp. 
1747
-
1762
, doi: .
Canfield
,
M.R.
,
Reveal
,
J.
,
Heinrich
,
B.
,
Kaufman
,
K.
,
Keller
,
J.
,
Kingdon
,
J.
,
Kitching
,
R.
,
Kramer
,
K.L.
,
Patton
,
J.L.
and
Perrine
,
J.D.
(
2011
),
Field Notes on Science and Nature
,
Harvard University Press
,
Cambridge, MA
.
Chao
,
T.C.
(
2014
), “
Enhancing metadata for research methods in data curation
”,
Proceedings of ASIS&T-AM
, Vol. 
51
No. 
1
, pp. 
1
-
4
, doi: .
Chao
,
T.C.
(
2015
), “
Mapping methods metadata for research data
”,
International Journal of Digital Curation
, Vol. 
10
No. 
1
, pp. 
82
-
94
, doi: .
Chatman
,
S.
(
1978
),
Story and Discourse: Narrative Structure in Fiction and Film
,
Cornell University Press
,
Ithaca, NY
.
Chavan
,
V.
and
Penev
,
L.
(
2011
), “
The data paper: a mechanism to incentivize data publishing in biodiversity science
”,
BMC Bioinformatics
, Vol. 
12
No. 
Suppl. 15
, p.
S2
, doi: .
Chavan
,
V.
,
Penev
,
L.
and
Hobern
,
D.
(
2013
), “
Cultural change in data publishing is essential
”,
BioScience
, Vol. 
63
No. 
6
, pp. 
419
-
420
, doi: .
Donnelly
,
V.
(
2016
), “
A Study in grey: grey Literature and archaeological Investigation in England 1990 to 2010
”,
[PhD thesis]
,
University of Oxford
,
available at:
 https://ora.ox.ac.uk/objects/uuid:9dc84f2d-af55-4d77-ae18-12fe2eefde1b
Du
,
J.T.
,
Liu
,
Y.-H.
,
Zhu
,
Q.
and
Chen
,
Y.
(
2013
), “
Modelling marketing professionals' information behaviour in the workplace: towards a holistic understanding
”,
Information Research
, Vol. 
18
No. 
1
,
paper 560
.
Faniel
,
I.M.
and
Jacobsen
,
T.E.
(
2010
), “
Reusing scientific data: how earthquake engineering researchers assess the reusability of colleagues' data
”,
Computer Supported Cooperative Work
, Vol. 
19
No. 
3
, pp. 
355
-
375
, doi: .
Faniel
,
I.M.
,
Kriesberg
,
A.
and
Yakel
,
E.
(
2016
), “
Social scientists' satisfaction with data reuse
”,
JASIST
, Vol. 
67
No. 
6
, pp. 
1404
-
1416
, doi: .
Faniel
,
I.M.
,
Frank
,
R.D.
and
Yakel
,
E.
(
2019
), “
Context from the data reuser's point of view
”,
Journal of Documentation
, Vol. 
75
No. 
6
, pp. 
1274
-
1297
, doi: .
Farina
,
A.
,
Marongiu
,
P.
,
Bru
,
M.
and
Borkowski
,
D.
(
2025
), “
When data meets the past: data collection, sharing, and reuse in ancient world studies
”,
Open Information Science
, Vol. 
9
No. 
1
, doi: .
Fisher
,
R.
,
van Staden
,
C.J.
and
Richards
,
G.
(
2020
), “
Watch that tone: an investigation of the use and stylistic consequences of tone in corporate accountability disclosures
”,
Accounting, Auditing and Accountability Journal
, Vol. 
33
No. 
1
, pp. 
77
-
105
, doi: .
Francke
,
H.
(
2008
), “
(Re)creations of scholarly journals: Document and information Architecture in open access journals
”,
[PhD thesis]. University of Borås/Valfrid, available at:
 http://hdl.handle.net/2320/1815
Fu
,
J.
,
Tian
,
L.
,
Zhang
,
C.
and
Li
,
J.
(
2023
), “
Opening research data contributes to the citations of related research articles: evidence from Data in Brief
”,
Learned Publishing
, Vol. 
36
No. 
3
, pp. 
426
-
438
, doi: .
García-García
,
A.
,
López-Borrull
,
A.
and
Peset
,
F.
(
2015
), “
Data journals: eclosión de nuevas revistas especializadas en datos
”,
El Profesional de la Información
, Vol. 
24
No. 
6
, pp. 
845
-
854
, doi: .
Gregory
,
K.
,
Ninkov
,
A.
,
Ripp
,
C.
,
Roblin
,
E.
,
Peters
,
I.
and
Haustein
,
S.
(
2023
), “Tracing data: a survey investigating disciplinary differences in data citation”, in
Quantitative Science Studies
,
Advance publication
, doi: .
Gustafsson
,
G.
(
2013
),
Funerary Practices in the 1990s
,
2.0
,
Svensk nationell datatjänst (distributor)
,
Göteborg
,
available at:
doi: .
Huang
,
P.P.
and
Jeng
,
W.
(
2022
), “
Data paper as a reward? Motivation, consideration, and perspective behind data paper submission
”,
Proceedings of the Association for Information Science and Technology
, Vol. 
59
No. 
1
, pp. 
437
-
441
, doi: .
Huggett
,
J.
(
2012
), “
Promise and paradox: accessing open data in archaeology
”,
Mills, C., Pidd, M. and Ward, E. (Eds)
,
Proceedings of the digital humanities congress 2012
,
Sheffield
,
6–8th September 2012
,
Humanities Research Institute
,
available at:
 https://www.dhi.ac.uk/books/dhc2012/promise-and-paradox/
Huggett
,
J.
(
2016
), “
Digital data realities
”,
Introspective Digital Archaeology
,
available at:
 https://introspectivedigitalarchaeology.wordpress.com/2016/06/29/digital-data-realities/ (
accessed
 26 November 2024).
Huvila
,
I.
(
2018
), “Ecology of archaeological information work”, in
Huvila
,
I.
(Ed.),
Archaeology and Archaeological Information in the Digital Society
,
Routledge
, pp. 
121
-
141
.
Huvila
,
I.
(
2020
), “
Information-making-related information needs and the credibility of information
”,
Information Research
, Vol. 
25
No. 
4
,
paper isic2002
, doi: .
Huvila
,
I.
(
2022
), “
Improving the usefulness of research data with better paradata
”,
Open Information Science
, Vol. 
6
No. 
1
, pp. 
28
-
48
, doi: .
Huvila
,
I.
,
Enwald
,
H.
,
Hirvonen
,
N.
and
Eriksson-Backa
,
K.
(
2019
), “
The concept of usefulness in library and information science research
”,
Information Research
, Vol. 
24
No. 
4
,
paper colis1907
.
Huvila
,
I.
,
Sköld
,
O.
and
Börjesson
,
L.
(
2021
), “
Documenting information making in archaeological field reports
”,
Journal of Documentation
, Vol. 
77
No. 
5
, pp. 
1107
-
1127
, doi: .
Huvila
,
I.
,
Börjesson
,
L.
and
Sköld
,
O.
(
2022a
), “
Archaeological information-making activities according to field reports
”,
Library and Information Science Research
, Vol. 
44
No. 
3
, 101171, doi: .
Huvila
,
I.
,
Börjesson
,
L.
and
Sköld
,
O.
(
2022b
), “
Citing methods literature: citations to field manuals as paradata on archaeological fieldwork
”,
Information Research
, Vol. 
27
No. 
3
, doi: .
Huvila
,
I.
,
Sköld
,
O.
and
Andersson
,
L.
(
2023a
), “Knowing-in-practice, its traces and ingredients”, in
Cozza
,
M.
and
Gherardi
,
S.
(Eds.),
The Posthumanist Epistemology of Practice Theory: Re-imagining Method in Organization Studies and beyond
,
Palgrave MacMillan
, pp.
37
-
69
.
Huvila
,
I.
,
Zengenene
,
D.
,
Sköld
,
O.
and
Andersson
,
L.
(
2023b
), “
Data papers as documentation of research processes and practices
”,
[Conference presentation], IST23 Conference: Information Science Perspectives to Documenting Processes and Practices
,
Uppsala, Sweden
.
Huvila
,
I.
,
Andersson
,
L.
and
Sköld
,
O.
(
2024
), “
Patterns in paradata preferences among the makers and reusers of archaeological data
”,
Data and Information Management
, Vol. 
8
No. 
4
, 100077, doi: .
Huvila
,
I.
,
Andersson
,
L.
,
Sköld
,
O.
and
Liu
,
Y.-H.
(
2025
), “
Data makers' and users' views on useful paradata: priorities in documenting data creation, curation, manipulation and use in archaeology
”,
International Journal of Digital Curation
, Vol. 
19
No. 
1
, pp. 
1
-
24
, doi: .
Jiao
,
C.
and
Darch
,
P.T.
(
2020
), “
The role of the data paper in scholarly communication
”,
Proceedings of the Association for Information Science and Technology. Association for Information Science and Technology
, Vol. 
57
No. 
1
, e316, doi: .
Jiao
,
H.
,
Qiu
,
Y.
,
Ma
,
X.
and
Yang
,
B.
(
2024
), “
Dissemination effect of data papers on scientific datasets
”,
Journal of the Association for Information Science and Technology
, Vol. 
75
No. 
2
, pp. 
115
-
131
, doi: .
Kansa
,
E.C.
and
Kansa
,
S.W.
(
2013
), “
We all know that a 14 is a sheep: data publication and professionalism in archaeological communication
”,
Journal of Eastern Mediterranean Archaeology and Heritage Studies
, Vol. 
1
No. 
1
, pp. 
88
-
97
, doi: .
Kansa
,
E.C.
and
Kansa
,
S.W.
(
2022
), “
Promoting data quality and reuse in archaeology through collaborative identifier practices
”,
Proceedings of the National Academy of Sciences of the United States of America
Vol. 
119
No. 
43
, e2109313118, doi: .
Khazraee
,
E.
(
2019
), “
Assembling narratives: tensions in collaborative construction of knowledge
”,
Journal of the Association for Information Science and Technology
, Vol. 
70
No. 
4
, pp. 
325
-
337
, doi: .
Kintigh
,
K.W.
,
Spielmann
,
K.A.
,
Brin
,
A.
,
Candan
,
K.S.
,
Clark
,
T.C.
and
Peeples
,
M.
(
2018
), “
Data integration in the service of synthetic research
”,
Advances in Archaeological Practice
, Vol. 
6
No. 
1
, pp. 
30
-
41
, doi: .
Knorr-Cetina
,
K.D.
(
1981
), “
The manufacture of knowledge: an essay on the constructivist and contextual nature of science
”,
Elsevier Science and Technology
.
Kratz
,
J.E.
and
Strasser
,
C.
(
2014
), “
Data publication consensus and controversies
”,
[version 3], F1000Research
, Vol. 
3
No. 
94
, p.
94
, doi: .
Kratz
,
J.E.
and
Strasser
,
C.
(
2015
), “
Researcher perspectives on publication and peer review of data
”,
PLoS One
, Vol. 
10
No. 
2
, e0117619, doi: .
Li
,
K.
and
Jiao
,
C.
(
2022
), “
The data paper as a sociolinguistic epistemic object: a content analysis on the rhetorical moves used in data paper abstracts
”,
Journal of the Association for Information Science and Technology
, Vol. 
73
No. 
6
, pp. 
834
-
846
, doi: .
Li
,
K.
,
Greenberg
,
J.
and
Dunic
,
J.
(
2020
), “
Data objects and documenting scientific processes: an analysis of data events in biodiversity data papers
”,
JASIST
, Vol. 
71
No. 
2
, pp. 
172
-
182
, doi: .
Li
,
S.
,
Wang
,
G.
and
Luo
,
Y.
(
2022
), “
Tone of language, financial disclosure, and earnings management: a textual analysis of form 20-F
”,
Financial Innovation
, Vol. 
8
No. 
43
, 43, doi: .
Marwick
,
B.
and
Pilaar Birch
,
S.E.
(
2018
), “
A standard for the scholarly citation of archaeological data as an incentive to data sharing
”,
Advances in Archaeological Practice
, Vol. 
6
No. 
2
, pp. 
125
-
143
, doi: .
McGillivray
,
B.
,
Marongiu
,
P.
,
Pedrazzini
,
N.
,
Ribary
,
M.
,
Wigdorowitz
,
M.
and
Zordan
,
E.
(
2022
), “
Deep impact: a study on the impact of data papers and datasets in the humanities and social sciences
”,
Publications
, Vol. 
10
No. 
4
, p.
39
, doi: .
McKeague
,
P.
,
Corns
,
A.
,
Larsson
,
Å.
,
Moreau
,
A.
,
Posluschny
,
A.
,
Daele
,
K.V.
and
Evans
,
T.
(
2020
), “
One archaeology: a manifesto for the systematic and effective use of mapped data from archaeological fieldwork and research
”,
Information: An International Interdisciplinary Journal
, Vol. 
11
No. 
4
, p.
222
, doi: .
Morgan
,
C.
and
Wright
,
H.
(
2018
), “
Pencils and pixels: drawing and digital media in archaeological field recording
”,
Journal of Field Archaeology
, Vol. 
43
No. 
2
, pp. 
136
-
151
, doi: .
Nusser
,
S.M.
,
Cutcher-Gershenfeld
,
J.E.
,
Mikytuck
,
A.M.
and
Korkmaz
,
G.
(
2021
),
Fostering Data Reusability: Increasing Impact and Ease in Sharing and Reusing Research Data – Workshop Report and Action Steps
, doi: .
Olaisen
,
J.
(
1990
), “Information quality factors and the cognitive authority of electronic information”, in
Wormell
,
I.
(Ed.),
Information Quality: Definitions and Dimensions
,
Graham Taylor
, pp. 
91
-
121
.
Parsons
,
M.A.
and
Fox
,
P.A.
(
2013
), “
Is data publication the right metaphor?
”,
Data Science Journal
, Vol. 
12
No. 
0
, pp. 
WDS32
-
WDS46
, doi: .
Pasquetto
,
I.V.
,
Randles
,
B.M.
and
Borgman
,
C.
(
2017
), “
On the reuse of scientific data
”,
Data Science Journal
, Vol. 
16
No. 
8
, pp. 
1
-
9
, doi: .
Pronk
,
T.E.
(
2019
), “
The time efficiency gain in sharing and reuse of research data
”,
Data Science Journal
, Vol. 
18
No. 
10
, pp. 
1
-
8
, doi: .
Richards
,
J.D.
,
Jakobsson
,
U.
,
Novák
,
D.
,
Štular
,
B.
and
Wright
,
H.
(
2021
), “
Digital archiving in archaeology: the state of the art. Introduction
”,
Internet Archaeology
, Vol. 
58
, doi: .
Rieh
,
S.Y.
(
2010
), “Credibility and cognitive authority of information”, in
Bates
,
M.J.
and
Maack
,
M.N.
(Eds),
Encyclopedia of Library and Information Sciences
, (3rd. ed.) ,
Taylor & Francis
, pp. 
1337
-
1344
, doi: .
Roof
,
J.
(
2020
),
Tone: Writing and the Sound of Feeling
,
Bloomsbury Academic
,
London
.
Rubin
,
V.L.
(
2022
),
Misinformation and Disinformation: Detecting Fakes with the Eye and AI
,
Springer
,
Cham
.
Schöpfel
,
J.
(
2011
), “
Towards a prague definition of grey literature
”,
The Grey Journal
, Vol. 
7
No. 
1
, pp.
5
-
18
.
Schöpfel
,
J.
,
Farace
,
D.
,
Prost
,
H.
and
Zane
,
A.
(
2019
), “
Data papers as a new form of knowledge organization in the field of research data
”,
Knowledge Organization
, Vol. 
46
No. 
8
, pp. 
622
-
638
, doi: .
Shen
,
Y.
(
2016
), “
Research data sharing and reuse practices of academic faculty researchers: a study of the Virginia Tech data landscape
”,
International Journal of Digital Curation
, Vol. 
10
No. 
2
, pp. 
157
-
175
, doi: .
Sköld
,
O.
(
2025
), “The concept of paradata”, in
Huvila
,
I.
,
Andersson
,
L.
,
Friberg
,
Z.
,
Liu
,
Y.-H.
and
Sköld
,
O.
(Eds),
Paradata: Documenting Data Creation, Curation and Use
,
Cambridge University Press
, pp.
11
-
39
.
Sköld
,
O.
and
Andersson
,
L.
(
2025
), “Paradata ‘in the wild’: how and where paradata emerges in research data documentation”, in
Huvila
,
I.
,
Andersson
,
L.
,
Friberg
,
Z.
,
Liu
,
Y.-H.
and
Sköld
,
O.
(Eds),
Paradata: Documenting Data Creation, Curation and Use
,
Cambridge University Press
, pp.
40
-
74
.
Sleath
,
P.
,
Butler
,
R.
and
Bond
,
C.
(
2024
), “
Linking geology and art: observations to interpretations of the Sanetsch fold, Helvetic Alps
”,
EGU24-11416, available at:
 https://doi.org/10.5194/egusphere-egu24-11416 (
accessed
 26 November 2024).
Sobotkova
,
A.
(
2018
), “
Sociotechnical obstacles to archaeological data reuse
”,
Advances in Archaeological Practice
, Vol. 
6
No. 
2
, pp. 
117
-
124
, doi: .
Stockwell
,
P.
(
2014
), “Atmosphere and tone”, in
Stockwell
,
P.
and
Whiteley
,
S.
(Eds.),
The Cambridge Handbook of Stylistics
,
Cambridge University Press
, pp.
360
-
374
.
Stuart
,
D.
(
2017
), “
Data bibliometrics: metrics before norms
”,
Online Information Review
, Vol. 
41
No. 
3
, pp. 
428
-
435
, doi: .
Stvilia
,
B.
and
Lee
,
D.J.
(
2024
), “
Data quality assurance in research data repositories: a theory-guided exploration and model
”,
Journal of Documentation
, Vol. 
80
No. 
4
, pp. 
793
-
812
, doi: .
Stvilia
,
B.
,
Gasser
,
L.
,
Twidale
,
M.B.
and
Smith
,
L.C.
(
2007
), “
A framework for information quality assessment
”,
Journal of the American Sociaety for Information Science and Technology
, Vol. 
58
No. 
12
, pp. 
1720
-
1733
, doi: .
Thelwall
,
M.
(
2020
), “
Data in Brief: can a mega-journal for data be useful?
”,
Scientometrics
, Vol. 
124
No. 
1
, pp. 
697
-
709
, doi: .
Van House
,
N.A.
(
2002
), “
Trust and epistemic communities in biodiversity data sharing
”,
Proceedings of the 2nd. ACM/IEEE-CS Joint Conference on Digital Libraries
,
ACM Press
, pp. 
231
-
239
, doi: .
Wallis
,
J.C.
,
Rolando
,
E.
and
Borgman
,
C.L.
(
2013
), “
If we share data, will anyone use them? Data sharing and reuse in the long tail of science and technology
”,
PLoS One
, Vol. 
8
No. 
7
, e67332, doi: .
Walters
,
W.H.
(
2020
), “
Data journals: incentivizing data access and documentation within the scholarly communication system
”,
Insights
, Vol. 
33
, pp. 
1
-
20
, doi: .
Wang
,
X.
,
Duan
,
Q.
and
Liang
,
M.
(
2021
), “
Understanding the process of data reuse: an extensive review
”,
Journal of the Association for Information Science and Technology
, Vol. 
72
No. 
9
, pp. 
1161
-
1182
, doi: .
Wilson
,
P.
(
1983
),
Second-hand Knowledge: An Inquiry into Cognitive Authority
,
Greenwood Press
,
Westport, CN
.
Yoon
,
A.
and
Lee
,
Y.Y.
(
2019
), “
Factors of trust in data reuse
”,
Online Information Review
, Vol. 
43
No. 
7
, pp. 
1245
-
1262
, doi: .
Published by Emerald Publishing Limited. This article is published under the Creative Commons Attribution (CC BY 4.0) licence. Anyone may reproduce, distribute, translate and create derivative works of this article (for both commercial and non-commercial purposes), subject to full attribution to the original publication and authors. The full terms of this licence may be seen at http://creativecommons.org/licences/by/4.0/legalcode

or Create an Account

Close Modal
Close Modal