Skip to article sections
Purpose

This study examines how data processing, analysis, documentation and sharing emerge across research projects, and why variation in these practices cannot be explained by discipline alone.

Design/methodology/approach

The study is based on 28 semi-structured interviews with established researchers purposively selected across macro-level disciplines. The data were analysed using theory-informed qualitative content analysis. Detlor’s (2010) information management (IM) framework was operationalised to structure the data lifecycle and actor levels, while Knorr Cetina’s (1999) theory of epistemic cultures was used to interpret project-level variation.

Findings

Research data took shape through interconnected processing, analysis, documentation and sharing phases, where earlier phases shaped later traceability and reuse. Sharing ranged from open to controlled forms. Variation occurred both across and within macro-level disciplinary contexts, suggesting that discipline alone is insufficient to account for variation in research data management (RDM) practices and that these practices should instead be examined in relation to project-specific epistemic configurations and organisational conditions. More communitarian configurations involved shared infrastructures, standardised practices and established sharing; hybrid configurations combined shared standards with project-specific documentation and controlled sharing; and more individualised configurations relied on local practices and negotiated sharing. The study also identifies tensions between institutional criteria for appropriate data management and researchers’ project-specific needs.

Practical implications

RDM support should combine common principles with project-specific guidance. Support should identify the relevant lifecycle stage, examine how data and research objects are constructed, clarify what different actors consider an adequate solution, and agree what can be shared, with whom and in what form. A single coordinating point of contact could help identify project-specific needs and bring together relevant RDM, IT, legal and data protection expertise.

Originality/value

The study combines IM and epistemic cultures to analyse RDM as a lifecycle shaped by organisational coordination objectives and epistemically differentiated project practices. It shows why support for documentation and sharing should be aligned with project-specific knowledge production practices.

Open science and research data management (RDM) policies emphasise the systematic organisation, documentation and sharing of research data to support transparency, reproducibility and reuse (e.g. European Commission, n.d.; Federation of Finnish Learned Societies, 2025; UNESCO, 2021). Many researchers also hold positive views of openness, although this is not necessarily reflected in everyday research practices (e.g. Tenopir et al., 2020). This tension cannot be explained solely by barriers such as workload, time constraints and legal uncertainty or by facilitators such as recognition, incentives and support services. Understanding it also requires examining how data, documentation and sharing are constituted within different knowledge-production practices.

Science and technology studies (STS) and data studies have shown that the ideal of open sharing and reuse, formulated largely from the premises of data-intensive research, cannot be unproblematically generalised to all research (e.g. Barrocas Ferreira and Borges, 2022; Leonelli, 2020; Reichmann, 2022). Different modes of knowledge production, instruments, divisions of labour and social organisation may generate different logics for data formation, documentation and sharing.

In this study, I examine how research data lifecycles take shape through data processing and analysis, documentation and sharing. I combine Detlor’s (2010) information management (IM) framework with Knorr Cetina’s (1999) perspective on epistemic cultures: the former is operationalised to structure lifecycle phases and actor levels, including organisational and support-service conditions, while the latter helps interpret how and why practices differentiate across projects. Since such differentiation does not necessarily follow the boundaries of faculties, departments or disciplines, but may relate to uses of data, research designs, knowledge production practices and local organisational contexts (Leonelli, 2020; Smith-Doerr et al., 2016), the research project is a critical level of analysis. By “research project”, I refer to individual research undertakings in which data are produced, processed, documented and potentially shared.

This study is based on interviews with leading and established researchers, most of whom correspond broadly to the R4 career stage and the rest to the R3 career stage in the European Framework for Research Careers (EURAXESS, n.d.). It examines how research data, documentation and sharing are constituted across different projects and disciplines, and how these differences relate to macro-level disciplinary contexts, epistemic configurations manifested at the project level and organisational conditions. In addition, the study examines tensions between researchers’ research-related objectives and the role-specific objectives of the organisation and support services in RDM. I ask five related questions. The first three concern the main lifecycle phases examined in the study, while the fourth and fifth address cross-cutting epistemic and organisational dimensions:

RQ1.

How do research data take shape in projects as part of the construction of the research object?

RQ2.

What kinds of documentation practices emerge in research projects, and how are they shaped by organisational conditions, support-service interfaces and actor-level epistemic practices?

RQ3.

On what grounds, under what conditions and through what mechanisms is data shared? How is sharing connected to the nature of data and to documentation practices?

RQ4.

How are differences in the nature of research data, documentation practices and conditions for sharing related to macro-level discipline and project-level epistemic configurations?

RQ5.

How do the institutionally typical objectives of organisations and support services converge with or diverge from researchers’ own research-related objectives in data management practices?

This literature review brings together RDM research on data lifecycle practices and STS and data studies research on how data are constituted within knowledge-production processes. RDM literature has examined practices particularly in relation to policies, infrastructures, competencies and support needs, including barriers and conditions affecting documentation, sharing and reuse across disciplines (e.g. Borghi and Van Gulick, 2021; Klingner et al., 2023; Rantasaari, 2021; Wiley, 2022). STS and data studies, in turn, emphasise how data take shape through research designs, practices and the construction of research objects, with consequences for subsequent documentation and sharing (e.g. Barrocas Ferreira and Borges, 2022; Leonelli, 2020; Reichmann, 2022).

In STS and data studies, data are not treated as neutral or ready-made objects but as shaped by research designs, questions and conceptual premises. This is compatible with Knorr Cetina’s (1999) argument that research objects are constructed within epistemic cultures, where instruments, social organisation and modes of empirical work shape what can become an object of knowledge. Bowker (2005, p. 184), Gitelman and Jackson (2013, pp. 1–3) and Leonelli (2020) similarly question the notion of “raw data”, emphasising that data emerge through choices, processing and packaging. Darch (2018) further argues that reproducibility cannot be reduced to providing data and code, because it depends on the broader organisation of knowledge production, evaluation and sharing.

Because data are formed through situated research practices, documentation becomes central to their later intelligibility, sharing and reuse. Previous RDM research has identified insufficient documentation for external reuse, limited familiarity with metadata creation and uneven use of metadata standards (Akers and Doty, 2013; Cheung et al., 2022; Senft et al., 2022). Field-specific challenges nevertheless differ: humanities and social sciences often face difficulties documenting context, provenance and interpretive complexity (Drucker, 2011; Edmond and Lehmann, 2021; Tóth-Czifra, 2019), whereas natural and biomedical sciences are more often shaped by instrument-based data production, technical heterogeneity and requirements for metadata and minimum information standards (Herres-Pawlis et al., 2022; Mittal et al., 2022; Syn and Kim, 2022).

Documentation also makes transformations between data collection, processing, analysis and reuse traceable. Borgman et al. (2014) and Pujol Priego et al. (2022) emphasise that documentation from earlier lifecycle stages is needed to understand what has been done to the data and thus to interpret and reuse them at later stages. At the same time, documentation sufficient for internal use may not be sufficient for external reuse, which can require additional documentation and negotiation (Borgman et al., 2014; Pujol Priego et al., 2022; Wallis et al., 2013). Recent work on paradata, or information about how data are created, processed and used, similarly emphasises the importance of such process information for reusability, while its relevance, sufficiency and disclosability remain context-dependent (Huvila, 2026).

Empirical RDM research has often examined disciplinary differences in data sharing through factors that support or constrain it. In natural, medical and data-intensive contexts, funder and journal requirements, repositories, standards and research infrastructures are important conditions, although their adoption remains uneven (Chen and Wu, 2017; Devare et al., 2023). In the humanities, social sciences and research involving sensitive personal data, ethical, data protection, ownership and contextual issues are recurrent constraints (Van den Eynden, 2016; Tóth-Czifra, 2019).

Barrocas Ferreira and Borges (2022), Leonelli (2020) and Reichmann (2022) criticise the assumption that open-sharing ideals derived from data-intensive research can be generalised to all research. Huvila and Sinnamon (2024) distinguish data-driven, non-data and data-decentred research cultures that differ in how materials are understood as data and how sharing is integrated into research. From an epistemic-cultures perspective, sharing is therefore selective and conditional. Pujol Priego et al. (2022) show that disciplinary epistemic cultures and individual incentives influence what is shared, with whom and when, while research on reuse and the transfer of context emphasises the role of trusted colleagues, networks and explanations from data producers (Borgman et al., 2014; Wallis et al., 2013).

Discipline alone is insufficient for understanding how processing and analysis, documentation and sharing are connected across research projects. Projects may differ in research objects, methods, divisions of labour, infrastructures and organisational contexts even within the same field. Leonelli (2020) emphasises that data practices are also use-specific, depending on the goals, commitments and tools through which data are interpreted in particular research situations. Smith-Doerr et al. (2016), in their study of the chemical sciences, show that shared disciplinary norms may nevertheless be enacted differently across laboratories, research centres and industry settings. Malazita et al. (2020) similarly describe laboratories as epistemic infrastructures that produce research objects, subjects and legitimate knowledge practices.

Knorr Cetina’s (1999) notion of epistemic cultures offers a conceptual tool for interpreting such variation. In this study, it is used to examine how configurations of data, research designs, infrastructures and divisions of labour shape processing and analysis, documentation and sharing across the research data lifecycle. Knorr Cetina’s (2007) later notion of knowledge cultures broadens this perspective to organisational and institutional environments that sustain, regulate and constrain knowledge practices.

The present study therefore combines RDM, STS and data studies perspectives by examining how processing and analysis shape the conditions for documentation and sharing, and how documentation and sharing function both as RDM practices and as parts of knowledge production. Detlor’s (2010) IM framework is operationalised to structure the research data lifecycle and analytical levels used in this study, while Knorr Cetina’s (1999) epistemic cultures perspective helps interpret project-level epistemic variation.

In this study, research data refer to materials collected, produced, observed, simulated or compiled in research, which may serve as evidence and whose meaning is constituted through the research design and interpretive context. Although data are never completely “raw” in a strict sense (Bowker, 2005, p. 184), I use the term “raw data” to refer to materials at an early stage of the research process, before they are processed into an analysable form. Following Briney (2015, pp. 16–17), I understand RDM as practices through which data are made easier to find, understand and use in both ongoing and future research. I use Detlor’s (2010) IM framework to structure three activities in the research data lifecycle. Creation and acquisition are operationalised here through data processing and analysis, organisation through documentation and distribution through sharing. Knorr Cetina’s (1999) perspective on epistemic cultures complements this lifecycle structure by helping to interpret why these practices vary across projects rather than following formal disciplinary, organisational or lifecycle boundaries in a straightforward way.

According to Detlor (2010), IM aims to help organisations and people access, process and use information efficiently and effectively. His process-oriented model describes information as moving through creation, acquisition, organisation, storage, distribution and use and examines IM from organisational, library and personal perspectives. The organisational perspective treats information as a strategic resource whose management supports organisational objectives. The library perspective concerns the acquisition, organisation, storage, retrieval, access and dissemination of information collections, while the personal perspective concerns how individuals manage information to advance their own objectives.

In this study, the organisational perspective refers to the university and its external operating environment. Detlor’s library perspective is broadened into a research support ecosystem that includes, for example, the library, research IT and legal services. I use the term “actor level” for the personal perspective because the analysis also includes formal and informal research groups.

Knorr Cetina (1999, p. 3) defines epistemic cultures as constellations of practices, arrangements and mechanisms through which the sciences construct their objects, instruments and communal forms. She illustrates their variation through high-energy physics and molecular biology, which differ in the construction of research objects, technical arrangements and social organisation. High-energy physics is characterised by sense-inaccessible objects and collective division of labour, whereas molecular biology centres more strongly on concrete objects and relatively independent laboratory researchers as epistemic subjects (Knorr Cetina, 1999, pp. 159–191, 216–240). In both, however, research objects are detached from their original contexts and transformed into analysable objects (Knorr Cetina, 1999, pp. 32–33). Smith-Doerr et al. (2016) further show that shared epistemic norms may be enacted differently within the same broad discipline across local and organisational settings.

I use the term epistemic configuration as an analytical term for patterns manifested at the project level, rather than treating projects as bounded epistemic cultures. This distinction reflects both that different lifecycle phases within the same project may display different epistemic features and that the analysis does not establish whether the observed features are specific to a particular project or extend to wider research communities, fields or organisational settings.

In the analysis, Detlor’s (2010) framework provides the basis for structuring the lifecycle phases and the organisational, research-support and actor levels, while Knorr Cetina’s (1999) perspective helps interpret variation through the construction of research objects, instruments and infrastructures, and the social organisation of work. Accordingly, Sections 5.1.1, 5.2.1 and 5.3.1 examine organisational and support-service conditions, while Sections 5.1.2, 5.2.2 and 5.3.2 examine researcher and research-group practices and their epistemic differentiation.

This qualitative semi-structured interview study draws on interviews with 28 established researchers at career stages R4 and R3 (EURAXESS, n.d.), purposively selected across macro-level disciplines following the OECD (2007) classification. Table 1 presents participants’ ID codes, positions, macro-level disciplines and research fields, reported at a level that describes their research contexts without making individuals unnecessarily identifiable. Half of the interviewees were professors, while the rest held other leading or established research positions.

Of the 44 researchers identified through the university website and invited by email, 28 participated, resulting in a response rate of 64%. Sampling aimed for variation across macro-level disciplines, positions and research fields rather than a balanced disciplinary distribution. Participants were recruited based on discipline and career stage, but the unit of analysis was the research project described in the interviews: the setting in which data were collected, produced, processed, documented and shared.

The interviews were conducted in spring 2025, 5 face-to-face and 23 remotely. They lasted 45–70 min, with an average of 58 min. Participants received the thematic interview guide and privacy notice with the invitation and gave oral informed consent to participation, recording and personal data processing as described in the notice. A risk assessment indicated low risks related to personal data processing, so a Data Protection Impact Assessment (DPIA) was not required.

The interviews covered RDM broadly, including the nature of data, storage, processing and analysis, documentation, preservation, sharing, reuse, and support and training needs. This article focuses on processing and analysis, documentation and sharing. Interviews were recorded and transcribed using a local, secure OpenWhisper-based AI application provided by my home university. After checking, transcripts were manually coded and thematised in NVivo 14 and compiled into Word files by theme and macro-level discipline. ChatGPT (5.x+) supported summarisation of researcher-prepared theme- and discipline-specific compilations and generation of theory-informed questions for the data. All AI-assisted compilations were checked against the original data, and AI was not used for independent coding or producing the findings.

The interview guide was not designed around Detlor’s (2010) IM framework or Knorr Cetina’s (1999) theory of epistemic cultures, but these frameworks were introduced only during analysis. The analysis followed theory-informed qualitative content analysis, drawing on directed content analysis, in which prior theory guides initial analysis while the data may refine and complement interpretation (Hsieh and Shannon, 2005). First, data were coded according to data management practices. Second, Detlor’s organisational, library and personal perspectives were operationalised as organisational, support-service and actor levels, while the selected lifecycle phases were operationalised as processing and analysis (creation and acquisition), documentation (organisation) and sharing (distribution). Third, Knorr Cetina’s perspective was used to interpret variation in relation to the construction of the research object, the nature of data, methods, infrastructures, division of labour and sharing norms. The third phase resulted in an ideal-typical continuum ranging from more communitarian through hybrid to more individualised epistemic configurations. Projects were characterised overall as more communitarian, hybrid or more individualised on the basis of their overall configuration of epistemic features, while individual lifecycle phases could occupy different positions on the continuum.

Observations and interpretations were cross-checked, when necessary, against the original interview recordings and transcripts and theme- and discipline-specific compilations were compared with original excerpts to ensure that the analysis remained grounded in the data. Because this is a qualitative, semi-structured interview study, the data do not allow strong quantitative generalisations about disciplinary profiles. Nor are disciplines used as units of analysis, but rather as contexts. Although the study discusses support services and their rationalities, support-service staff were not interviewed. These should therefore be understood as analytical constructions inferred from researchers’ interviews and the institutionally typical tasks of relevant support-service actors.

5.1.1 Organisation and support services

The analysis shows that the organisation both enables and constrains data processing through infrastructures, ownership arrangements, data protection and ethical requirements. University guidelines, national and international infrastructures, and funder and journal requirements led researchers to consider reproducibility and traceability already during data collection and processing. Here, reproducibility means reproducing results using the same data and methods, while traceability refers to transparent documentation from raw data to research results.

In my data, ownership varied according to project funding and data source, while regulation constrained what personal data could be collected, processed and shared. From Detlor’s (2010) perspective, the organisation thus both enables and directs data acquisition, organisation and reuse.

Support services acted as interpreters and intermediaries of information by assisting researchers particularly with intellectual property rights, agreements and data protection issues, but in the interviews their support appeared partly insufficient. The burden was especially visible in relation to sensitive data, where researchers described data protection issues as time-consuming and difficult to manage.

Perhaps those data protection issues have been a terrible headache, and they have taken an unreasonable amount of time. (…) And yet it is a very great responsibility, what we are supposed to do. (H2, Humanities)

Researchers therefore wished for easily accessible, interactive support beyond basic training, addressing the entire data processing workflow rather than individual documents.

5.1.2 Actors and epistemic cultures

Researchers’ raw data ranged from experimental, observational and compiled materials to computational models and their output files. Processing into an analysable form depended on data type and research design: for example, personal data functioned either as objects of analysis or as means of distinguishing research units. At Detlor’s (2010) actor level, these practices transform raw data into information meaningful to the research questions.

In quantitative, infrastructure-driven projects, processing and analysis were often based on software, quality control and partly automated analysis pipelines, as well as collective or multidisciplinary division of labour. By analysis pipelines, I refer here to established sequences of work steps, software and methods through which raw data are processed and analysed into information that can answer the research questions. By contrast, in many qualitative projects in the humanities and social sciences, processing and analysis were intertwined, and the same researcher could be largely responsible for handling the data. In such cases, the construction of the research object, as defined by Knorr Cetina (1999), proceeded through close reading, note -taking, transcription and iterative writing.

In my data, these practices can be interpreted through Knorr Cetina’s (1999) epistemic cultures perspective by examining how the construction of the research object, technical arrangements and division of labour were reflected in data processing and analysis.

On this basis, more communitarian features appeared especially in fields that use shared infrastructures and established analysis pipelines, such as astronomy, computational genetics and some data analytics projects, where data processing was more strongly embedded in standardised formats, shared tools and collectively readable workflows.

In terms of the analysis, we typically use fairly standard tools. These are written by observatories or developed by people over decades. […] The observatory says, we know this instrument well, we suggest that the standard way of dealing with the data is using this pipeline. (L5, Natural sciences)

The excerpt shows that standardisation was embedded in community-maintained infrastructure rather than devised locally within the project.

More individualised practices were emphasised in quantum optics, ecology and many qualitative projects in the humanities and social sciences, where processing and analysis were more closely tied to locally produced, assembled or interpreted materials, situated interpretation and the researcher’s own experimental or interpretive setting.

Yes, [the analysis] is indeed not very software-intensive, but quite quickly becomes close reading and iterative writing, and that is perhaps my most common research method. (Y1, Social sciences)

Here, the research object was constructed through close reading and iterative interpretation rather than primarily through a standardised analytical pipeline.

Biomedical consortia and laboratory-based fields such as pharmaceutical chemistry and chemistry-related metabolomics occupied intermediate, hybrid positions, combining shared quality norms with project- or group-specific datasets and workflows.

From the perspective of describing the research data lifecycle, standardised analysis pipelines may also support more systematic documentation. By contrast, local solutions rely more on context and tacit knowledge, making their description more dependent on the researcher’s manual documentation.

5.2.1 Organisation and support services

Organisational and external boundary conditions guided documentation through data policy, infrastructures, data management plans, privacy notices and open-sharing requirements. Many researchers had developed documentation practices through project naming conventions, data dictionaries, version control and electronic laboratory notebooks.

From Detlor’s (2010) perspective, the organisation thus set objectives for documentation, but researchers did not always feel that they received sufficient support for implementing them. Several interviewees wanted more concrete, interactive support, for example in describing procedures, quality assurance, recording parameters and linking code with data.

At the same time, some researchers experienced documentation as an established practice. In instrument-based projects, organisational technical infrastructure automatically produced part of the documentation, reducing project uncertainty and supporting reproducibility.

The mass spectrometer itself documents a great deal: the file contains the most critical parameters and often also the details. When it is opened with the right software, one can see directly what kind of analysis it was and what was measured. (L3, Natural sciences)

The excerpt illustrates how, in instrument-based research, part of the documentation is generated by the technical infrastructure as part of the measurement data.

Also, some researchers using more manual documentation considered their current practices sufficient, although these often served internal manageability rather than the external intelligibility expected by the organisation.

5.2.2 Actors and epistemic cultures

In my interview data, documentation took three main forms: metadata, log and parameter data partly produced by instruments; code- and version-control-based documentation; and narrative, manual documentation emphasising context.

In quantitative, infrastructure-driven projects, such as astronomy and metabolomics (the study of small molecules involved in metabolism), documentation was often systematic, partly instrument-produced and code-based. Recording parameters, software versions, folder structures and analysis stages was integral to constructing data and supporting reproducibility.

In other projects, especially qualitative projects in the humanities and social sciences, documentation consisted of manual notes and records of transcription, thematisation, interpretation and context. Here, the purpose of documentation was not primarily external reproducibility but the interpretive construction of the research object and internal manageability within the project.

Documentation was motivated particularly by the continuity of one’s own work, traceability and openness, but constrained by lack of time, lack of expertise and its perceived laboriousness.

More communitarian, hybrid and more individualised epistemic configurations were also visible in documentation practices. In more communitarian projects, such as astronomy, computational genetics and some psychology projects, documentation was tied to standardised procedures, version control, codebooks and shared practices.

If by a codebook you mean a description of what each variable represents, we make one at the latest when we open the data, because otherwise nobody else would be able to use it. (Y5, Social sciences)

Here, documentation was explicitly oriented towards making data intelligible and useable beyond the immediate project.

In more individualised projects, such as ecology, quantum optics and many humanities fields, documentation more often served the researcher’s own project through local records or narrative notes, although some aimed for external intelligibility and reproducibility.

A bit like a lab notebook […] quite sparsely, because I largely do this work on my own. […] [For an external user] I would need to compile something separately. (L6, Natural sciences)

In interviewee L6’s quantum optics project, the paper laboratory notebook primarily supported the researcher’s own experimental work, while making the documentation intelligible to an external user would have required a separate compilation.

In hybrid projects, such as biomedical cohorts and human geography, documentation was shaped by shared standards and metadata practices, but their application remained partly project- and situation-specific.

At Detlor’s (2010) actor level, documentation can be interpreted as the organisation of information that makes data manageable within the project and, where applicable, for later sharing and reuse. From Knorr Cetina’s (1999, pp. 1–25) perspective, documentation also participates in the construction of the research object by delimiting, selecting and making visible what is considered relevant in the project.

Standardised documentation may facilitate later sharing and reuse, whereas local and narrative documentation depends more on research context and may require further work for external use.

5.3.1 Organisation and support services

Data sharing was guided by legal and administrative boundary conditions as well as by the requirements of funders and publishers. The university, major funders and many publishers recommend or require sharing data, or at least metadata, to increase reproducibility, transparency and research impact. From the perspective of Detlor’s (2010) distribution phase, such requirements connect data sharing to the organisation’s information flows.

In this study, sharing was most established in projects in astronomy, computational genetics and molecular plant biology, as well as in some data analytics projects and social science projects in psychology and human geography. Sharing typically took place through discipline-specific or general repositories, data publications, supplementary materials or, in the case of biological materials, on request. In biomedical cohort studies, some education projects and many humanities projects, sharing was more often permission-based, internal or planned for later archiving.

The role of support services was twofold: they were expected to interpret conditions related to intellectual property rights, data protection and documentation and to support practical data sharing. The interviews nevertheless emphasised that researchers’ and support services’ views on data ownership, data protection and the possibilities of sharing did not always coincide, and some researchers felt left alone especially with legal and data protection questions.

The biggest help that could happen is to start getting some consensus around this because every lawyer you ask or every ethics committee you ask gives a different answer to the question of what data can we actually share. (L1, Natural sciences)

Many researchers therefore wished for more active and interactive support and a more consistent line on data sharing, preservation, and further use and reuse. They also wished for a designated contact point and step-by-step guidance, for example on documentation, anonymisation and preservation.

5.3.2 Actors and epistemic cultures

Based on the interviews, researchers usually shared either raw data supplemented with analysis code and variables or processed data. By contrast, notes, paper laboratory notebooks, sensitive data that were difficult to anonymise or had been created before the GDPR, and large datasets often remained for researchers’ own use. If the actual data could not be shared, metadata or identifiers often could. From Detlor’s (2010) distribution perspective, data were often seen not only as internal project resources but also as institutionally valuable, although organisational constraints could limit external sharing.

Regardless of discipline, sharing was motivated by advancing science, collaboration, resource saving, visibility, transparency and reproducibility. From Detlor’s (2010) perspective, these motives frame sharing as supporting knowledge production, dissemination and collaboration.

Well, the benefit is of course that not all wisdom resides with us. So, if we think about what causes type 1 diabetes, and there is no preventive treatment for it, it does not matter whether I discover it or someone else does. Of course, it is good if someone discovers it, and if our data are useful for that, then good. (LT1, Medical and health sciences)

Sharing was framed as contributing to a collective research problem rather than primarily as compliance with an openness requirement.

Some researchers were also motivated by evaluation practices that recognised data production and sharing as meriting research outputs, connecting sharing to researchers’ research-related and professional objectives (see Detlor, 2010). Sharing was also constrained by practical resources, including workload, available time, technical expertise and inadequate infrastructures, as well as by concerns that data producers might not receive sufficient recognition. The latter illustrates an asymmetry in epistemic credit, in which some researchers risk remaining background producers of data while others may gain merit from data production or reuse. This resonates with Knorr Cetina’s (1999, pp. 192–215) discussion of the tension between communitarian and actorial orders, where collective knowledge production coexists with individual- and group-level interests.

In more communitarian projects, such as astronomy, computational genetics and some data analytics projects, sharing formed part of established knowledge-production practices and relied on infrastructures commonly used within the research community. In one data analytics project, the researcher described sharing as a reciprocal foundational principle of the field:

Our field rests on the principle that data can be shared. […] You want someone else to be able to use your data because you want to use other people’s data. […] In our field, it is very common for people to share almost everything. (T1, Data analytics)

Data were published, for example, in the Hugging Face repository, which the researcher described as a shared ecosystem in the field for both data sharing and reuse.

In more individualised projects, such as qualitative projects in the humanities, social sciences and health sciences, data were more closely tied to the researcher’s interpretive context and project-specific conditions, making sharing case-specific, permission-based or request-based.

I cannot share them in a way that they would just be there and any researcher [could access them], but of course I could share them upon request, for example if a question arose about how these results were actually reached. […] But it would have to go through me. (Y6, Social sciences)

Sharing was not ruled out in principle, but the identifiability of the data and the need to preserve the research context made open, independent reuse problematic and shifted sharing towards researcher-mediated, case-specific access.

In hybrid projects, such as biomedical cohort and consortium studies and some education projects, standards and collaborative structures enabled sharing, but it remained controlled, conditional and dependent on anonymisation, permissions or repository-specific requirements.

From the perspective of epistemic conditions, the possibilities for sharing and reuse were thus linked to the projects’ knowledge-production processes: the ways in which data were processed, analysed and documented shaped how easily they could subsequently be shared and reused.

5.4.1 Summary of findings in light of the theoretical frameworks

The findings show how processing and analysis, documentation and sharing varied across projects and disciplines. Detlor’s (2010) IM framework was operationalised to structure the analysis of lifecycle phases and analytical levels, while Knorr Cetina’s (1999) epistemic cultures perspective helped interpret variation in project-level epistemic configurations.

Table 2 summarises the ideal-typical continuum used to interpret project-level epistemic configurations across the research data lifecycle.

The lifecycle phases were interdependent but not necessarily aligned. In interviewee T2’s multidisciplinary data analytics projects, processing and documentation drew on field-specific standards, shared directory structures, README files and GitHub or GitLab version control, and analyses were designed to be reproducible. Sharing was more conditional: the data were often owned by collaborating partners and could contain sensitive personal or register data, so sharing was limited by ownership, access rights and data protection requirements. Thus, more communitarian analysis and documentation did not necessarily lead to equally open sharing. The analysis of the projects described by T2 therefore illustrates why Table 2 should be used to examine lifecycle phases separately rather than assume that a project occupies the same position across the continuum in all phases: No single standardised or open practice therefore determined the project’s overall epistemic configuration.

Variation also occurred within macro-level disciplines and more specific fields. In educational research, cases Y3 and Y6 illustrate this clearly. In Y3’s longitudinal study involving several universities, data management rested on collective division of labour and shared procedures: survey and register data were harmonised through shared forms, data-merging instructions, codebooks, metadata descriptions and version control. The organisational, research-support and actor levels intersected through contractual and data protection arrangements, expert support and shared data processing and use practices. Overall, the project was a hybrid positioned towards the more communitarian end of the continuum because external access required a research plan and steering-group approval. By contrast, Y6’s video data project constructed the research object through context-dependent interpretation, selective transcription and data sessions. The identifiability and multimodality of the videos made anonymisation problematic and sharing researcher-mediated and request-based, while data protection requirements and legal uncertainty were experienced as burdensome. Organisational and research-support conditions therefore interacted differently with Y6’s local, interpretive process than with Y3’s more collectively organised and standardised setting.

Conversely, similar communitarian features appeared across otherwise distant fields. Both molecular plant biology and digital language research relied on data and infrastructures shared within their research communities and made data or methods available through disciplinary repositories, GitHub or Hugging Face.

Together, these findings suggest that project-specific configurations of research objects, methods, division of labour and infrastructures were more important in accounting for variation than disciplinary labels alone. This is consistent with Smith-Doerr et al. (2016), who show that apparently coherent disciplinary epistemic practices can be enacted differently in local organisational contexts, and with Malazita et al. (2020), who similarly foreground situated research environments and infrastructures as constitutive of epistemic objects, subjects and practices.

Variation was also shaped by funder, journal and other organisational requirements, which could make local practices more standardised, documented or shareable. Epistemic practices and project-level configurations should therefore not be treated as fixed or independent of organisational change: such requirements may gradually become part of field- or project-specific data practices, but their effects are mediated by research objects, data types, infrastructures and divisions of labour.

The next section examines how these differences relate to organisational and support-service requirements and actors’ views of appropriate data management, drawing selectively on organisational and knowledge-culture concepts as interpretive resources rather than additional coding frameworks.

5.4.2 Information management and epistemic cultures in the university context

Detlor’s (2010) and Knorr Cetina’s (1999) frameworks complement each other but also reveal a tension in how the research data lifecycle is understood. Detlor’s IM framework conceptualises it as a strategic and coordinated organisational information process, whereas Knorr Cetina’s perspective on epistemic cultures helps interpret why it is realised differently across projects and why accepted forms of knowledge, evidence and practice may differ from organisational rationalities. Research projects and the support services that frame them can therefore be interpreted as operating within a broader university knowledge culture that sustains, regulates and constrains epistemic practices (Knorr Cetina, 2007).

These perspectives point to different rationalities. Rather than functioning as a unified IM organisation, the university may be understood as a polycentric knowledge environment in which relatively autonomous subsystems are loosely coupled and draw on partly different institutional logics (Friedland and Alford, 1991; Weick, 1976). Their objectives, time horizons and understandings of sufficient data management and acceptable evidence may therefore diverge. In the present study, some researchers expected legal services to help enable lawful sharing under controlled conditions, whereas they perceived legal advice as emphasising risk minimisation and compliance with intellectual property and data protection requirements. Practical alignment thus takes place at the interfaces between research services, legal services, the library, IT services and research projects.

Table 3 is an analytical ideal-typical framework. Its “possible objectives” and “acceptable evidence” are informed partly by researchers’ accounts and partly by the institutionally typical tasks of administrative and support-service actors. Because support-service staff were not interviewed, the table should not be read as a direct empirical description of individual actors’ objectives, but as an interpretive structuring of role-specific rationalities.

The table shifts attention from the availability of resources, guidelines and support to how the IM lifecycle is shaped by role-specific rationalities, criteria for successful practice and forms of acceptable evidence (Friedland and Alford, 1991; Knorr Cetina, 2007). A researcher may assess a project through methods and analyses reported in a peer-reviewed article, while the library may emphasise a well-prepared DMP, metadata and FAIR principles. In a loosely coupled university context (Weick, 1976), formally established policies and support processes may therefore leave unresolved practical problems of documentation, anonymisation or sharing. For example, some researchers perceived DMP questions and guidelines as poorly aligned with the practical needs of their projects, while others found DMPs useful for structuring data management and collaboration. Researchers also sometimes questioned the value of data sharing when it was unlikely to support meaningful reuse or serve a meaningful research purpose. Yet role-specific emphases may also overlap when documentation, sharing and the management of intellectual property or data protection risks support the research objectives of the project or broader field.

5.4.3 Relationship to previous research

5.4.3.1 Layered documentation and external reuse

In my data, “sufficient documentation” often meant internal intelligibility within the project, revealing a gap between documentation sufficient for project members and for external reuse. Extending local documentation, such as instrument logs or notes, for external users required time and resources and could also be constrained by dataset size or sensitivity. This resonates with previous research emphasising provenance and context as conditions for reuse: Leonelli (2020) highlights the mutability of data across lifecycle phases, while Borgman et al. (2014) emphasise that later stages depend on understanding preceding processing steps. Huvila (2026) similarly shows that paradata are context-dependent and not always disclosable, supporting layered sharing and controlled access. The present study further shows that this gap varied across project-level epistemic configurations: more standardised documentation in projects displaying more communitarian features supported sharing, whereas local and narrative documentation in projects displaying more individualised features made external reuse more dependent on contextual knowledge.

5.4.3.2 Established and negotiated sharing

Although universities, funders and publishers encourage or require the open sharing of data, or at least metadata, data in my study were often shared primarily within the project. Beyond the project, sharing was directed mainly to collaborators or researchers in the same field, either openly or under controlled conditions. This aligns with previous research showing that sharing is shaped by conventions and circles of trust and may be routine in some cultures but negotiated case by case in others (Darch, 2018; Pujol Priego et al., 2022; Reichmann, 2022). Similarly, open sharing was most established in more communitarian configurations in my data, whereas sharing in hybrid and more individualised configurations was often request-based. The need for producer explanations and selective documentation in transferring context has also been emphasised by Borgman et al. (2014) and Wallis et al. (2013), while Barrocas Ferreira and Borges (2022) caution against one-size-fits-all assumptions of open science.

Routine data sharing was also linked to researchers’ commitment to open science and to funder and journal requirements. Across project-level epistemic configurations, many researchers also wanted interactive, project-specific support for planning and implementing sharing, consistent with Pujol Priego et al.’s (2022) call for intermediary and coordinating support.

5.4.4 Theoretical implications

Combining Detlor’s (2010) IM framework with Knorr Cetina’s (1999) theory of epistemic cultures conceptualises the research data lifecycle as simultaneously organisationally structured and epistemically differentiated. Documentation makes this especially visible: its purpose, sufficiency and burden vary according to whether it supports the interpretive construction of the research object, infrastructure-based traceability or external sharing and reuse. This helps explain why positive attitudes towards openness do not necessarily lead to systematic documentation or extensive data sharing. The study therefore refines critiques of universal openness: beyond the legal and ethical limits captured by “as open as possible, as closed as necessary”, the meaningfulness of sharing also depends on the epistemic conditions of research (e.g. Khan et al., 2024; Reichmann, 2022).

The analysis further extends differentiation beyond research projects to the organisational setting: appropriate RDM may be defined differently across project-level epistemic configurations and between research projects and support services. Detlor’s (2010) IM framework makes visible organisational coordination aims, while Knorr Cetina’s (1999) epistemic cultures perspective helps explain project-level epistemic differentiation. At the same time, the broader analysis suggests that assessments of appropriate RDM within university support services may reflect role-specific rationalities, objectives and criteria of acceptable evidence. Overall, the analysis suggests that the research data lifecycle in the university context is not a uniformly coordinated organisational process, but is shaped by organisational coordination aims, project-level epistemic differentiation and the interfaces between research projects and support services.

In this study, data did not appear as a ready-made starting point for research but took shape as part of the construction of the research object through selection, processing and analysis. This process was shaped by the projects’ methods, instruments, infrastructures and division of labour, as well as by the extent to which these knowledge-production practices were standardised or interpretive. Earlier lifecycle stages, in turn, partly conditioned the subsequent possibilities for documentation, sharing and reuse (RQ1).

The documentation of these knowledge-production processes ranged from standardised instrument-, metadata- and code-based records to local, manual and narrative notes. These practices were shaped by both the projects’ knowledge-production practices and organisational requirements, as well as by the interfaces with support services encountered by researchers. The systematicity of documentation was further constrained by available time, expertise and support, even when researchers considered better documentation desirable. Documentation served both organisational data management objectives and the epistemic needs of research, such as the construction of the research object and the traceability of analysis, as well as intelligibility within the project (RQ2).

With regard to sharing, data were shared openly, under controlled conditions, on request, or only within the project, depending on the role of the data in knowledge production; the sharing practices established within the research community; and the conditions imposed by ownership, data protection, funders and publishers. Sharing was also shaped by researchers’ motivations and by the time, expertise and support available to them. The nature of the data and the adequacy of documentation affected how readily the data could be made intelligible and reusable to external users: standardised documentation that was intelligible to outsiders facilitated independent reuse, whereas context-dependent documentation more often required additional explanation and was associated with more conditional or researcher-mediated forms of sharing (RQ3).

The nature of research data, documentation practices and conditions for sharing were shaped differently across projects and disciplines. Three ideal-typical forms can be identified in my data: (1) more communitarian epistemic configurations, in which shared infrastructures, relatively standardised practices and established sharing were emphasised; (2) hybrid configurations, in which shared standards and quality requirements were combined with project- or group-specific datasets and workflows in processing and analysis, while documentation and sharing could remain partly project-specific, conditional or negotiated; and (3) more individualised configurations, in which local and project-specific practices of analysis, documentation and sharing were foregrounded. This variation can be understood through modes of knowledge production, especially the degree of standardisation of practices, the use of infrastructures and the social organisation of work (RQ4).

When considering the objectives of the organisation and support services in relation to those of research projects, Detlor’s IM framework highlights how the organisation seeks to coordinate the research data lifecycle as a strategic information process, whereas Knorr Cetina’s (1999) perspective on epistemic cultures helps to interpret how this lifecycle differs across projects. Juxtaposing the two perspectives makes visible tensions between organisational and support-service understandings of appropriate data management, on the one hand; and project-specific practices of data processing and analysis, documentation and sharing, on the other. These objectives converged when documentation, sharing and risk management also supported the research objectives of the project, but diverged when organisational requirements or support-service emphases were perceived by some researchers as poorly aligned with project-specific knowledge-production practices (RQ5).

In practical terms, this means that because projects differ, RDM support cannot be organised around a one-size-fits-all approach. Common guidelines and principles are still needed, but they should be complemented by interactive, staged and project-specific support.

  1. First, identify the lifecycle stage at which support is needed.

  2. Second, examine how the data and the research object are constructed, which parts of the workflow follow common practices or are project-specific, and what role data sharing plays in the particular research context.

  3. Third, make explicit what different actors consider to constitute an adequate solution.

  4. Fourth, agree on what can be shared, with whom and in what form.

In practice, a single coordinating point of contact could help identify project-specific needs and bring together the RDM, IT, legal, data protection or other expertise required to address them.

The study showed differences in data management practices both between and within macro-level disciplines. However, because of the number and distribution of the interviews, the study does not allow strong conclusions about systematic disciplinary profiles. Instead, it illustrates how practices are shaped by project-specific epistemic configurations and modes of knowledge production, as well as by the role-specific objectives of the university organisation, its operating environment and support services.

Formal ethics review was not required under the Finnish National Board on Research Integrity (TENK) guidelines for ethical review in the human sciences. The study involved voluntary semi-structured interviews with adult professional researchers and did not include research design elements requiring prior ethical review. Participants received information about the study and the processing of personal data through a data protection notice. The interviews were pseudonymised before analysis, direct identifiers were stored separately from the research data, and a risk assessment concluded that the processing involved only low and appropriately managed risks.

During data analysis, ChatGPT (OpenAI, GPT-5.x, Thinking) was used to support the summarisation of researcher-prepared theme- and discipline-specific compilations and to generate theory-informed questions for the data. During manuscript preparation, it was also used for copy-editing, including language refinement, condensation and critical feedback on author-written text.

I would like to express my sincere thanks to Professor Kristina Eriksson-Backa of Åbo Akademi University and Professor Isto Huvila of Uppsala University for reading several drafts of this study and for their valuable comments, suggestions and encouragement throughout the writing process.

Akers
,
K.G.
and
Doty
,
J.
(
2013
), “
Disciplinary differences in faculty research data management practices and perspectives
”,
International Journal of Digital Curation
, Vol. 
8
No. 
2
, pp. 
5
-
26
, doi: .
Barrocas Ferreira
,
B.
and
Borges
,
M.M.
(
2022
), “
The epistemic cultures of the digital humanities and their relation to open science: contributions to the open humanities discourse
”,
Central European Journal of Educational Research
, Vol. 
4
No. 
2
, pp. 
1
-
7
, doi: .
Borghi
,
J.A.
and
Van Gulick
,
A.E.
(
2021
), “
Data management and sharing: practices and perceptions of psychology researchers
”,
PLoS One
, Vol. 
16
No. 
5
, e0252047, doi: .
Borgman
,
C.L.
,
Darch
,
P.T.
,
Sands
,
A.E.
,
Wallis
,
J.C.
and
Traweek
,
S.
(
2014
), “
The ups and downs of knowledge infrastructures in science: implications for data management
”,
Proceedings of the ACM/IEEE Joint Conference on Digital Libraries
, pp. 
257
-
266
, doi: .
Bowker
,
G.C.
(
2005
),
Memory Practices in the Sciences
,
MIT Press
,
Cambridge, MA
.
Briney
,
K.
(
2015
),
Data Management for Researchers: Organize, Maintain and Share Your Data for Research Success
,
Pelagic Publishing
,
Exeter
.
Chen
,
X.
and
Wu
,
M.
(
2017
), “
Survey on the needs for chemistry research data management and sharing
”,
The Journal of Academic Librarianship
, Vol. 
43
No. 
4
, pp. 
346
-
353
, doi: .
Cheung
,
M.
,
Cooper
,
A.
,
Dearborn
,
D.
,
Hill
,
E.
,
Johnson
,
E.
,
Mitchell
,
M.
and
Thompson
,
K.
(
2022
), “
Practices before policy: research data management behaviours in Canada
”,
Partnership: The Canadian Journal of Library and Information Practice and Research
, Vol. 
17
No. 
1
, pp. 
1
-
80
, doi: .
Darch
,
P.T.
(
2018
), “
Limits to the pursuit of reproducibility: emergent data-scarce domains of science
”,
Lecture Notes in Computer Science
, Vol. 
10766
, pp. 
164
-
174
, doi: .
Detlor
,
B.
(
2010
), “
Information management
”,
International Journal of Information Management
, Vol. 
30
No. 
2
, pp. 
103
-
108
, doi: .
Devare
,
M.
,
Arnaud
,
E.
,
Antezana
,
E.
and
King
,
B.
(
2023
), “Governing agricultural data: challenges and recommendations”, in
Towards Responsible Plant Data Linkage: Data Challenges for Agricultural Research and Development
,
Springer
,
Cham
, pp. 
201
-
222
, doi: .
Drucker
,
J.
(
2011
), “
Humanities approaches to graphical display
”,
Digital Humanities Quarterly
, Vol. 
5
No. 
1
, doi: .
Edmond
,
J.
and
Lehmann
,
J.
(
2021
), “
Digital humanities, knowledge complexity, and the five ‘aporias’ of digital research
”,
Digital Scholarship in the Humanities
, Vol. 
36
No. 
Supplement_2
, pp. 
ii95
-
ii108
, doi: .
EURAXESS
(
n.d.
), “
Career development
”,
available at:
 Link to the website (
accessed
 11 June 2026).
European Commission
(
n.d.
), “
Open science in horizon Europe
”,
available at:
 Link to the website (
accessed
 29 May 2026).
Federation of Finnish Learned Societies
(
2025
),
The Declaration For Open Science and Research 2025-2030
,
Federation of Finnish Learned Societies
,
available at:
 Link to the DOI (
accessed
 12 June 2026).
Friedland
,
R.
and
Alford
,
R.R.
(
1991
), “Bringing society back in: symbols, practices, and institutional contradictions”, in
Powell
,
W.W.
and
DiMaggio
,
P.J.
(Eds),
The New Institutionalism in Organizational Analysis
,
University of Chicago Press
,
Chicago, IL
, pp. 
232
-
263
.
Gitelman
,
L.
and
Jackson
,
V.
(
2013
), “Introduction”, in
Gitelman
,
L.
(Ed.), “
Raw Data” is an Oxymoron
,
MIT Press
,
Cambridge, MA
, pp. 
1
-
14
.
Herres-Pawlis
,
S.
,
Bach
,
F.
,
Bruno
,
I.J.
,
Chalk
,
S.J.
,
Jung
,
N.
,
Liermann
,
J.C.
,
McEwen
,
L.R.
,
Neumann
,
S.
,
Steinbeck
,
C.
,
Razum
,
M.
and
Koepler
,
O.
(
2022
), “
Minimum information standards in chemistry: a call for better research data management practices
”,
Angewandte Chemie International Edition
, Vol. 
61
No. 
51
, e202203038, doi: .
Hsieh
,
H.F.
and
Shannon
,
S.E.
(
2005
), “
Three approaches to qualitative content analysis
”,
Qualitative Health Research
, Vol. 
15
No. 
9
, pp. 
1277
-
1288
, doi: .
Huvila
,
I.
(
2026
), “
Navigating limits and extents of open process transparency in research documentation: a paradata perspective
”,
Journal of Documentation
, Vol. 
82
No. 
7
, pp. 
358
-
375
, doi: .
Huvila
,
I.
and
Sinnamon
,
L.S.
(
2024
), “
When data sharing is an answer and when (often) it is not: acknowledging data-driven, non-data, and data-decentered cultures
”,
Journal of the Association for Information Science and Technology
, Vol. 
75
No. 
13
, pp. 
1515
-
1530
, doi: .
Khan
,
S.
,
Hirsch
,
J.S.
and
Zeltzer-Zubida
,
O.
(
2024
), “
A dataset without a code book: ethnography and open science
”,
Frontiers in Sociology
, Vol. 
9
, 1308029, doi: .
Klingner
,
C.M.
,
Denker
,
M.
,
Grün
,
S.
,
Hanke
,
M.
,
Oeltze-Jafra
,
S.
,
Ohl
,
F.W.
,
Radny
,
J.
,
Rotter
,
S.
,
Scherberger
,
H.
,
Stein
,
A.
,
Wachtler
,
T.
,
Witte
,
O.W.
and
Ritter
,
P.
(
2023
), “
Research data management and data sharing for reproducible research: results of a community survey of the German National Research Data Infrastructure Initiative Neuroscience
”,
eNeuro
, Vol. 
10
No. 
2
, doi: .
Knorr Cetina
,
K.
(
1999
),
Epistemic Cultures: How the Sciences Make Knowledge
,
Harvard University Press
,
Cambridge, MA
, doi: .
Knorr Cetina
,
K.
(
2007
), “
Culture in global knowledge societies: knowledge cultures and epistemic cultures
”,
Interdisciplinary Science Reviews
, Vol. 
32
No. 
4
, pp. 
361
-
375
, doi: .
Leonelli
,
S.
(
2020
), “Learning from data journeys”, in
Leonelli
,
S.
and
Tempini
,
N.
(Eds),
Data Journeys in the Sciences
,
Springer
,
Cham
, pp. 
1
-
24
,
available at:
 Link to the website
Malazita
,
J.W.
,
Teboul
,
E.J.
and
Rafeh
,
H.
(
2020
), “
Digital humanities as epistemic cultures: how DH Labs make knowledge, objects, and subjects
”,
DHQ: Digital Humanities Quarterly
, Vol. 
14
No. 
3
, doi: .
Mittal
,
D.
,
Mease
,
R.
,
Kuner
,
T.
,
Flor
,
H.
,
Kuner
,
R.
and
Andoh
,
J.
(
2022
), “
Data management strategy for a collaborative research center
”,
GigaScience
, Vol. 
12
, pp. 
1
-
25
, doi: .
OECD
(
2007
), “
Revised field of science and technology (FOS) classification in the Frascati manual
”,
available at:
 Link to the website (
accessed
 31 May 2026).
Pujol Priego
,
L.
,
Wareham
,
J.
and
Romasanta
,
A.K.S.
(
2022
), “
The puzzle of sharing scientific data
”,
Industry and Innovation
, Vol. 
29
No. 
2
, pp. 
219
-
250
, doi: .
Rantasaari
,
J.
(
2021
), “
Doctoral students' educational needs in research data management: perceived importance and current competencies
”,
International Journal of Digital Curation
, Vol. 
16
No. 
1
, p.
36
, doi: .
Reichmann
,
S.
(
2022
), “
The narrow road to data: data sharing for global, unspecified reuse
”,
SocArXiv preprint
. doi: .
Senft
,
M.
,
Stahl
,
U.
and
Svoboda
,
N.
(
2022
), “
Research data management in agricultural sciences in Germany: we are not yet where we want to be
”,
PLoS One
, Vol. 
17
No. 
9
, e0274677, doi: .
Smith-Doerr
,
L.
,
Croissant
,
J.
,
Vardi
,
I.
and
Sacco
,
T.
(
2016
), “Epistemic cultures of collaboration: coherence and ambiguity in interdisciplinarity”, in
Investigating Interdisciplinary Collaboration: Theory and Practice across Disciplines
,
Rutgers University Press
,
New Brunswick, NJ
, pp. 
65
-
83
, doi: .
Syn
,
S.Y.
and
Kim
,
S.
(
2022
), “
Characterizing the research data management practices of NIH biomedical researchers indicates the need for better support at laboratory level
”,
Health Information and Libraries Journal
, Vol. 
39
No. 
4
, pp. 
347
-
356
, doi: .
Tenopir
,
C.
,
Rice
,
N.M.
,
Allard
,
S.
,
Baird
,
L.
,
Borycz
,
J.
,
Christian
,
L.
,
Grant
,
B.
,
Olendorf
,
R.
and
Sandusky
,
R.J.
(
2020
), “
Data sharing, management, use, and reuse: practices and perceptions of scientists worldwide
”,
PLoS One
, Vol. 
15
No. 
3
, e0229003, doi: .
Tóth-Czifra
,
E.
(
2019
), “
The risk of losing thick description: data management challenges arts and humanities face in the evolving FAIR data ecosystem
”,
HAL-SHS
,
available at:
 Link to the website
UNESCO
(
2021
),
UNESCO Recommendation on Open Science
,
UNESCO
, doi: .
Van den Eynden
,
V.
,
Knight
,
G.
,
Vlad
,
A.
,
Radler
,
B.
,
Tenopir
,
C.
,
Leon
,
D.
,
Manista
,
F.
,
Whitworth
,
J.
and
Corti
,
L.
(
2016
), “
Towards open research: practices, experiences, barriers and opportunities
”,
Wellcome Trust
. doi: .
Wallis
,
J.C.
,
Rolando
,
E.
and
Borgman
,
C.L.
(
2013
), “
If we share data, will anyone use them? Data sharing and reuse in the long tail of science and technology
”,
PLoS One
, Vol. 
8
No. 
7
, e67332, doi: .
Weick
,
K.E.
(
1976
), “
Educational organizations as loosely coupled systems
”,
Administrative Science Quarterly
, Vol. 
21
No. 
1
, pp. 
1
-
19
, doi: .
Wiley
,
C.
(
2022
), “
Research data management: a case study examining aerospace, industrial and mechanical science engineering faculty research practices
”,
Science and Technology Libraries
, Vol. 
42
No. 
3
, pp. 
391
-
398
, doi: .
Published by Emerald Publishing Limited. This article is published under the Creative Commons Attribution (CC BY 4.0) licence. Anyone may reproduce, distribute, translate and create derivative works of this article (for both commercial and non-commercial purposes), subject to full attribution to the original publication and authors. The full terms of this licence may be seen at Link to the terms of the CC BY 4.0 licence.

Data & Figures

Table 1

Interviewed researchers by position, macro-level discipline and research field

IDPositionMacro-level disciplineResearch field
L1Academy Research FellowNatural sciencesMetabolomics in chemistry
T1ProfessorEngineering and technologyData analytics
H1University teacherHumanitiesHistory
Y1ProfessorSocial sciencesInformation studies
L2Senior researcherNatural sciencesPhysiology and genetics
Y2ProfessorSocial sciencesPsychology
LT1ProfessorMedical and health sciencesBiomedicine
H2ProfessorHumanitiesHistory and archaeology
T2ProfessorEngineering and technologyData analytics
L3University lecturerNatural sciencesPharmaceutical chemistry
Y3Senior researcherSocial sciencesEducation
LT2ProfessorMedical and health sciencesNursing science
Y4ProfessorSocial sciencesAccounting and finance
Y5Associate professorSocial sciencesGeography
L4ProfessorNatural sciencesMolecular plant biology
H3Senior researcherHumanitiesCultural history
Y6ProfessorSocial sciencesEducation
Y7Senior researcherSocial sciencesEntrepreneurship
LT3ProfessorMedical and health sciencesDentistry
L5Senior researcherNatural sciencesAstronomy
Y8ProfessorSocial sciencesEducation
L6Senior researcherNatural sciencesQuantum optics
L7ProfessorNatural sciencesEcology and evolutionary biology
H4University lecturerHumanitiesPolitical history
H5University lecturerHumanitiesCultural studies
Y9Senior researcherSocial sciencesEntrepreneurship
H6ProfessorHumanitiesDigital linguistics
Y10Senior researcherSocial sciencesEconomic sociology
Source(s): Author’s own work
Table 2

Ideal-typical continuum of epistemic configurations in research data lifecycle practices

Lifecycle phase/dimensionMore communitarianHybridMore individualised
Processing and analysisShared infrastructures, collective or distributed division of labour, established pipelines; data often treated as a communal resourceShared standards or quality requirements combined with project- or group-specific datasets and bounded multidisciplinary workLocal tools and methods, individual or small-project responsibility; data constituted primarily as the material of one’s own project
DocumentationStandardised, code-, instrument- or pipeline-based documentation integrated into the workflowShared standards combined with contextual or project-specific additionsNarrative, local or manual documentation, often oriented to internal intelligibility and interpretive work
SharingExpected or routine sharing through established infrastructures or repositoriesControlled, conditional or permission-based sharingNegotiated, case-specific or request-based sharing; data often remain internal to the project
Source(s): Author’s own work
Table 3

Actor-specific rationalities, possible objectives and acceptable evidence in research data management

ActorPossible objectiveAcceptable evidence
Research servicesAligning funder requirements, research ethics and open science principles in the data management of research projectsDMPs, reports and other documents approved by funders; agreements; ethical statements and processes; compliance with open science policies
Legal servicesManaging legality, data protection and IPR risks, and minimising legal risks in the processing and sharing of research dataPrivacy notice, DPIA, agreements, access permits, legal conditions for data sharing and reuse
LibrarySupporting data findability, documentation, description, sharing and reusabilityDMP, metadata, README, repository selection, FAIR assessment, data citation and identifier practices
Research ITProviding a secure, controlled and useable technical environmentApproved systems, access rights, information security practices, logs, backups, and sufficient storage, computing and distribution solutions
ResearchersAddressing research questions, scientific autonomy, appropriate use of data, possible sharing and academic meritReporting of methods, analyses and data processing; peer-reviewed publication; internally sufficient project documentation; where appropriate, open, controlled or request-based data sharing

Note(s): DMP = data management plan; DPIA = data protection impact assessment; IPR = intellectual property rights; README = a file that provides essential information about a dataset or software; FAIR = findable, accessible, interoperable and reusable

Source(s): Author’s own work

Supplements

References

Akers
,
K.G.
and
Doty
,
J.
(
2013
), “
Disciplinary differences in faculty research data management practices and perspectives
”,
International Journal of Digital Curation
, Vol. 
8
No. 
2
, pp. 
5
-
26
, doi: .
Barrocas Ferreira
,
B.
and
Borges
,
M.M.
(
2022
), “
The epistemic cultures of the digital humanities and their relation to open science: contributions to the open humanities discourse
”,
Central European Journal of Educational Research
, Vol. 
4
No. 
2
, pp. 
1
-
7
, doi: .
Borghi
,
J.A.
and
Van Gulick
,
A.E.
(
2021
), “
Data management and sharing: practices and perceptions of psychology researchers
”,
PLoS One
, Vol. 
16
No. 
5
, e0252047, doi: .
Borgman
,
C.L.
,
Darch
,
P.T.
,
Sands
,
A.E.
,
Wallis
,
J.C.
and
Traweek
,
S.
(
2014
), “
The ups and downs of knowledge infrastructures in science: implications for data management
”,
Proceedings of the ACM/IEEE Joint Conference on Digital Libraries
, pp. 
257
-
266
, doi: .
Bowker
,
G.C.
(
2005
),
Memory Practices in the Sciences
,
MIT Press
,
Cambridge, MA
.
Briney
,
K.
(
2015
),
Data Management for Researchers: Organize, Maintain and Share Your Data for Research Success
,
Pelagic Publishing
,
Exeter
.
Chen
,
X.
and
Wu
,
M.
(
2017
), “
Survey on the needs for chemistry research data management and sharing
”,
The Journal of Academic Librarianship
, Vol. 
43
No. 
4
, pp. 
346
-
353
, doi: .
Cheung
,
M.
,
Cooper
,
A.
,
Dearborn
,
D.
,
Hill
,
E.
,
Johnson
,
E.
,
Mitchell
,
M.
and
Thompson
,
K.
(
2022
), “
Practices before policy: research data management behaviours in Canada
”,
Partnership: The Canadian Journal of Library and Information Practice and Research
, Vol. 
17
No. 
1
, pp. 
1
-
80
, doi: .
Darch
,
P.T.
(
2018
), “
Limits to the pursuit of reproducibility: emergent data-scarce domains of science
”,
Lecture Notes in Computer Science
, Vol. 
10766
, pp. 
164
-
174
, doi: .
Detlor
,
B.
(
2010
), “
Information management
”,
International Journal of Information Management
, Vol. 
30
No. 
2
, pp. 
103
-
108
, doi: .
Devare
,
M.
,
Arnaud
,
E.
,
Antezana
,
E.
and
King
,
B.
(
2023
), “Governing agricultural data: challenges and recommendations”, in
Towards Responsible Plant Data Linkage: Data Challenges for Agricultural Research and Development
,
Springer
,
Cham
, pp. 
201
-
222
, doi: .
Drucker
,
J.
(
2011
), “
Humanities approaches to graphical display
”,
Digital Humanities Quarterly
, Vol. 
5
No. 
1
, doi: .
Edmond
,
J.
and
Lehmann
,
J.
(
2021
), “
Digital humanities, knowledge complexity, and the five ‘aporias’ of digital research
”,
Digital Scholarship in the Humanities
, Vol. 
36
No. 
Supplement_2
, pp. 
ii95
-
ii108
, doi: .
EURAXESS
(
n.d.
), “
Career development
”,
available at:
 Link to the website (
accessed
 11 June 2026).
European Commission
(
n.d.
), “
Open science in horizon Europe
”,
available at:
 Link to the website (
accessed
 29 May 2026).
Federation of Finnish Learned Societies
(
2025
),
The Declaration For Open Science and Research 2025-2030
,
Federation of Finnish Learned Societies
,
available at:
 Link to the DOI (
accessed
 12 June 2026).
Friedland
,
R.
and
Alford
,
R.R.
(
1991
), “Bringing society back in: symbols, practices, and institutional contradictions”, in
Powell
,
W.W.
and
DiMaggio
,
P.J.
(Eds),
The New Institutionalism in Organizational Analysis
,
University of Chicago Press
,
Chicago, IL
, pp. 
232
-
263
.
Gitelman
,
L.
and
Jackson
,
V.
(
2013
), “Introduction”, in
Gitelman
,
L.
(Ed.), “
Raw Data” is an Oxymoron
,
MIT Press
,
Cambridge, MA
, pp. 
1
-
14
.
Herres-Pawlis
,
S.
,
Bach
,
F.
,
Bruno
,
I.J.
,
Chalk
,
S.J.
,
Jung
,
N.
,
Liermann
,
J.C.
,
McEwen
,
L.R.
,
Neumann
,
S.
,
Steinbeck
,
C.
,
Razum
,
M.
and
Koepler
,
O.
(
2022
), “
Minimum information standards in chemistry: a call for better research data management practices
”,
Angewandte Chemie International Edition
, Vol. 
61
No. 
51
, e202203038, doi: .
Hsieh
,
H.F.
and
Shannon
,
S.E.
(
2005
), “
Three approaches to qualitative content analysis
”,
Qualitative Health Research
, Vol. 
15
No. 
9
, pp. 
1277
-
1288
, doi: .
Huvila
,
I.
(
2026
), “
Navigating limits and extents of open process transparency in research documentation: a paradata perspective
”,
Journal of Documentation
, Vol. 
82
No. 
7
, pp. 
358
-
375
, doi: .
Huvila
,
I.
and
Sinnamon
,
L.S.
(
2024
), “
When data sharing is an answer and when (often) it is not: acknowledging data-driven, non-data, and data-decentered cultures
”,
Journal of the Association for Information Science and Technology
, Vol. 
75
No. 
13
, pp. 
1515
-
1530
, doi: .
Khan
,
S.
,
Hirsch
,
J.S.
and
Zeltzer-Zubida
,
O.
(
2024
), “
A dataset without a code book: ethnography and open science
”,
Frontiers in Sociology
, Vol. 
9
, 1308029, doi: .
Klingner
,
C.M.
,
Denker
,
M.
,
Grün
,
S.
,
Hanke
,
M.
,
Oeltze-Jafra
,
S.
,
Ohl
,
F.W.
,
Radny
,
J.
,
Rotter
,
S.
,
Scherberger
,
H.
,
Stein
,
A.
,
Wachtler
,
T.
,
Witte
,
O.W.
and
Ritter
,
P.
(
2023
), “
Research data management and data sharing for reproducible research: results of a community survey of the German National Research Data Infrastructure Initiative Neuroscience
”,
eNeuro
, Vol. 
10
No. 
2
, doi: .
Knorr Cetina
,
K.
(
1999
),
Epistemic Cultures: How the Sciences Make Knowledge
,
Harvard University Press
,
Cambridge, MA
, doi: .
Knorr Cetina
,
K.
(
2007
), “
Culture in global knowledge societies: knowledge cultures and epistemic cultures
”,
Interdisciplinary Science Reviews
, Vol. 
32
No. 
4
, pp. 
361
-
375
, doi: .
Leonelli
,
S.
(
2020
), “Learning from data journeys”, in
Leonelli
,
S.
and
Tempini
,
N.
(Eds),
Data Journeys in the Sciences
,
Springer
,
Cham
, pp. 
1
-
24
,
available at:
 Link to the website
Malazita
,
J.W.
,
Teboul
,
E.J.
and
Rafeh
,
H.
(
2020
), “
Digital humanities as epistemic cultures: how DH Labs make knowledge, objects, and subjects
”,
DHQ: Digital Humanities Quarterly
, Vol. 
14
No. 
3
, doi: .
Mittal
,
D.
,
Mease
,
R.
,
Kuner
,
T.
,
Flor
,
H.
,
Kuner
,
R.
and
Andoh
,
J.
(
2022
), “
Data management strategy for a collaborative research center
”,
GigaScience
, Vol. 
12
, pp. 
1
-
25
, doi: .
OECD
(
2007
), “
Revised field of science and technology (FOS) classification in the Frascati manual
”,
available at:
 Link to the website (
accessed
 31 May 2026).
Pujol Priego
,
L.
,
Wareham
,
J.
and
Romasanta
,
A.K.S.
(
2022
), “
The puzzle of sharing scientific data
”,
Industry and Innovation
, Vol. 
29
No. 
2
, pp. 
219
-
250
, doi: .
Rantasaari
,
J.
(
2021
), “
Doctoral students' educational needs in research data management: perceived importance and current competencies
”,
International Journal of Digital Curation
, Vol. 
16
No. 
1
, p.
36
, doi: .
Reichmann
,
S.
(
2022
), “
The narrow road to data: data sharing for global, unspecified reuse
”,
SocArXiv preprint
. doi: .
Senft
,
M.
,
Stahl
,
U.
and
Svoboda
,
N.
(
2022
), “
Research data management in agricultural sciences in Germany: we are not yet where we want to be
”,
PLoS One
, Vol. 
17
No. 
9
, e0274677, doi: .
Smith-Doerr
,
L.
,
Croissant
,
J.
,
Vardi
,
I.
and
Sacco
,
T.
(
2016
), “Epistemic cultures of collaboration: coherence and ambiguity in interdisciplinarity”, in
Investigating Interdisciplinary Collaboration: Theory and Practice across Disciplines
,
Rutgers University Press
,
New Brunswick, NJ
, pp. 
65
-
83
, doi: .
Syn
,
S.Y.
and
Kim
,
S.
(
2022
), “
Characterizing the research data management practices of NIH biomedical researchers indicates the need for better support at laboratory level
”,
Health Information and Libraries Journal
, Vol. 
39
No. 
4
, pp. 
347
-
356
, doi: .
Tenopir
,
C.
,
Rice
,
N.M.
,
Allard
,
S.
,
Baird
,
L.
,
Borycz
,
J.
,
Christian
,
L.
,
Grant
,
B.
,
Olendorf
,
R.
and
Sandusky
,
R.J.
(
2020
), “
Data sharing, management, use, and reuse: practices and perceptions of scientists worldwide
”,
PLoS One
, Vol. 
15
No. 
3
, e0229003, doi: .
Tóth-Czifra
,
E.
(
2019
), “
The risk of losing thick description: data management challenges arts and humanities face in the evolving FAIR data ecosystem
”,
HAL-SHS
,
available at:
 Link to the website
UNESCO
(
2021
),
UNESCO Recommendation on Open Science
,
UNESCO
, doi: .
Van den Eynden
,
V.
,
Knight
,
G.
,
Vlad
,
A.
,
Radler
,
B.
,
Tenopir
,
C.
,
Leon
,
D.
,
Manista
,
F.
,
Whitworth
,
J.
and
Corti
,
L.
(
2016
), “
Towards open research: practices, experiences, barriers and opportunities
”,
Wellcome Trust
. doi: .
Wallis
,
J.C.
,
Rolando
,
E.
and
Borgman
,
C.L.
(
2013
), “
If we share data, will anyone use them? Data sharing and reuse in the long tail of science and technology
”,
PLoS One
, Vol. 
8
No. 
7
, e67332, doi: .
Weick
,
K.E.
(
1976
), “
Educational organizations as loosely coupled systems
”,
Administrative Science Quarterly
, Vol. 
21
No. 
1
, pp. 
1
-
19
, doi: .
Wiley
,
C.
(
2022
), “
Research data management: a case study examining aerospace, industrial and mechanical science engineering faculty research practices
”,
Science and Technology Libraries
, Vol. 
42
No. 
3
, pp. 
391
-
398
, doi: .

Languages

or Create an Account

Close subscription notice
Close access options