Purpose

As HR practitioners balance strategic responsibilities with administrative work in daily practice, understanding the role of generative AI (GenAI) in supporting their work becomes critical. This study examines how GenAI, specifically ChatGPT, affects productivity in completing HR tasks, namely routine, junior-level, written deliverables that support HR service delivery. We treat ChatGPT primarily as a support tool for HR tasks and examine whether different ways of using it, raw adoption or iterative editing, are associated with different productivity outcomes.

Design/methodology/approach

A randomized controlled experiment was conducted with 91 final-year business students, simulating routine, junior-level HR tasks across three functional domains: (1) internal communication and employer branding, (2) recruitment, and (3) selection. Participants were randomly assigned to a control or treatment group. Productivity was assessed using three measures: effectiveness (expert-rated deliverable quality), completion time and efficiency (quality per unit of time). A delayed-treatment design enabled within- and between-group comparisons. Tasks were validated by HR practitioners, who also assessed outputs on written expression, content quality and originality as indicators of effectiveness.

Findings

In this experimental setting, ChatGPT utilization was associated with significant productivity improvements: +35.8% in effectiveness, −53.9% in completion time and +391.7% in efficiency, operationalized as expert-rated quality per unit of completion time. No significant differences were observed between usage patterns, suggesting that, in this experimental setting, basic use of GenAI was associated with important gains in HR tasks completion. Delayed access to the tool resulted in high efficiency increases, a pattern we discuss in relation to contrast effects and initial task engagement without tool support.

Practical implications

This study provides actionable insights for HR practitioners on integrating GenAI into operational workflows to enhance output in everyday HR work, where written deliverables are produced under time pressure, amidst ongoing interactions with line managers and employees and in a fragmented manner.

Originality/value

The study offers empirical evidence on the effects of GenAI tools in Human Resource Management practice, extending prior research through a controlled experimental design. By focusing on daily tasks rather than strategic activities, it deepens understanding of GenAI’s role in an overlooked yet important HR work domain currently undergoing transformation.

Human resource management (HRM) is undergoing profound transformation driven by the adoption of advanced technologies, such as artificial intelligence (AI). AI has become vital for streamlining internal processes, enhancing employee efficiency (Budhwar et al., 2022; Nankervis et al., 2021) and reshaping workplaces (Bick et al., 2024). Such technological advancements also impact the daily work life of HR practitioners (Hermann et al., 2025; Malik et al., 2023). However, most existing literature highlights the strategic domain in the context of generative AI (GenAI), often overlooking routine work (Hermann et al., 2025).

What remains insufficiently understood is how these technologies reshape the everyday work and routines of HR practitioners (Wallo and Coetzer, 2023) and affect their productivity (Noy and Zhang, 2023). HR work is enacted in practice through everyday activities, interactions and micro-decisions that shape organizational outcomes (Wallo and Coetzer, 2023). The HRM-as-practice perspective places attention on HR practices, HR praxis and HR practitioners as intertwined and inseparable in everyday organizational settings (Björkman et al., 2014). In this view, HR practices are enacted through HR praxis, conceptualized as the flow of situated activities in context, through which practitioners “go about doing HR work” in real situations, rather than how HR work is formally or abstractly described (Espegren and Hugosson, 2025). Focusing on HR praxis as the enactment of abstract activities also resonates with process perspectives that distinguish between intended and implemented HRM, which may differ substantially (Nishii and Wright, 2008). HR praxis often includes fragmented and time-pressured HR tasks, through which service is delivered to line managers, employees and other stakeholders. Exploring AI is particularly relevant to written HR tasks, and its impact on HR practitioners’ productivity is likely to depend on how it is used within everyday workflows.

Addressing this research gap gains urgency in light of the accelerating integration of GenAI into work environments and responds to growing calls to examine AI in HR practitioners’ everyday work. Empirical studies in other professions, such as consulting (Dell’ Acqua et al., 2023), creative industries (Zhou and Lee, 2024) and pediatric neurology (Karakas et al., 2023), suggest notable productivity gains. However, there is limited evidence that GenAI impacts routine tasks in the HR context (Noy and Zhang, 2023). In particular, we know little about whether GenAI changes the speed and quality with which HR practitioners produce the written outputs that support daily service delivery. Limited research is of concern as HR tasks often involve high volumes of written deliverables such as announcements, templates and onboarding documents that could be significantly affected by GenAI’s capabilities. These outputs form a substantial part of daily HR work. They support consistency and compliance across HR processes as well as coordination with line managers and employees. In producing them, HR practitioners often work under time pressure and interruptions, hence even modest time savings or quality enhancements can enhance their productivity. Without empirical evidence, organizations risk either overestimating or underutilizing the potential of these tools in the HR domain, particularly in routine, junior-level work where such written outputs are produced at scale.

To advance understanding, this study employs a controlled experimental design using a vignette methodology, designed to provide rigorous evidence on the effect of GenAI access in a simulated HR work setting. We focus on routine, junior-level, written HR tasks, through which HR service is delivered, such as internal announcements, recruitment texts, selection materials, policy drafts and other HR documentation. For brevity, we refer to these as HR tasks throughout the paper. Consistent with an HRM-as-practice perspective that foregrounds HR praxis, ChatGPT is treated as a task support tool and examined at the level of enactment, that is the completion of these written deliverables (Björkman et al., 2014; Nishii and Wright, 2008). In this study, we develop three vignettes featuring HR tasks related to internal communication and employer branding, recruitment and selection. We examine how GenAI impacts productivity across three outcome metrics, task completion time, expert-rated deliverable quality (effectiveness) and efficiency. By focusing on routine, junior-level HR tasks, the study examines how GenAI may change everyday HR service delivery and coordination, rather than more senior or strategic HR work. Consistent with this scope, we do not test more complex HR tasks.

This study makes two contributions. First, it extends our understanding of the impact of GenAI on productivity across professions by applying this lens to the HRM field. Our perspective focuses on the often-overlooked, yet important micro-level reality of HR work, characterized by routine, segmented tasks and work overload (Wallo and Coetzer, 2023). By examining GenAI’s impact on HR tasks in a simulated work setting and estimating its effects on speed, quality and efficiency, the study advances understanding of Gen-AI-driven productivity effects.

Second, this work provides new insights into the relationship between productivity gains from GenAI and usage patterns, namely raw adoption and iterative editing. No prior empirical study has explicitly examined this issue in the context of HR tasks. These usage patterns constitute different manifestations of HR praxis, representing ways of “going about” these HR tasks. This inquiry disentangles whether GenAI-driven productivity gains in HR tasks depend on access to the tool or on how it is enacted in day-to-day HR practice. Additionally, we uncover a sequential pattern in tool use, such that late users performed better, which we attribute to either contrast effects (Mussweiler and Strack, 2000) or deeper initial task engagement at earlier stages.

HRM is conceptualized as the design and implementation of HR policies and practices within the daily operations of organizations, with its strategic nature underscoring the alignment with business strategies and the potential to generate a sustainable competitive advantage (Collins and Clark, 2003; Lo et al., 2015; Paauwe and Boon, 2018). However, the actual day-to-day work of an HR professional is quite different. It tends to be segmented and overloaded with routine tasks, encompassing ongoing interactions and interpersonal communications with line managers, employees and job candidates (Persson and Wallo, 2024), service requests and document work (Wallo and Coetzer, 2023) and people-related decision-making (Narzary et al., 2025; Espegren and Hugosson, 2025). A large part of everyday HR work is organized around written tasks that support these activities. These HR tasks include, for example, internal messages and announcements, job advertisement texts and selection materials, adapted to local conditions and contextual expectations (Häll et al., 2023). Producing these outputs is often embedded in a reactive workday shaped by stakeholder requests and frequent meetings, leaving limited opportunities for uninterrupted work (Wallo and Coetzer, 2023). This micro-level reality contrasts with the strategic vision of HR roles and yet remains underexplored in empirical research (Narzary et al., 2025; Wallo and Coetzer, 2023).

Viewed through an HRM-as-practice lens, HR work as an organizational phenomenon does not exist until it is enacted; both individual actions and structural factors acquire meaning when they are manifested in practice (Björkman et al., 2014). Specifically, this perspective suggests that a coherent understanding of HR work involves focusing on “3Ps”, namely HR practices, HR praxis, HR practitioners and intersections among them (Espegren and Hugosson, 2025). HR practices can be conceptualized as abstract general activities, patterns and principles (e.g. high-performance work practices; Espegren and Hugosson, 2025) or norms, processes and procedures (Björkman et al., 2014). In contrast, HR praxis captures the actual work, underscoring how work is practically carried out (Espegren and Hugosson, 2025). The distinction between HR practices and HR praxis aligns with process views of HRM that differentiate between intended and implemented HRM (Nishii and Wright, 2008). It also resonates with the rhetoric and reality tensions in HRM, where strategic role narratives can diverge from how HR work is actually performed under constraints and competing demands (Legge, 2005). The third pillar of this perspective refers to HR practitioners, who are the actors directly involved in this process of enactment. They are often described as being “at the beck and call” of managers (Wright, 2008), as they must continuously shift across short cycles of requests, follow-ups and coordination with limited opportunities for uninterrupted work.

Taken together, these insights illustrate that HRM practitioners often engage in HR praxis within a fragmented, daily reality in which they are expected to be increasingly productive. Productivity captures both the effectiveness of work completion (i.e. the quality of outputs) and the efficiency of delivery (i.e. amount of output relative to time input), and HR practitioners’ responsibilities extend to both dimensions. In this study, we conceptualize productivity as comprising three measurable components: (1) effectiveness, operationalized as the quality of task outputs; (2) time, representing the speed of task completion; and (3) efficiency, computed as the ratio of quality to time. This aligns with literature that views productivity as a balance between work quality and resource utilization (e.g. Sink and Tuttle, 1989; Koss and Lewis, 1993).

The contemporary landscape featuring technological innovations, such as AI applications, creates an unprecedented opportunity to reshape the work and productivity of HR practitioners. Such technologies may signal a turning point for HR practitioners and their productivity, as HR tasks are enacted under time pressure and interruptions that constrain the checking, tailoring and consistency work performed before outputs are shared with stakeholders (Wallo and Coetzer, 2023). Conversely, the use of e-HRM technologies can be “insidious” as tools become embedded in routines over time and will have undeclared and perhaps unexpected effects on how HRM is conducted in the organization in the future (Bondarouk and Brewster, 2016). To date, however, little research has examined HR praxis, the enactment of daily tasks by HR practitioners (Ferm et al., 2024; Häll et al., 2023; Wallo and Coetzer, 2023) and its relationship with productivity in the context of GenAI.

The emergence of GenAI and its tools (e.g. ChatGPT, Gemini, Claude, DALL-E) has begun to transform HR functions, including practices and tasks (Kim et al., 2021). This integration will revolutionize HRM operations by increasing the efficacy and efficiency of HR departments (Nankervis et al., 2021), necessitating that HR practitioners take on new responsibilities and develop new skills (Cantoni and Mangia, 2019). Capable of producing human-like work outputs (Dell’ Acqua et al., 2023), GenAI accelerates the shift toward “intelligent” automation in HR tasks. In the recruitment process, for example, AI applications operate as bots and respond to candidates’ queries or provide information about the requirements of the position (Almaraghi, 2024; Kambur and Akar, 2022; Pan and Froese, 2023). AI applications help with screening candidates, with reduced human interaction and involvement (Rabenu and Baruch, 2025). AI applications in HR extend to payroll, administration and worktime management (Balasundaram and Venkatagiri, 2020; Upadhyay and Khandelwal, 2018). In the field of analytics, AI technology is considered as a strategic tool which supports data-driven decision-making (Vrontis et al., 2022), delivered with greater speed (Niehueser and Boak, 2020; Albert, 2019) and scale (Davenport and Mittal, 2022). Overall, AI and related intelligent technologies can be catalysts in increasing the productivity of HR practitioners and strengthening HRM functions and activities (Budhwar et al., 2022).

Recent empirical studies across various professional domains highlight the transformative impact of GenAI, particularly advanced chatbots like ChatGPT, on employee productivity. Noy and Zhang (2023) demonstrated substantial productivity gains in medium-level professional writing tasks (among consultants, data analysts, marketers, HR staff focused only on communication tasks). Similarly, Dell’ Acqua et al. (2023) found that ChatGPT use significantly improved both speed and quality in complex consulting work, yet revealed limits in tasks beyond AI’s frontier. In creative industries, Zhou and Lee (2024) demonstrated a 25% rise in artistic output with text to image AI. Brynjolfsson et al. (2023) reported a 14% productivity boost among customer support agents using conversational AI, in less experienced staff. Karakas et al. (2023) discussed the GenAI’s role in improving administrative efficiency in pediatric neurology, while Choi et al. (2023) found reduced time and greater satisfaction in legal analysis tasks. At firm level, Kim et al. (2022) linked AI adoption to improved cost structures in diverse sectors, and Dell’ Acqua (2022) highlighted how recruiter engagement shifted under AI guidance.

Despite these advances, the impact of GenAI on productivity within the HR profession itself remains underexplored, particularly in routine, written outputs that support everyday HR interactions and micro-decisions (e.g. drafting and tailoring texts that coordinate HR service delivery). GenAI is treated here as augmentation for routine written work. The main goal is to automate or at least reduce the time required for repetitive, rule-based tasks, such as drafting and structuring texts, reducing costs while improving quality (Wirtz et al., 2019). From a practice-oriented perspective, the value of this technology is manifested only when it becomes part of HR praxis, enacted by practitioners and shaped by contextual factors and by alignment to users’ needs (Espegren and Hugosson, 2025). Accordingly, this research aims to answer the following research question:

RQ1.

How does the use of GenAI affect productivity in HR tasks?

The advent of GenAI has not only transformed the content and the nature of HR work (Malik et al., 2023) but also introduced diverse ways that practitioners interact with technology in routine HR work. From an HRM-as-practice perspective, a digital tool matters through how it is used in routine work, within ongoing demands, roles and local adjustments (Espegren and Hugosson, 2025). In this sense, “use” is not a single behavior as practitioners may submit outputs with no edits or use the tool iteratively while retaining responsibility for tailoring and final checks. Extant literature has focused on the influence of AI on productivity in various professions, but very few studies have directly examined whether productivity depends on how AI tools are used in practice, that is how their use is enacted within routine workflows.

A key distinction exists between directly adopting AI-generated outputs and refining them through iterative editing. Noy and Zhang (2023) documented usage patterns on ChatGPT (e.g. writing drafts, summarizing texts, brainstorming) and observed that the majority of study participants submitted ChatGPT’s outputs without further editing. Interestingly, they found no evidence that users who edited or redefined the AI-generated outputs achieved higher quality than those who submitted raw outputs. This suggests that more intensive engagement with the tool does not automatically produce better outcomes, making it important to test usage patterns directly rather than assume their value. This insight also leaves open critical questions about the actual value of human intervention in AI-generated outputs raising concerns about potential job replacement and substitution of human labor (Kambur and Akar, 2022).

In the recruitment context, a field experiment further underscores that the way AI tools are used can critically impact outcomes (Dell’ Acqua, 2022). Recruiters who passively accepted highly accurate AI recommendations performed worse than those who confronted AI outputs with less reliability. This contrasts with the findings of Noy and Zhang (2023), suggesting that much more research is needed to elucidate the productivity implications of each approach. Dell’ Acqua et al. (2023) drew a further conceptual distinction between “centaur” practices, where humans and AI handle separate subtasks, and “cyborg” practices, involving deeply intertwined, iterative collaboration at the subtask level. However, this study does not provide evidence of which practice leads to greater productivity.

Overall, this body of evidence suggests that productivity gains from GenAI may depend less on whether a technology exists and more on how these tools are actually used in routine HR tasks, hence the relevance of examining usage patterns rather than treating “use” as a single, uniform behavior. Consistent with HR praxis, outputs may require contextual tailoring and professional judgement to fit local expectations and stakeholders, although a productivity advantage in this type of usage should not be taken for granted. To our knowledge, no empirical study has explicitly examined whether different usage patterns affect productivity in the context of routine HR tasks. Accordingly, this study addresses the following research question:

RQ2.

Do different usage patterns of GenAI (e.g. raw adoption vs. iterative editing) differentially influence productivity outcomes?

This study builds on recent frameworks emphasizing the operational activities of HR work (Wallo and Coetzer, 2023) and the emerging influence of GenAI on productivity (Noy and Zhang, 2023; Budhwar et al., 2022). Given the aim of estimating the effect of GenAI use (ChatGPT) on productivity in a simulated HR work setting, participants were randomly assigned to a treatment or control condition, and an experimental design was employed (Gray, 2021). Here, “use” refers to assignment to a condition in which ChatGPT is available for use during task completion. This design permits a causal interpretation of the effect of ChatGPT use within the experimental setting, while implications beyond junior, routine HR tasks are treated as a boundary condition. The primary methodological approach was an experimental vignette design involving concise, systematically constructed descriptions of scenarios that represent a deliberate combination of specific characteristics (Brewer, 2000).

In this experiment, the vignettes are presented as realistic routine, operational HR writing tasks, each designed to be completed within a maximum of 30 min. Experimental design is particularly well-suited to achieving the objectives of this study and has been widely used in related fields (Aguinis and Bradley, 2014). To further strengthen internal validity (Gray, 2021), participants were assigned to treatment and control groups (Hedrick et al., 1993). All tasks were completed individually in supervised computer-lab sessions under identical technical conditions. Participants were instructed to rely on standard web resources and, where applicable, ChatGPT, and to avoid collaboration or the use of other generative tools during task completion.

The three vignette-based tasks assigned to participants were designed to simulate distinct functional HRM areas. These are routine HR tasks that junior HR staff commonly perform across internal communication/employer branding, recruitment and selection (see supplementary material for a sample). The first task simulated an internal HR communication addressing staff and clients (design of Christmas cards), reflecting key activities in internal communication, organizational culture and employer branding. The task aimed to reinforce core company values and test participants’ ability to produce creative, values-driven communication. The second task required the development of a recruitment advertisement, based on a full job analysis and structured around the AIDA (attention, interest, desire, action) model. This task targeted core HRM activities within resourcing and assessed the ability to translate job requirements into persuasive communication aligned with the organization’s profile.

Finally, the third task required participants to propose a set of behavioral and/or situational interview questions based on a full competency-based job analysis. For each question, participants were asked to specify the targeted competency and explain the rationale behind its design. This task is related to the selection function in HRM and was intended to assess participants’ ability to apply job analysis results and align assessment methods with role-specific requirements.

Initially, a preliminary exploratory study was conducted with HR practitioners, aimed at developing three different vignettes representing realistic HR tasks along with appropriate completion times for an HR administration internship position. The vignettes derived from the preliminary research were subsequently used in the main experimental study.

In the first phase, three HR managers (currently employed in HR roles that include supervising interns/junior HR employees) provided examples of real tasks previously assigned to business school interns, focusing on routine written deliverables that require no specific training and are relatively simple enough to be completed within approximately 30 min. The tasks were intentionally designed to be completed in a tight schedule, reflecting the time pressures typical of a real business environment.

Next, ten HR practitioners were recruited through professional networks. Eligibility required a current HR role and at least three years of HR experience involving routine written deliverables (e.g. internal communication, recruitment texts, selection documentation). In the resulting panel, experience ranged from 3 to 15 years (M = 7.3, SD = 3.4). These draft tasks were then evaluated through a brief questionnaire. Participants rated realism and appropriateness on a 5-point Likert scale and provided qualitative feedback. Based on these evaluations, minor adjustments were made to enhance clarity and precision before inclusion in the main experiment.

Convenience sampling was used to select the participants for the experiment, primarily due to accessibility (Rahi, 2017) and relevance to the research context. Such sampling allowed the recruitment of final-year management students poised to enter internships or entry-level HR positions, ensuring alignment with study’s objectives (Saunders et al., 2012).

The aim was to measure the effect of AI on productivity with three productivity measures: (1) effectiveness (quality of outputs), (2) completion time and (3) efficiency (quality per unit of time). This study was reviewed and approved by the Ethics Review Board of a public university in Greece. The ethics clearance number is available upon request. All participants provided informed consent prior to their involvement in the study.

Participants were selected using convenience sampling (Rahi, 2017). The participants were 91 final-year students eligible to undertake their internship from the Department of Management Science and Technology at a public university in Northern Greece. This sample was chosen because they represent individuals who are at the threshold of entering the professional workforce and are likely to engage in routine, junior-level operational HR writing tasks. Therefore, this group provides an appropriate context for testing how GenAI tools impact productivity in routine HR writing tasks relevant to entry-level HR roles.

Participants were randomly assigned to either the control or the treatment group using a lottery method, with participants labeled as either number 1 for group A (control group) or number 2 for group B (treatment group). Each participant was assigned a unique identifier to maintain anonymity throughout the process and ensure confidentiality. Prior to the start of the experimental procedure, participants received standardized instructions regarding the structure of the process, along with guidance on what actions they were expected to complete.

Students who agreed to participate were asked to complete three realistic HR tasks, similar to those they would perform in an HR administration internship role. The experiment took place in parallel across three computer labs with equivalent technology and internet access, and the experimental material was available online through Google Forms. A researcher supervised all sessions to ensure procedural consistency and resolve any technical issues. The completion time for each task was recorded through Google Forms, with each submission carrying a timestamp.

Initially, participants were provided with instructions for the experiment procedure and afterward were asked to give their consent. They then answered a brief questionnaire capturing their demographic characteristics. Right after, participants completed the first assigned task within a maximum time frame of 30 min, followed by a 10-min break.

After the break, the students in group B (treatment) registered for ChatGPT accounts as needed. Prior to the second task, the treatment group underwent a brief training session on the basic functions and capabilities of ChatGPT 3.5. Then, they proceeded with the second task, which also had a 30-min time limit.

After another 10-min break, participants continued with the third task. Participants in group A (control) registered in their ChatGPT accounts at this stage, as both groups used the tool during the third task. This allowed for within-group comparisons of performance between the first two non-AI-supported tasks and the final AI-supported task. They received the same training that the treatment group had undertaken in the previous step of the experimental procedure, becoming a delayed treatment group (Elliott and Brown, 2002). The third task was likewise limited to 30 min, and after completing it, both groups completed a questionnaire regarding their experience with ChatGPT. The questionnaire captured patterns of use, including whether participants relied on ChatGPT for raw output or engaged in more iterative editing processes. After each phase, the experiment concluded. The structure and timeline of the experimental procedure are summarized in Figure 1 (experimental design).

Figure 1
A flowchart illustrating the experimental design process involving three tasks, breaks, and group assignments.The flowchart illustrates the experimental design process. The process begins with an introduction phase that includes instructions, consent, and demographic data collection. This is followed by Task 1, which lasts for 30 minutes. After Task 1, there is a 10-minute break. Participants are then divided into two groups: Group A, the control group, and Group B, the treatment group, which involves registering or signing up on a ChatGPT account. Both groups proceed to Task 2, which also lasts for 30 minutes. After Task 2, there is another 10-minute break. Group A becomes the delayed treatment group and also registers or signs up on a ChatGPT account. Both groups then proceed to Task 3, which lasts for 30 minutes. Following Task 3, participants answer questions about ChatGPT usage. The survey ends, and the results are evaluated by experts.

Experimental design

Figure 1
A flowchart illustrating the experimental design process involving three tasks, breaks, and group assignments.The flowchart illustrates the experimental design process. The process begins with an introduction phase that includes instructions, consent, and demographic data collection. This is followed by Task 1, which lasts for 30 minutes. After Task 1, there is a 10-minute break. Participants are then divided into two groups: Group A, the control group, and Group B, the treatment group, which involves registering or signing up on a ChatGPT account. Both groups proceed to Task 2, which also lasts for 30 minutes. After Task 2, there is another 10-minute break. Group A becomes the delayed treatment group and also registers or signs up on a ChatGPT account. Both groups then proceed to Task 3, which lasts for 30 minutes. Following Task 3, participants answer questions about ChatGPT usage. The survey ends, and the results are evaluated by experts.

Experimental design

Close modal

The quality of the work was assessed through blind review by experienced HR practitioners who evaluate similar written deliverables in their daily work. The panel comprised ten HR practitioners (who also participated in the preliminary study) and three academics with professional HR experience. Each deliverable was evaluated by three different raters. All submissions were anonymized and presented in random order. Raters were blind to experimental condition and participant identity. Additionally, each participant was evaluated by at least five different raters in total across all three tasks. The assessment was based on three criteria: quality of written expression, quality of content and originality, each rated on a 7-point Likert scale in line with Noy and Zhang’s (2023) study. To assess inter-rater reliability, Fleiss’ Kappa was calculated separately for the three evaluation criteria across the three raters assigned to each deliverable. The obtained kappa coefficients ranged from 0.75 to 0.85 and were statistically significant in all cases (p < 0.001), indicating substantial agreement among raters and supporting the consistent application of the evaluation criteria. As an additional robustness check, Krippendorff’s alpha for ordinal ratings was also estimated and produced satisfactory results across the three evaluation criteria, further confirming the consistency of the raters’ evaluations.

The initial sample consisted of 91 undergraduate students, who were randomly assigned to either control group (group A, n = 46) or treatment group (group B, n = 45). Participants were similar in terms of gender distribution (group A: 37% male; group B: 44% male), with no significant differences between groups on demographic characteristics (p > 0.05). This suggests that the randomization procedure was successful and that the observed group differences in task performance are unlikely to be due to demographic differences.

To evaluate the impact of ChatGPT usage on participants’ productivity, three primary measures aligned with a multidimensional productivity framework were analyzed:

(1) Effectiveness, assessed via evaluators’ expert ratings of deliverable quality (i.e. the quality of work), (2) completion time, representing the duration taken to complete each task (in minutes), and (3) efficiency, calculated as the quality score divided by completion time. Higher efficiency values indicate greater productivity per unit of time. This operationalization allows for a more nuanced assessment of productivity, distinguishing between doing work well (effectiveness), doing it fast (time) and doing it well and fast (efficiency). Reported percentage changes refer to improvement from non-AI-supported to ChatGPT-supported task completion across the delayed-treatment design, rather than to a pure Task 3 treatment-versus-control contrast.

For the statistical analysis, IBM SPSS Statistics version 30.0 was used. Statistical analyses were conducted to assess both between-group effects (Control vs. Treatment group) and within-group effects (i.e. changes across tasks in the same group).

Table 1 provides a summary of descriptive statistics across all measures of productivity and tasks. As shown, the two groups exhibited differences across productivity indicators and tasks.

Table 1

Mean scores and standard deviations across all measures of productivity and tasks

MeasureTaskGroup A (mean ± SD)Group B (mean ± SD)Total (mean ± SD)
EffectivenessTask 13.41 ± 0.693.09 ± 1.053.24 ± 0.90
EffectivenessTask 23.29 ± 0.693.82 ± 1.043.57 ± 0.92
EffectivenessTask 34.95 ± 1.353.90 ± 1.284.40 ± 1.41
TimeTask 128.65 ± 1.8424.65 ± 4.1326.55 ± 3.80
TimeTask 226.68 ± 3.2117.06 ± 8.9221.65 ± 8.33
TimeTask 38.74 ± 4.7815.41 ± 8.7412.23 ± 7.83
EfficiencyTask 10.12 ± 0.020.13 ± 0.040.12 ± 0.03
EfficiencyTask 20.13 ± 0.030.36 ± 0.320.25 ± 0.26
EfficiencyTask 30.81 ± 0.610.39 ± 0.380.59 ± 0.54

Note(s): Values in Table 1 represent raw descriptive means and standard deviations. Means reported in the inferential analyses are estimated marginal means from the mixed-design ANOVA and may therefore differ slightly from the descriptive means reported in the table

To examine effectiveness across all three tasks and evaluate how participants’ scores evolved within and between experimental groups, a 2 (Group: A vs. B) × 3 (Task: 1, 2, 3) mixed-design ANOVA was conducted. The within-subjects factor was Task, comprising three levels and the between-subjects factor was Group.

The multivariate analysis revealed a significant main effect of Task, Wilks’ Λ = 0.573, F(2, 62) = 23.07, p < 0.001, η2p = 0.427, indicating that effectiveness varied significantly across tasks. The means reported in the following pairwise comparisons are estimated marginal means from the mixed model and therefore differ slightly from the raw descriptive means reported in Table 1, which presents unadjusted means and standard deviations. Follow-up pairwise comparisons of estimated marginal means showed that overall effectiveness increased from Task 1 (M = 3.25, SE = 0.11) to Task 2 (M = 3.56, SE = 0.11), p = 0.075 (marginal) and significantly from Task 2 to Task 3 (M = 4.43, SE = 0.16), p < 0.001. The Task 2–Task 3 mean difference was 0.87 points (95% CI [0.47, 1.27]). A significant Task × Group interaction emerged, Wilks’ Λ = 0.704, F(2, 62) = 13.02, p < 0.001, η2p = 0.296, indicating that the effectiveness trajectory differed between groups.

As shown in Figure 2, participants in Group B, who used ChatGPT from Task 2, demonstrated a marked improvement between Task 1 and Task 2, as indicated by paired-samples t-tests, t(43) = −3.92, p < 0.001. However, this was followed by a plateau in Task 3, t(33) = −0.33, p = 0.740. In contrast, Group A, who used ChatGPT only in Task 3, showed no improvement in Task 2, t(39) = 0.41, p = 0.684. However, Group A exhibited a substantial gain in Task 3, t(30) = −7.41, p < 0.001.

Figure 2
A line graph showing estimated marginal means of effectiveness across three tasks by experimental group.A line graph showing estimated marginal means of effectiveness across three tasks by experimental group. The x axis represents two groups, Group A and Group B. The y axis represents estimated marginal means ranging from 3.000 to 5.500. There are three data lines representing three tasks, labeled as 1, 2, and 3. Task 1 is represented by a purple line, Task 2 by an orange line, and Task 3 by a blue line. For Group A, Task 1 has a mean of 3.290, Task 2 has a mean of 3.409, and Task 3 has a mean of 4.953. For Group B, Task 1 has a mean of 3.088, Task 2 has a mean of 3.824, and Task 3 has a mean of 3.902. Error bars representing 95% confidence intervals are included for each data point.

Estimated marginal means of effectiveness across three tasks by experimental group (with 95% confidence intervals)

Figure 2
A line graph showing estimated marginal means of effectiveness across three tasks by experimental group.A line graph showing estimated marginal means of effectiveness across three tasks by experimental group. The x axis represents two groups, Group A and Group B. The y axis represents estimated marginal means ranging from 3.000 to 5.500. There are three data lines representing three tasks, labeled as 1, 2, and 3. Task 1 is represented by a purple line, Task 2 by an orange line, and Task 3 by a blue line. For Group A, Task 1 has a mean of 3.290, Task 2 has a mean of 3.409, and Task 3 has a mean of 4.953. For Group B, Task 1 has a mean of 3.088, Task 2 has a mean of 3.824, and Task 3 has a mean of 3.902. Error bars representing 95% confidence intervals are included for each data point.

Estimated marginal means of effectiveness across three tasks by experimental group (with 95% confidence intervals)

Close modal

To investigate how tasks’ completion time evolved across all three tasks within and between experimental groups, a 2 (Group: A vs. B) × 3 (Task: 1, 2, 3) mixed-design ANOVA was conducted. The within-subjects factor was Task, comprising three levels, and the between-subjects factor was Group.

The multivariate analysis revealed a significant main effect of Task, Wilks’ Λ = 0.175, F(2, 62) = 146.19, p < 0.001, η2p = 0.825, indicating that completion time significantly decreased over time across tasks. As noted above, the pairwise comparisons are based on estimated marginal means from the mixed model, which may differ slightly from the raw descriptive means reported in Table 1. Pairwise comparisons of estimated marginal means showed a drop from Task 1 (M = 26.65, SE = 0.40) to Task 2 (M = 21.87, SE = 0.85), p < 0.001, and further to Task 3 (M = 12.08, SE = 0.89), p < 0.001. The decrease between Task 2 and Task 3 was statistically significant, p < 0.001. All differences between tasks were statistically significant, with the largest decrease observed between Task 2 and Task 3 (Mdiff = 9.79 min, 95% CI [7.05, 12.54]). A significant Task × Group interaction was also observed, Wilks’ Λ = 0.512, F(2, 62) = 29.52, p < 0.001, η2p = 0.488, indicating that the reduction in time varied by group.

In Group A, completion time decreased modestly between Task 1 and Task 2 (M1 = 28.65, SD = 1.84; M2 = 26.68, SD = 3.21), t(39) = 3.03, p = 0.004, but decreased substantially from Task 2 to Task 3 (M3 = 8.74, SD = 4.78), t(30) = 18.46, p < 0.001. The total decrease from Task 1 to Task 3 was also significant, t(34) = 21.23, p < 0.001. In contrast, Group B, which had access to ChatGPT from Task 2, first exhibited a decrease in completion time from Task 1 to Task 2 (M1 = 24.65, SD = 4.13; M2 = 17.06, SD = 8.92), t(43) = 5.98, p < 0.001, followed by a non-significant change from Task 2 to Task 3 (M3 = 15.41, SD = 8.74), t(33) = 0.85, p = 0.402. For interpretive clarity, the group-specific M and SD values reported in this paragraph represent descriptive statistics, whereas the t-values are based on within-group paired-samples tests.

These findings suggest that the introduction of ChatGPT access is linked to substantial reductions in completion time in routine tasks, with the largest reduction at first exposure. As shown in Figure 3, Group A benefited most when ChatGPT was first introduced in Task 3, whereas Group B showed little change from Task 2 to Task 3 after earlier time gains.

Figure 3
A line graph showing estimated marginal means of completion time across three tasks by experimental group with 95% confidence intervals.A line graph showing estimated marginal means of completion time across three tasks by experimental group with 95% confidence intervals. The horizontal axis represents the groups, labeled as Group A and Group B. The vertical axis represents the estimated marginal means, ranging from 5 to 30. There are three data lines, each representing a different task: Task 1 in purple, Task 2 in orange, and Task 3 in blue. Each line shows the trend of completion times for the respective task across the two groups. Task 1 starts at approximately 29 for Group A and decreases to approximately 25 for Group B. Task 2 starts at approximately 27 for Group A and decreases to approximately 17 for Group B. Task 3 starts at approximately 9 for Group A and increases to approximately 15 for Group B. Error bars representing 95% confidence intervals are included for each data point.

Estimated marginal means of completion time across three tasks by experimental group (with 95% confidence intervals)

Figure 3
A line graph showing estimated marginal means of completion time across three tasks by experimental group with 95% confidence intervals.A line graph showing estimated marginal means of completion time across three tasks by experimental group with 95% confidence intervals. The horizontal axis represents the groups, labeled as Group A and Group B. The vertical axis represents the estimated marginal means, ranging from 5 to 30. There are three data lines, each representing a different task: Task 1 in purple, Task 2 in orange, and Task 3 in blue. Each line shows the trend of completion times for the respective task across the two groups. Task 1 starts at approximately 29 for Group A and decreases to approximately 25 for Group B. Task 2 starts at approximately 27 for Group A and decreases to approximately 17 for Group B. Task 3 starts at approximately 9 for Group A and increases to approximately 15 for Group B. Error bars representing 95% confidence intervals are included for each data point.

Estimated marginal means of completion time across three tasks by experimental group (with 95% confidence intervals)

Close modal

To study the sequential impact of ChatGPT usage on task efficiency (defined as effectiveness score ÷ completion time), a 2 (Group: A vs. B, between-subjects) × 3 (Task: 1, 2, 3, within-subjects) mixed-design ANOVA was conducted. Task served as the within-subjects factor, comprising three levels, and Group as the between-subjects factor.

Multivariate analysis revealed a significant main effect of Task, Wilks’ Λ = 0.468, F(2, 62) = 35.21, p < 0.001, η2p = 0.532, indicating that task efficiency increased significantly across tasks. As noted above, the pairwise comparisons are based on estimated marginal means from the mixed model, which may differ slightly from the raw descriptive means reported in Table 1. Pairwise comparisons of estimated marginal means showed a consistent improvement from Task 1 (M = 0.123, SE = 0.004) to Task 2 (M = 0.243, SE = 0.029), p < 0.001, and further to Task 3 (M = 0.598, SE = 0.062), p < 0.001. Both the linear trend, F(1, 63) = 57.81, p < 0.001, η2p = 0.479, and the quadratic trend, F(1, 63) = 7.92, p = 0.007, η2p = 0.112, were statistically significant, indicating a nonlinear acceleration of efficiency gains over time. A significant Task × Group interaction also emerged, Wilks’ Λ = 0.695, F(2, 62) = 13.63, p < 0.001, η2p = 0.305, suggesting that the trajectory of efficiency differed between groups. Contrast tests further confirmed that both the linear, F(1, 63) = 11.75, p = 0.001, η2p = 0.157, and quadratic components, F(1, 63) = 27.50, p < 0.001, η2p = 0.304, significantly interacted with group.

Descriptive statistics indicated that participants in Group A showed minimal change in efficiency from Task 1 to Task 2 (M1 = 0.119, SD = 0.02; M2 = 0.125, SD = 0.03), followed by a substantial increase in Task 3 (M = 0.809, SD = 0.61). In contrast, participants in Group B demonstrated a sharp improvement between Task 1 and Task 2 (M1 = 0.127, SD = 0.04; M2 = 0.360, SD = 0.32), but only a modest gain in Task 3 (M = 0.388, SD = 0.38), indicating a plateau following early exposure to ChatGPT.

Although the overall between-subjects effect of Group on mean efficiency was not statistically significant, F(1, 63) = 1.65, p = 0.203, η2p = 0.026, the significant interaction effects highlight distinct developmental trajectories in efficiency across conditions.

Paired-samples t-tests reinforced these findings. Across the total sample, efficiency improved significantly between Task 1 and Task 2 and also between Task 2 and Task 3. This pattern suggests that while ChatGPT support led to immediate efficiency gains for Group B, longer-term improvements were more pronounced in Group A, who encountered ChatGPT for the first time in Task 3. The significant quadratic trend underscores the nonlinear nature of improvement, with efficiency accelerating rapidly after ChatGPT was introduced. This pattern is illustrated in Figure 4.

Figure 4
A line graph showing estimated marginal means of efficiency across three tasks by experimental group.A line graph showing estimated marginal means of efficiency across three tasks by experimental group. The x axis represents the groups, labeled as Group A and Group B. The y axis represents the estimated marginal means, ranging from 0.0000 to 1.0000. There are three data lines, each representing a different task. Task 1 is represented by a purple line, Task 2 by an orange line, and Task 3 by a blue line. Each data point is marked with its value, and error bars indicate 95% confidence intervals. Group A shows values of 0.1194 for Task 1, 0.1254 for Task 2, and 0.8088 for Task 3. Group B shows values of 0.1270 for Task 1, 0.3596 for Task 2, and 0.3879 for Task 3.

Estimated marginal means of efficiency across three tasks by experimental group (with 95% confidence intervals)

Figure 4
A line graph showing estimated marginal means of efficiency across three tasks by experimental group.A line graph showing estimated marginal means of efficiency across three tasks by experimental group. The x axis represents the groups, labeled as Group A and Group B. The y axis represents the estimated marginal means, ranging from 0.0000 to 1.0000. There are three data lines, each representing a different task. Task 1 is represented by a purple line, Task 2 by an orange line, and Task 3 by a blue line. Each data point is marked with its value, and error bars indicate 95% confidence intervals. Group A shows values of 0.1194 for Task 1, 0.1254 for Task 2, and 0.8088 for Task 3. Group B shows values of 0.1270 for Task 1, 0.3596 for Task 2, and 0.3879 for Task 3.

Estimated marginal means of efficiency across three tasks by experimental group (with 95% confidence intervals)

Close modal

To investigate whether the usage pattern of ChatGPT, namely raw adoption of the generated output versus iterative editing, influenced productivity outcomes, participants were classified based on their responses to the post-task questionnaire. Specifically, after completing the AI-supported tasks, they indicated whether they had used ChatGPT’s output directly without further editing or whether they had edited it before submission. A series of analyses were then conducted comparing effectiveness, completion time and efficiency scores between these two groups. The analyses revealed no statistically significant differences across the outcome measures, suggesting that the method of ChatGPT use did not significantly influence participants’ performance in this setting. Specifically, non-significant differences were observed for Task 2 (p = 0.711) and Task 3 (p = 0.972), while efficiency differences were also non-significant across the corresponding analyses (p-values ranging from 0.214 to 0.658).

The purpose of this study was to investigate the impact of GenAI, and especially the ChatGPT tool, on productivity outcomes in routine, junior-level HR writing tasks, in the context of a work-simulated process through an experimental methodology. We examine ChatGPT as a tool enacted in completing these deliverables and assess not only the effect of tool usage but also whether usage patterns, raw adoption versus iterative editing, differentiate productivity outcomes. In doing so, the discussion links the findings back to the two research questions and to the practice-based view of HR work.

In this experimental setting, access to ChatGPT increased efficiency (+391.7%), alongside shorter completion time (−53.9%) and higher effectiveness (quality of deliverables, +35.8%). These results extend previous research that demonstrated positive effects of GenAI tools across various professional domains, including consulting (Dell’ Acqua et al., 2023), pediatric neurology (Karakas et al., 2023), creative industries (Zhou and Lee, 2024), customer support (Brynjolfsson et al., 2023), legal analysis (Choi et al., 2023), software development (Peng et al., 2023) and general professional writing tasks (Noy and Zhang, 2023). Further, these findings highlight the value of treating productivity as a multidimensional construct, operationalized through three complementary measures: effectiveness (output quality), completion time and efficiency (output per unit of time). This framing enables a more nuanced understanding of how GenAI tools influence productivity across distinct dimensions. Regarding RQ1, the findings indicate that tool usage in HR tasks primarily improves efficiency while also enhancing output quality, rather than accelerating work at the expense of quality. This is important because these tasks are produced within a fragmented and pressured daily routine; hence, gains in speed and quality can drastically improve HR service delivery in everyday HR praxis.

No evidence supported differences between usage patterns (raw adoption and iterative editing). Beyond aggregate effects, this finding aligns with Noy and Zhang (2023), but should be interpreted cautiously because usage patterns were classified post hoc among participants who used ChatGPT, rather than experimentally assigned. In this exploratory analysis, basic use of ChatGPT was associated with early productivity gains in routine, junior-level tasks, although more complex HR work may require more advanced prompting and professional review. In terms of RQ2, findings suggest that, for junior-level HR tasks completed under time constraints, the main distinction is whether participants use GenAI or not, whereas differences in how they work with the tool are not reflected in measurable outcome differences in this setting. From a practice-oriented view, the key issue in the context of these HR tasks is whether the tool is embedded and utilized in the workflow, rather than how it is used.

A more detailed observation of the results reveals an unexpected pattern. Participants in Group A, who first used ChatGPT in Task 3, improved more than twice as much as Group B who first used it in Task 2. This first exposure boost suggests that, in this setting, productivity, especially efficiency, peaks at initial adoption, while early adopters showed little further improvement; once the tool’s basic affordances are learned, further gains diminish. As presented in Figure 1 (experimental design), Group A, initially the control group, later became a delayed treatment group that also received the intervention. This pattern is consistent with delayed-treatment designs (Elliott and Brown, 2002), where delaying access does not blunt the effect and performance typically catches up. In this case, Group A’s performance before accessing ChatGPT improves (although not statistically significant) and, in the third task, not only converges with but exceeds Group B’s performance, suggesting that first exposure effects may be shaped by contrast-based evaluations and by prior task engagement without tool support. Here, a contrast effect may be relevant, whereby evaluations of a tool or environment are influenced by previous experience (Mussweiler and Strack, 2000). Group A’s exposure to Tasks 1 and 2 without ChatGPT appears to have enhanced their perception and use of the tool when it was finally introduced, producing contrast-based amplification that magnified ChatGPT’s benefits both subjectively and objectively.

Another possible explanation is deeper and more active engagement with the task content by Group A during the first two tasks. The assignment outlines included the principles and values of the fictitious company. Therefore, since they were asked to complete the tasks without the use of the GenAI tool, they may have developed a clearer grasp of the contextual cues embedded in the brief (e.g. values, tone and intended audience). This deeper engagement with the context may have helped participants translate these cues into more targeted written outputs once ChatGPT was introduced, incorporating the organizational values more consistently in the third task. Consistent with the HRM as practice perspective overcoming agency and structure dualism, findings support that the effects are not only an outcome of tool usage but also depend on how users (practitioners) understood everyday demands, role expectations and local conditions (Espegren and Hugosson, 2025). This interpretation aligns with experimental evidence showing that deep semantic processing enhances immediate task performance, not only delayed retention (Chang, 2017).

Finally, the findings can be situated within HR transformation debates about how new tools and arrangements redistribute work and intensify tensions between standardization and local adaptation (Häll et al., 2023). In routine HR tasks, GenAI can reduce drafting time, potentially creating capacity for other work within everyday HR service delivery or opportunities for strategic involvement. However, consistent with the HRM-as-practice view, the value of HR outputs also depends on contextualization and fit with local conditions and stakeholder expectations. This implies that as the use of GenAI becomes commonplace, productivity gains may reconfigure HR work rather than reduce it, shifting effort toward supervision and quality assurance of machine assisted outputs. Further, when first drafts can be produced quickly, work can be redistributed in different ways. For example, HR may produce more standardized drafts in less time, or drafting may shift to line managers or other employees, with HR providing review. In either case, once ChatGPT is embedded in organizational routines, coordination is needed to ensure consistent and compliant outputs. However, boosting productivity in junior-level HR tasks through GenAI may carry risks for entry-level employment. Emerging evidence indicates that early career workers in AI-exposed roles have experienced employment declines, particularly in positions dominated by routine cognitive work (Brynjolfsson et al., 2025). Our findings lend empirical support to this broader pattern, by documenting sizeable productivity gains. As GenAI technologies become embedded in organizational routines, they may be difficult to reverse and could reshape how HRM is enacted over time (Bondarouk and Brewster, 2016), potentially reducing traditional entry points to organizations.

The study advances HRM-as-practice theorizing (Björkman et al., 2014) by demonstrating that productivity gains from GenAI are best understood at the level of HR praxis. While improvements from GenAI usage alone cannot be understated, from a theoretical perspective, the study shifts focus from abstract technological affordances per se toward micro-level enactments of HR work. It reinforces the view that GenAI’s actual value materializes through its integration in everyday praxis, where the intended design of HR policies meets the practical constraints, interruptions and stakeholder demands of everyday service delivery, underscoring the implementation of practices in everyday work (Nishii and Wright, 2008). In other words, effects emerge at the level of task enactment, produced in situated activity, which is often carried out in a fragmented fashion amid frequent interruptions, under time pressure and competing priorities (Wallo and Coetzer, 2023).

However, the study also suggests that theorizing about HR praxis is inseparable from HR practitioners, who are the carriers of practice in a certain context (Espegren and Hugosson, 2025). Our findings are consistent with the possibility that practitioners’ situational characteristics, such as their experience with GenAI, prior task engagement and contextual understanding of organizational features, shape how GenAI is enacted. Put differently, practitioners remain central even in GenAI-driven tasks, interpreting and contextualizing machine-generated material. This aligns with the HRM-as-practice dual emphasis on agency and structure but adds nuance, suggesting that task enactment by practitioners may represent a plausible pathway linking technological affordances and productivity.

Relatedly, these findings suggest that GenAI both enables and constraints HR praxis simultaneously. It standardizes some aspects of output production, while amplifying the importance of contextual conditions. This duality supports the view of HR praxis as emerging from ongoing negotiation between technological structures and practitioner agency (Björkman et al., 2014). In this sense, GenAI acts as a tool or structuring condition that reshapes but does not determine HR enactment. We further suggest that contemporary HR praxis requires scholars to reconceptualize it as a form of hybrid enactment, distributed across humans and intelligent tools.

Finally, the lack of performance differences between raw adoption and iterative editing indicates that, for routine HR written deliverables, access to the tool itself matters more than how participants work with it. This indicates that, in structured tasks, increased user involvement does not necessarily lead to higher-quality outputs (Noy and Zhang, 2023).

Evidently, the large productivity improvements derived from ChatGPT usage in routine, junior-level written deliverables suggest that GenAI can support organizations and practitioners in everyday HR tasks, particularly in communication, recruitment and selection activities. Large productivity gains raise an organizational design question about whether reduced time on routine HR tasks is used to reduce headcount or reallocated toward higher value-added HR work. The absence of performance differences between raw and edited AI outputs suggests that organizations may obtain gains even with basic onboarding and straightforward use of GenAI in the production of junior-level HR tasks.

However, routine HR tasks still require final checks for contextual fit and consistency before deliverables are released. If drafting becomes faster and better, organizations should decide whether some HR tasks remain primarily an HR responsibility or shift to line managers or non-specialists with HR review and what governance is needed to maintain consistency and fit across departments or business units. In day-to-day terms, governance means defining who develops a draft, who validates context and compliance and who is accountable for the final message released to stakeholders. This is also what Shotter and Tsoukas (2014) describe as phronesis in practice, in this case engaged judgment about what to accept, how to tailor outputs to the specific context and how to protect fairness and trust in everyday interactions.

Finally, HR practitioners should identify the HR tasks best suited for GenAI support, offer necessary onboarding and establish clear guidelines for professional review (e.g. mandatory checks for contextual fit, consistency and stakeholder appropriateness before outputs are released). Managers may also consider structuring workflows so that employees first engage with core task content manually, at least initially, before introducing GenAI support, combining deeper understanding of the task with the efficiency and effectiveness gains GenAI provides.

First, the use of a student sample may limit the external validity of the findings. Replications involving actual HR practitioners could provide clearer insights. Second, the tasks were limited to 30-min writing exercises, which may not fully reflect the complexities of longer, multi-hour HR workflows where GenAI tools might perform differently. Finally, the evaluation relied on expert scoring. Future research could enhance validity by incorporating objective business performance metrics (e.g. KPIs). In addition, future research could capture tool use in more detail (e.g. number of iterations, types of edits and verification checks) and combine self-reports with digital trace data (e.g. prompt logs, revision history) to better observe how work unfolds during task completion.

This study shows that ChatGPT can substantially improve productivity in routine, junior-level HR tasks by increasing output quality while reducing completion time. The findings suggest that, although much of the debate on these technologies has focused on strategic HRM, GenAI is also consequential for HRM as a task-support tool embedded in everyday HR praxis. For these routine, junior-level tasks, the absence of clear differences between raw adoption and iterative editing suggests that initial productivity gains may be achievable with basic onboarding and straightforward use, provided that outputs remain subject to professional review.

During the preparation of this work, the authors used ChatGPT (OpenAI) and DeepL as language editing tools to improve the language and readability. All edits were applied to existing, author-created material. After using these tools, the authors reviewed and edited the content as needed and take full responsibility for the final version of the publication.

The study is part of the research project “Mapping the terrain and exploring the potential of generative artificial intelligence technology for firms and employees in Greece”. The research project is implemented in the framework of the Hellenic Foundation for Research and Innovation (H.F.R.I.) call “3rd Call for H.F.R.I.’s Research Projects to Support Faculty Members & Researchers” (H.F.R.I. project number: 23498).

The supplementary material for this article can be found online.

Aguinis
,
H.
and
Bradley
,
K.J.
(
2014
), “
Best practice recommendations for designing and implementing experimental vignette methodology studies
”,
Organizational Research Methods
, Vol. 
17
No. 
4
, pp. 
351
-
371
, doi: .
Albert
,
E.A.
(
2019
), “
AI in talent acquisition: a review of AI-applications used in recruitment and selection
”,
Strategic HR Review
, Vol. 
18
No. 
5
, pp. 
215
-
221
, doi: .
Almaraghi
,
D.M.Q.
(
2024
), “
Artificial intelligence and the future of human resource management: data analysis and decision guidance to enhance work productivity and employee motivation
”,
Pakistan Journal of Life and Social Sciences (PJLSS)
, Vol. 
22
No. 
2
, pp. 
12738
-
12748
, doi: .
Balasundaram
,
S.
and
Venkatagiri
,
S.
(
2020
), “
A structured approach to implementing Robotic Process Automation in HR
”,
Journal of Physics: Conference Series
, Vol. 
1427
No. 
1
, 012008, doi: .
Bick
,
A.
,
Blandin
,
A.
and
Deming
,
D.
(
2024
), “
The rapid adoption of generative AI
”,
working paper No. 32966, National Bureau of Economic Research, Cambridge, MA, September
.
Björkman
,
I.
,
Ehrnrooth
,
M.
,
Makela
,
K.
,
Smale
,
A.
and
Sumelius
,
J.
(
2014
), “
From HRM practices to the practice of HRM: setting a research agenda
”,
Journal of Organizational Effectiveness. People and Performance
, Vol. 
1
No. 
2
, pp. 
122
-
140
, doi: .
Bondarouk
,
T.
and
Brewster
,
C.
(
2016
), “
Conceptualising the future of HRM and technology research
”,
International Journal of Human Resource Management
, Vol. 
27
No. 
21
, pp. 
2652
-
2671
, doi: .
Brewer
,
M.
(
2000
), “Research design and issues of validity”, in
Reis
,
H.T.
and
Judd
,
C.M.
(Eds),
Handbook of Research Methods in Social and Personality Psychology
,
Cambridge University Press
,
Cambridge
, pp. 
3
-
16
.
Brynjolfsson
,
E.
,
Li
,
D.
and
Raymond
,
L.
(
2023
), “
Generative AI at work
”,
working paper No. 31161, National Bureau of Economic Research, Cambridge, MA, November
.
Brynjolfsson
,
E.
,
Chandar
,
B.
and
Chen
,
R.
(
2025
), “
Canaries in the coal mine: six facts about the recent employment effects of artificial intelligence
”,
working paper, Stanford Digital Economy Lab, Stanford, CA, 13 November
.
Budhwar
,
P.
,
Malik
,
A.
,
De Silva
,
M.T.
and
Thevisuthan
,
P.
(
2022
), “
Artificial intelligence - challenges and opportunities for international HRM: a review and research agenda
”,
International Journal of Human Resource Management
, Vol. 
33
No. 
6
, pp. 
1065
-
1097
, doi: .
Cantoni
,
F.
and
Mangia
,
G.
(
2019
),
Human Resource Management and Digitalization
,
Routledge
,
London; NY
.
Chang
,
S.H.
(
2017
), “
The effects of test trial and processing level on immediate and delayed retention
”,
Europe's Journal of Psychology
, Vol. 
13
No. 
1
, pp. 
129
-
142
, doi: .
Choi
,
J.H.
,
Monahan
,
A.
and
Schwarcz
,
D.
(
2023
), “
Lawyering in the age of artificial intelligence
”,
Minnesota Legal Studies Research Paper No. 23-31, University of Minnesota, Minneapolis, MN, 7 November
.
Collins
,
C.J.
and
Clark
,
K.D.
(
2003
), “
Strategic human resource practices, top management team social networks, and firm performance: the role of human resource practices in creating organizational competitive advantage
”,
Academy of Management Journal
, Vol. 
46
No. 
6
, pp. 
740
-
751
, doi: .
Davenport
,
T.H.
and
Mittal
,
N.
(
2022
), “
How generative AI is changing creative work
”,
Harvard Business Review
,
Published online
 14 November 2022,
available at:
 Link to the website (
accessed
 30 July 2025).
Dell' Acqua
,
F.
(
2022
), “
Falling asleep at the wheel: human/AI collaboration in a field experiment on HR recruiters
”,
available at:
 Link to the website (
accessed
 30 July 2025).
Dell’ Acqua
,
F.
,
McFowland
,
E.
,
Mollick
,
E.
,
Lifshitz-Assaf
,
H.
,
Kellogg
,
K.C.
,
Rajendran
,
S.
,
Krayer
,
L.J.
,
Candelon
,
F.
and
Lakhani
,
K.R.
(
2023
), “
Navigating the jagged technological Frontier: field experimental evidence of the effects of AI on knowledge worker productivity and quality
”,
Working paper No. 24-013, Social Science Research Network, Rochester, NY, 22 September
.
Elliott
,
S.A.
and
Brown
,
J.S.L.
(
2002
), “
What are we doing to waiting list controls?
”,
Βehaviour Research and Therapy
, Vol. 
40
No. 
9
, pp. 
1047
-
1052
, doi: .
Espegren
,
Y.
and
Hugosson
,
M.
(
2025
), “
HR analytics-as-practice: a systematic literature review
”,
Journal of Organizational Effectiveness. People and Performance
, Vol. 
12
No. 
5
, pp. 
83
-
111
, doi: .
Ferm
,
L.
,
Wallo
,
A.
,
Reineholm
,
C.
and
Lundqvist
,
D.
(
2024
), “
Defender, disturber or driver? The ideal-typical professional identities of HR practitioners
”,
Personnel Review
, Vol. 
53
No. 
6
, pp. 
1524
-
1541
, doi: .
Gray
,
D.E.
(
2021
),
Doing Research in the Real World
, (5th ed.) ,
SAGE Publications
,
London
.
Häll
,
A.
,
Tengblad
,
S.
,
Oudhuis
,
M.
and
Dellve
,
L.
(
2023
), “
How hard can it be? A qualitative study following an HRT implementation in a global industrial corporate group
”,
Personnel Review
, Vol. 
52
No. 
5
, pp. 
1632
-
1646
, doi: .
Hedrick
,
T.E.
,
Bickman
,
L.
and
Rog
,
D.J.
(
1993
),
Applied Research Design: A Practical Guide
,
SAGE Publications
,
London
.
Hermann
,
E.
,
Puntoni
,
S.
and
Morewedge
,
C.K.
(
2025
), “
GenAI and the psychology of work
”,
Trends in Cognitive Sciences
, Vol. 
29
No. 
9
, pp. 
802
-
813
, doi: .
Kambur
,
E.
and
Akar
,
C.
(
2022
), “
Human resource developments with the touch of artificial intelligence: a scale development study
”,
International Journal of Manpower
, Vol. 
43
No. 
1
, pp. 
168
-
205
, doi: .
Karakas
,
C.
,
Brock
,
D.
and
Lakhotia
,
A.
(
2023
), “
Leveraging ChatGPT in the pediatric neurology clinic: practical considerations for use to improve efficiency and outcomes
”,
Pediatric Neurology
, Vol. 
148
, pp. 
157
-
163
, doi: .
Kim
,
S.
,
Wang
,
Y.
and
Boon
,
C.
(
2021
), “
Sixty years of research on technology and human resource management: looking back and looking forward
”,
Human Resource Management
, Vol. 
60
No. 
1
, pp. 
229
-
247
, doi: .
Kim
,
T.
,
Park
,
Y.
and
Kim
,
W.
(
2022
), “
The impact of artificial intelligence on firm performance
”,
Proceedings of the 2022 Portland International Conference on Management of Engineering and Technology (PICMET)
,
Portland, OR, USA
, pp. 
1
-
10
, doi: .
Koss
,
E.
and
Lewis
,
D.A.
(
1993
), “
Productivity or efficiency– measuring what we really want
”,
National Productivity Review
, Vol. 
12
No. 
2
, pp. 
273
-
284
, doi: .
Legge
,
K.
(
2005
),
Human Resource Management: Rhetorics and Realities
, (Anniversary ed.) ,
Red Globe Press
,
London
.
Lo
,
K.
,
Macky
,
K.
and
Pio
,
E.
(
2015
), “
The HR competency requirements for strategic and functional HR practitioners
”,
International Journal of Human Resource Management
, Vol. 
26
No. 
18
, pp. 
2308
-
2328
, doi: .
Malik
,
A.
,
Budhwar
,
P.
and
Kazmi
,
B.A.
(
2023
), “
Artificial intelligence (AI)-assisted HRM: towards an extended strategic framework
”,
Human Resource Management Review
, Vol. 
33
No. 
1
, 100940, doi: .
Mussweiler
,
T.
and
Strack
,
F.
(
2000
), “Consequences of social comparison: selective accessibility, assimilation, and contrast”, in
Suls
,
J.
and
Wheeler
,
L.
(Eds),
Handbook of Social Comparison: Theory and Research
,
Kluwer Academic
,
Dordrecht
, pp. 
253
-
270
.
Nankervis
,
A.
,
Connell
,
J.
,
Cameron
,
R.
,
Montague
,
A.
and
Prikshat
,
V.
(
2021
), “
Are we there yet? Australian HR professionals and the fourth industrial revolution
”,
Asia Pacific Journal of Human Resources
, Vol. 
59
No. 
1
, pp. 
3
-
19
, doi: .
Narzary
,
S.
,
Makhecha
,
U.P.
,
Budhwar
,
P.
,
Malik
,
A.
and
Kumar
,
S.
(
2025
), “
A temporal evolution of human resource management and technology research: a retrospective bibliometric analysis
”,
Personnel Review
, Vol. 
54
No. 
3
, pp. 
800
-
823
, doi: .
Niehueser
,
W.
and
Boak
,
G.
(
2020
), “
Introducing artificial intelligence into a human resources function
”,
Industrial and Commercial Training
, Vol. 
52
No. 
2
, pp. 
121
-
130
, doi: .
Nishii
,
L.H.
and
Wright
,
P.M.
(
2008
), “Variability within organizations: implications for strategic human resources management”, in
Smith
,
D.B.
(Ed.),
The People Make the Place: Dynamic Linkages between Individuals and Organizations
,
Taylor & Francis Group/Lawrence Erlbaum Associates
,
New York
, pp. 
225
-
248
.
Noy
,
S.
and
Zhang
,
W.
(
2023
), “
Experimental evidence on the productivity effects of generative artificial intelligence
”,
Science
, Vol. 
381
No. 
6654
, pp. 
187
-
192
, doi: .
Paauwe
,
J.
and
Boon
,
C.
(
2018
), “Strategic HRM: a critical review”, in
Collins
,
D.
,
Wood
,
G.
and
Szamosi
,
L.
(Eds),
Human Resource Management: A Critical Approach
, (2nd ed.) ,
Routledge
,
London
, pp. 
49
-
73
.
Pan
,
Y.
and
Froese
,
F.J.
(
2023
), “
An interdisciplinary review of AI and HRM: challenges and future directions
”,
Human Resource Management Review
, Vol. 
33
No. 
1
, 100924, doi: .
Peng
,
S.
,
Kalliamvakou
,
E.
,
Cihon
,
P.
and
Demirer
,
M.
(
2023
), “
The impact of AI on developer productivity: evidence from GitHub Copilot
”,
available at:
 Link to the website (
accessed
 30 July 2025).
Persson
,
M.
and
Wallo
,
A.
(
2024
), “
Digital automation and working life of HR practitioners: a gender analysis of the implications for workforce and work practices
”,
Gender, Technology and Development
, Vol. 
28
No. 
3
, pp. 
408
-
427
, doi: .
Rabenu
,
E.
and
Baruch
,
Y.
(
2025
), “
Cyborging HRM theory: from evolution to revolution – the challenges and trajectories of AI for the future role of HRM
”,
Personnel Review
, Vol. 
54
No. 
1
, pp. 
174
-
198
, doi: .
Rahi
,
S.
(
2017
), “
Research design and methods: a systematic review of research paradigms, sampling issues, and instruments development
”,
International Journal of Economics and Management Sciences
, Vol. 
6
No. 
2
, 1000403, doi: .
Saunders
,
M.
,
Lewis
,
P.
and
Thornhill
,
A.
(
2012
),
Research Methods for Business Students
,
Pearson Education
,
Harlow
.
Shotter
,
J.
and
Tsoukas
,
H.
(
2014
), “
Performing phronesis: on the way to engaged judgment
”,
Management Learning
, Vol. 
45
No. 
4
, pp. 
377
-
396
, doi: .
Sink
,
D.S.
and
Tuttle
,
T.C.
(
1989
),
Planning and Measurement in Your Organisation of the Future
,
Industrial Engineering and Management Press
,
Norcross, GA
, pp. 
170
-
184
,
ch. 5
.
Upadhyay
,
A.K.
and
Khandelwal
,
K.
(
2018
), “
Applying artificial intelligence: implications for recruitment
”,
Strategic HR Review
, Vol. 
17
No. 
5
, pp. 
255
-
258
, doi: .
Vrontis
,
D.
,
Christofi
,
M.
,
Pereira
,
V.
,
Tarba
,
S.
,
Makrides
,
A.
and
Trichina
,
E.
(
2022
), “
Artificial intelligence, robotics, advanced technologies and human resource management: a systematic review
”,
International Journal of Human Resource Management
, Vol. 
33
No. 
6
, pp. 
1237
-
1266
, doi: .
Wallo
,
A.
and
Coetzer
,
A.
(
2023
), “
Understanding and conceptualising the daily work of human resource practitioners
”,
Journal of Organizational Effectiveness. People and Performance
, Vol. 
10
No. 
2
, pp. 
180
-
198
, doi: .
Wirtz
,
B.W.
,
Weyerer
,
J.C.
and
Geyer
,
C.
(
2019
), “
Artificial intelligence and the public sector— applications and challenges
”,
International Journal of Public Administration
, Vol. 
42
No. 
7
, pp. 
596
-
615
, doi: .
Wright
,
C.
(
2008
), “
Reinventing human resource management: business partners, internal consultants and the limits to professionalization
”,
Human Relations
, Vol. 
61
No. 
8
, pp. 
1063
-
1086
, doi: .
Zhou
,
E.
and
Lee
,
D.
(
2024
), “
Generative artificial intelligence, human creativity, and art
”,
PNAS Nexus
, Vol. 
3
No. 
3
, pgae052, doi: .
Published by Emerald Publishing Limited. This article is published under the Creative Commons Attribution (CC BY 4.0) license. Anyone may reproduce, distribute, translate and create derivative works of this article (for both commercial and non-commercial purposes), subject to full attribution to the original publication and authors. The full terms of this license may be seen at Link to the terms of the CC BY 4.0 licence.

Supplementary data

or Create an Account

Close Modal
Close Modal