This study critically examines how gender inequalities manifest in performance evaluations within professional workplaces. While initially focused on whether and how women's evaluations change after maternity, the review broadens to gendered mechanisms more generally, with sustained attention to motherhood-related dynamics such as flexibility stigma, ideal-worker norms and care responsibilities. By synthesizing 54 empirical studies published between 2009 and 2024, the paper identifies structural and interactional processes that reproduce inequalities, evaluates the evidence on motherhood penalties, and highlights gaps in longitudinal and sectoral coverage to guide future research.
The review followed PRISMA 2020 guidelines to ensure transparency and replicability. Searches were conducted in Scopus and Web of Science (2009–2024) using “performance evaluation,” “gender,” and “work,” restricted to English and Italian publications in the social sciences, management and economics. Only empirical studies addressing individual evaluations in hierarchical workplaces were included; exclusions covered non-hierarchical, clinical, or educational contexts and studies lacking a gender perspective. After screening 1,005 records and applying eligibility criteria, 54 studies were retained. An inductive coding process identified five recurrent themes structuring the synthesis.
The review shows that performance evaluations, though formally gender-neutral, consistently reproduce inequalities. Bias enters at multiple levels: criteria setting and managerial incentives, organizational practices shaped by ideal-worker norms, evaluation tools and rating scales, and socio-demographic dynamics in evaluator–evaluatee interactions. Mothers face distinctive penalties through flexibility stigma, time norms and stereotype-driven interpretations of behavior and feedback. Evidence on pre-/post-maternity evaluations remains scarce, limiting causal generalizability. Overall, the findings underscore that motherhood penalties constitute a distinct structural pathway of disadvantage, requiring targeted research and policy interventions.
The review is limited to peer-reviewed empirical studies published in English between 2009 and 2024, potentially excluding relevant insights from grey literature or non-English contexts. Most included studies are based in Western countries, raising concerns about global generalizability. Future research should expand geographical coverage and investigate intersectional dimensions more systematically.
Organizations should critically assess the design and implementation of performance evaluation tools to mitigate embedded gender biases. Practices such as structured feedback, joint evaluations and transparency in promotion decisions may reduce disparities. HR departments must also be aware of how motherhood and care responsibilities are systematically penalized in current systems.
Gender-biased performance evaluations contribute to systemic inequalities in income, career advancement and job security. This review underscores the need for more inclusive evaluation frameworks that recognize diverse forms of labor and challenge masculine norms embedded in organizational practices, particularly in relation to care and work-life balance.
This paper provides the first systematic review synthesizing empirical evidence on gender inequalities in performance evaluation with explicit attention to motherhood. By integrating 54 studies across sociology, psychology, management and economics, it highlights how structural, cultural and interactional mechanisms converge to produce persistent disparities. The focus on motherhood penalties—ideal-worker time norms, flexibility stigma and stereotype-driven feedback—advances understanding beyond general gender bias. The review also identifies key blind spots, including limited longitudinal evidence and sectoral coverage, offering a comprehensive agenda for future research and practical implications for equitable evaluation systems.
1. Introduction
Performance evaluations are widely portrayed as neutral and meritocratic devices that reward effort and talent, linking individual merit to career advancement and financial benefits (Castilla, 2008; Alon and Tienda, 2007). Yet research consistently shows that they reproduce inequalities, with tangible consequences for pay, promotion and careers (Correll et al., 2007; Bowers and Prato, 2018; Abraham, 2020; Rivera, 2020; Botelho and Gertsberg, 2022; Gorbatai et al., 2023). This tension raises an urgent question: if evaluation systems formally rest on objectivity and fairness, why do gendered disparities persist across occupations and countries?
Although appraisal systems are designed to minimize bias (Castilla, 2008; Castilla and Benard, 2010; Joshi et al., 2015; Rivera and Tilcsik, 2019), their supposed neutrality often masks persistent gendered and motherhood-related biases. Even the most formalized and transparent systems can reinforce patriarchal norms, leading to what Castilla and Benard (2010) termed the “paradox of meritocracy” (Koskinen-Sandberg, 2017; Kelley et al., 2024). As a result, women—particularly mothers—continue to experience systematic disadvantage in professional outcomes (England et al., 2020; Stainback and Tomaskovic-Devey, 2012, 2016).
Motherhood has therefore emerged as a paradigmatic test case of how organizational and social role expectations intersect to penalize women (Hodges and Budig, 2010; Budig and England, 2001; Waldfogel, 1997, 1998b). Yet our review shows that this phenomenon is embedded in a broader architecture of gendered mechanisms that affect performance evaluation more generally. Studies have linked the motherhood penalty to factors such as education (Budig et al., 2021; Anderson et al., 2003), timing of maternity (Amuedo-Dorantes and Kimmel, 2005; Taniguchi, 1999), age cohorts (Avellar and Smock, 2003; Waldfogel, 1998a), and race or minority status (Glauber, 2008; Korenman and Neumark, 1992). Some authors also highlight the interaction with pay-for-performance systems, where women's lower ratings—actual or perceived—amplify disparities (Castilla, 2008, 2015; Roth et al., 2012; Bono et al., 2017; Rivera and Tilcsik, 2019). However, evidence remains fragmented, and causality between performance evaluations and maternity leave is difficult to establish.
Our initial research question targeted changes in evaluations after maternity. During synthesis, however, we observed that directly maternity-focused studies were limited, while a larger body of evidence addresses broader gendered mechanisms. To ensure conceptual coherence, we broadened our question to “How and where do gendered mechanisms shape individual performance evaluations, with particular attention to motherhood-related dynamics?”.
Accordingly, our coding identified five macro-themes that capture these mechanisms, alongside a specific section on motherhood dynamics as a distinctive axis of inequality. The review, conducted following the PRISMA – Preferred Reporting Items for Systematic Reviews and Meta-Analyses – Statement (Page et al., 2021), systematically collates and synthesizes studies to frame performance evaluation within the broader landscape of organizational inequalities, while highlighting how flexibility stigma, ideal-worker norms and care responsibilities continue to penalize mothers and point to the need for structural change.
2. Theoretical framework
This framework situates performance evaluation studies within broader theories of gendered organizations and inequality regimes (Acker, 2006), offering the conceptual lens for interpreting the empirical findings. Joan Acker's (1990) theory of gendered organizations challenges the assumption of organizational neutrality, showing how norms and hierarchies are inherently gendered and privilege those who conform to the “ideal worker” model (Williams, 2001). From this perspective, performance evaluations—often presented as objective and merit-based—function as organizational practices that reproduce inequality. Motherhood is especially significant, as care responsibilities are constructed as incompatible with dominant norms of productivity and commitment.
Complementing this structural account, social role theory (Eagly and Karau, 2002; Eagly et al., 1992; Biddle, 1979) explains how descriptive and prescriptive stereotypes shape evaluative contexts. Expectations about competence, leadership and availability generate systematic double standards: behaviors rewarded in men are often penalized in women, while caregiving reinforces assumptions of lower ambition. These interactional mechanisms illustrate why even formalized systems remain vulnerable to bias. As Bernhardt et al. (1995) note, focusing only on average outcomes obscures how recognition and rewards are redistributed through reliance on gendered stereotypes.
The motherhood penalty exemplifies the convergence of structural and role-based dynamics. Evidence shows that mothers face systematic disadvantages in wages, promotions and evaluations (Meldgin et al., 2024; Wald et al., 2024; Budig and Hodges, 2010; Budig and England, 2001; Waldfogel, 1997, a, b; Harkness and Waldfogel, 1999), persisting even when controlling for individual and occupational characteristics (Anderson et al., 2003; Korenman and Neumark, 1992; Lundberg and Rose, 2000). Some have argued that mothers trade wages for family-friendly arrangements (Becker, 1992; Glass, 2004), yet studies show penalties endure even in organizations with generous benefits (Budig and Hodges, 2010). This persistence underscores the central role of employer discrimination—conscious or not—across hiring, pay and career progression (Correll et al., 2007).
Taken together, these perspectives provide a comprehensive framework: gendered organizations theory highlights the structural embedding of inequality, social role theory illuminates micro-level mechanisms, and the motherhood penalty demonstrates how they converge in practice. This synthesis clarifies why evaluation systems, though formally gender-neutral, often legitimize disparities. It also provides the conceptual basis for our systematic review and informs the categorization of findings into five thematic areas: normative frames of fairness, organizational practices, methodologies and tools, evaluation biases and care responsibilities.
3. Methodology and review context
We conducted the review following the PRISMA 2020 framework (Preferred Reporting Items for Systematic Reviews and Meta-Analyses; Page et al., 2021). Originally developed in medical research, PRISMA is now widely applied across disciplines for its emphasis on transparency, replicability and rigor. Its checklist and flow diagram (Figure 1) guided our reporting and ensured clarity, particularly in documenting inclusion and exclusion criteria. Compared to scoping reviews—often used for exploratory purposes—PRISMA imposes greater methodological stringency (Munn et al., 2018), aligning with our aim to produce not only a descriptive overview but also a reproducible synthesis of gendered mechanisms in performance evaluations.
Studies were eligible if they (1) were empirical, (2) examined individual performance evaluations in hierarchical workplace settings and (3) explicitly addressed gendered mechanisms or motherhood-related dynamics. We included peer-reviewed journal articles and, when indexed in Scopus or Web of Science, conference proceedings of comparable quality. Exclusion criteria ruled out non-hierarchical contexts (education, sport, clinical, self-employment), studies of diversity without a link to evaluations and research lacking a gender perspective.
The review covered 2009–2024 and was restricted to English and Italian publications in Social Sciences, Business, Management and Accounting, and Economics. Searches in Scopus and Web of Science using the keywords “performance evaluation,” “gender,” and “work” yielded 1,005 records. After deduplication (n = 68), 937 titles/abstracts were screened; 736 were excluded, leaving 201 articles for full-text review. Fifty-four empirical studies were retained, along with one book chapter (Caleo and Heilman, 2014) and two indexed conference proceedings (Chicu and Nedelcu, 2017; Marchegiani et al., 2018). Full details appear in the PRISMA flow diagram (Figure 1).
The final corpus addresses the research question: “How and where do gendered mechanisms shape performance evaluations, and what is known about motherhood-specific dynamics?” The empirical focus is on professional workplaces with hierarchical structures and standardized performance management processes, often codified in formal metrics or collective bargaining agreements (Botelho, 2024).
The thematic synthesis was conducted manually after full-text reading of all included records. An inductive coding strategy was adopted: initial open codes captured recurrent mechanisms, evaluative contexts and gendered dynamics, later consolidated into five macro-themes: (1) accountability, transparency, justice and meritocracy; (2) organizational practices; (3) methodologies, tools and techniques; (4) sociodemographic characteristics of evaluator and evaluated, including bias and (5) care responsibilities. Coding was carried out without software, relying on manual interpretation to capture nuances. We treated coding as a conceptual device rather than a rigid classificatory tool (Ezzy, 2002), consistent with grounded theory and thematic analysis traditions. Categories remained open and iteratively refined. Following Schreier (2012), overlaps across categories are not errors but reflect the ambiguity of social phenomena. Accordingly, some studies appear under multiple themes, not redundantly but because they illuminate different mechanisms depending on the analytic lens. This approach avoids oversimplification while enabling a more nuanced synthesis of gendered mechanisms in evaluations.
Theme definitions are detailed in Table 1 (Appendix 1), while Figure 2 illustrates the distribution of studies across macro-themes.
4. Systematic literature review (SLR)
The final sample includes 54 studies published between 2009 and 2024 across 41 journals (Appendix 2a and 2b). The sample includes one book chapter (Caleo and Heilman, 2014) and two conference proceedings indexed in Scopus/Web of Science (Chicu and Nedelcu, 2017; Marchegiani et al., 2018). Disciplinary coverage spans management, psychology, sociology and gender studies (Figure 3). Publication trends show growth until 2017, followed by a decline in recent years (Figure 4). This suggests a potential shift of scholarly attention toward adjacent topics, despite the persistent relevance of gender inequalities in performance evaluation.
Findings are organized around the five macro-themes, indicating for each (1) where in the evaluation process it operates and (2) how it produces differential ratings. A study-by-theme mapping is provided in Appendix 2. While these themes capture general mechanisms, we also devote a separate section to motherhood-specific dynamics, in line with our research question and with the recognition that maternal status represents a distinct axis of inequality (Collins, 1998).
4.1 Normative frames of fairness
This macro-theme captures the normative frames of fairness that formally underpin performance evaluations—namely, accountability, transparency, justice and meritocracy. While these ideals are meant to guarantee neutral rewards for effort and talent, the reviewed studies show how they often coexist with, or even mask, gendered biases.
Inesi and Cable (2015), analyzing the US military, found that male evaluators high in social dominance orientation downgraded female subordinates when perceiving threats to gender hierarchy, suggesting that more objective practices could mitigate this bias. However, as will be discussed later, many other studies have refuted this hypothesis.
Castilla (2015), using personnel data from a large firm (n = 9,000 employees), observed that accountability policies reduced gender disparities, as managers became more cautious in assessments. Yet Dobbin et al. (2015) reported that few accountability reforms produced lasting effects on segregation and inequality. Earlier, Castilla (2008) showed significant penalties for women, minorities and immigrants in bonus allocation but not in promotions or layoffs, attributing differences partly to the greater visibility of the latter. From a managerial perspective, the importance of transparency for equitable performance evaluations and the value attributed to meritocracy have also been explored. However, the work by Doan-Goss et al. (2023) focused narrowly on Canada's highly gender-segregated transport sector, with a very small sample (n = 9).
A broader analysis was conducted in 2010 by Castilla and Bernard on a sample of 445 participants with managerial experience, who were asked to make recommendations on bonuses, promotions and dismissals for various employee profiles. The authors identified a significant discrepancy in bonus allocations to the detriment of women, leading to theorize what they termed the «paradox of meritocracy». More recently, Amis et al. (2020) revisit this concept, classifying meritocracy among what they call “institutional myths”—alongside efficiency and positive globalization. They emphasize how the presumed neutrality inherent to organizational structures directly and indirectly contributes to the reproduction of inequality within and by organizations. This is why Treviño et al. (2018) question whether these systems are better understood as “masculinities,” given their calibration to patriarchal, male-dominated norms.
Taken together, these studies suggest that bias arises mainly in criteria setting and incentive alignment, while accountability and transparency can mitigate gaps by limiting discretion and clarifying standards.
4.2 Organizational practices
The analysis by Amis et al. (2020) is supported by other studies focused on organizational practices and their role in preventing and counteracting the (re)production of gender inequalities. Abraham et al. (2024) identify three factors: prevailing beliefs, the design of evaluation processes and evaluators' characteristics. Castilla (2008, 2011, 2015) similarly showed that disparities emerge in evaluation processes, from which pay and promotion gaps derive.
Performance management practices have also been extensively analyzed to understand whether gender differences can be observed in their perception. For instance, Festing et al. (2015), in a cross-country meta-analysis, found that female managers preferred group-based over individual evaluations, suggesting a different orientation toward collaboration. Consistently, Dalmia and Filiz-Ozbay (2021) argue that HR practices promoting collaboration between high- and low-performing employees can boost women's motivation. Experimental evidence from Bohnet et al. (2016) – albeit experimentally at the Harvard Decision Science Lab – indicates that joint evaluations can eliminate gender disparities, though outcomes may depend on how individual contributions are recognized (Meuris and Elias, 2023), as discussed in the next section.
The adoption of objective evaluation systems has long been advocated in performance management (Inesi and Cable, 2015); however, Koskinen-Sandberg (2017) argues that so-called gender-neutral evaluation systems are legitimized by the patriarchal culture that permeates organizational structures. The author coins the term “gender-neutral legitimacy” to describe how these systems, when intertwined with external societal mechanisms, contribute to the persistence of gender pay and career gaps. It is therefore crucial that the adoption of organizational practices favouring disadvantaged groups be carried out with caution and awareness of their potential side effects. This concern is highlighted in a meta-analysis by Leslie et al. (2014) and more recently confirmed by Kelley et al. (2024), which shows that well-intentioned initiatives may backfire, lowering perceptions of women's competence and motivation—a dynamic linked to Castilla's (2008) “paradox of meritocracy”.
Hence, performance evaluation systems should be aligned with organizational inclusion policies (Kelley et al., 2024). Additionally, according to the thirty-year meta-analysis by Joshi et al. (2015), performance pay reinforces hierarchies in high-prestige contexts, penalizing those unable to meet “ideal worker” expectations of 24/7 availability. Family-friendly practices may mitigate these effects: Boyar et al. (2016) and Katou (2022) show that caregiver support improves well-being and organizational outcomes, while Crowley (2013) documents how unattainable goals and part-time exclusions disadvantage working mothers.
Nonetheless, the lack of longitudinal evidence makes it difficult to confirm whether family-friendly policies durably reduce disparities (Hing et al., 2023). Overall, findings suggest that time norms, eligibility rules and pay–performance linkages define what counts as performance, systematically disadvantaging women—particularly mothers—under rigid “ideal worker” regimes.
4.3 Methodology, tools and techniques
Although merit is often invoked to counter gender discrimination—despite critiques of its paradoxical nature (Amis et al., 2020; Castilla, 2011)—specific performance management techniques can promote more equitable outcomes (Hing et al., 2023). Suggested tools include structured interviews, 360-degree feedback, transparent salary negotiations and mentorship programs. Abraham et al. (2024) confirm that evaluation design strongly influences inequality. Although not empirically tested, other studies confirm disparities emerge early in the evaluation process: in multistage evaluations, both Botelho and Abraham (2017) and Castilla (2011) find that double standards tend to appear primarily in the early steps of the process, gradually diminishing as the evaluator acquires more information about the employee. However, according to Castilla (2011), these managerial mechanisms are potentially operative across many phases of an employee's career—not only in performance evaluations, but also in pay, training opportunities, career advancement and task assignments—thus contributing to the persistence of workplace stratification.
Contrary to the findings of Botelho and Abraham (2017) and Castilla (2011), a recent longitudinal study by Brewer et al. (2020), which analyzed the performance of 2,765 medical residents, found that evaluation disparities widened with increasing experience and seniority. According to the authors, role expectations and implicit biases that may be triggered over time significantly contribute to the production of gender inequality.
The varying results obtained, and the limited scope of existing studies, make it difficult to clearly identify a single phase of the evaluation process where double standards consistently emerge.
Nonetheless, all three studies emphasize the importance of information in promoting more equitable performance evaluations: Meuris and Elias (2023) note that in the absence of concrete data on individual contributions, supervisors fall back on stereotypes, while Bohnet et al. (2016) experimentally demonstrate that joint evaluations can neutralize disparities—though only in lab settings. Consequently, the authors urge organizations to rethink their evaluation systems by incorporating multiple sources of information, including 360-degree feedback, as also recommended by Hing et al. (2023). In any case, when implemented with due consideration, group evaluation appears promising.
However, evidence on 360-degree feedback is mixed: Blanch-Hartigan (2011) warns that women's self-undervaluation can depress scores, whereas Grund and Soboll (2024) find no differences when monetary incentives are considered. Chicu and Nedelcu (2017) even propose integrating work–life balance and family feedback, though without empirical support.
Beyond the evaluation methodology (such as 360-degree, joint or peer evaluation) and techniques (such as interviews, information gathering and informal assessments), the literature reviewed in this systematic analysis has also explored the importance of the evaluation tools themselves.
Specifically, Rivera and Tilcsik (2019) show that a 6-point rather than 10-point scale reduced gaps in male-dominated fields. The authors attribute this result to cultural differences and stereotypes that evaluators associate with specific numerical scales.
Moreover, Rivera et al. (2021) find that real-time feedback apps altered patterns of who rated whom and shaped future scores. Informal assessments, too, can disadvantage women when they reduce mentorship and sponsorship opportunities (Bono et al., 2017). The authors also note that positive feedback had a stronger impact on future evaluations than negative feedback.
All the studies examined—despite offering highly diverse methodological approaches, samples and findings—converge on the importance of clarity in evaluation methodologies, tools and techniques as essential for ensuring fairness. Ambiguity, which according to Caleo and Heilman (2014) may concern (1) the quantity and quality of information, (2) the evaluation criteria, (3) the structure of the assessment and (4) group work, creates excessive room for stereotypes to operate. This in turn exacerbates what Koskinen-Sandberg (2017) terms gender-neutral legitimated gaps, which continue to (re)produce inequalities.
Evaluation designs that reduce ambiguity—such as clearer rating scales, joint evaluations with individual-contribution data and carefully implemented multi-source feedback—can attenuate gender differentials in scores.
4.4 Sociodemographic characteristics of evaluators and evaluatees/evaluation bias
As previously noted, gender stereotypes play a decisive role in shaping performance evaluations (Abraham et al., 2024). Castilla (2011) identifies three mechanisms through which managers influence employee assessments: (1) the influence of the social network among current and former managers; (2) manager-to-manager homophily (horizontal); and (3) manager-to-employee homophily (vertical). His longitudinal analysis shows that surrounding contexts—more than individual stereotypes—distort outcomes. Vignette-based studies corroborate this, revealing pro-male bias in ratings and recommendations (Bloomfield et al., 2021; Chattopadhyay, 2021), as well as harsher treatment of women who violate prescriptive stereotypes (He et al., 2024). Gendered interpretations extend to “potential”: men are more often rated as high-potential, while women's expressions of passion are penalized.
Further evidence highlights paradoxical double standards. Using the same vignette-based method as Bloomfield et al. (2021), but with a larger sample, Bode et al. (2022) find that socially impactful work may hurt men's promotion chances but not women's. Equally paradoxical—because it stems from the reinforcement of a stereotype—is the discovery of a double standard favouring women, who are evaluated more positively than men when they show adaptability to teamwork (Carpini et al., 2023) and less positively when they display high levels of agreeableness (Nandkeolyar et al., 2022).
These findings are relatively new within the field of performance management research, which has emphasized the disadvantages of such systems for marginalized groups. This dimension has also been explored through an intersectional lens (Chatman et al., 2022; Bohlmann and Zacher, 2020), showing that bias compounds with age, disadvantaging older women.
Gaining more experience may also prove detrimental to gender equality in performance evaluations. The study by Brewer et al. (2020), conducted on more than 2,000 medical residents, found that men and women were perceived as equally competent at the start of their specialization, when the student role predominates. However, by the third year—when the role of colleague becomes more central—men were perceived as more competent than women. Additionally, the authors observed that, when controlling for resident performance, women received harsher and more critical feedback than men after committing medical errors of similar severity, indicating a lack of supervisory support and, thereby, reducing their chances of promotion (Bono et al., 2017).
The literature has also focused on the individual characteristics of the employee being evaluated. A meta-analysis by Blanch-Hartigan (2011) identified differences in self-evaluations between men and women, highlighting that women's tendency to underestimate themselves may constitute a potentially harmful factor. Another important element concerns the value attributed to merit-based rewards: according to the findings of Froese et al. (2019), such rewards are considered more important by male employees and those with higher levels of education. This aligns with earlier findings by Jonnergård et al. (2010), who observed that performance evaluations are perceived differently: men emphasize what is evaluated, reflecting hierarchical perspectives, while women emphasize who evaluates and how. These perceptions also influence career trajectories: by analyzing the characteristics of newly hired and recently certified auditors, the authors found that women had achieved fewer outcomes and displayed lower levels of ambition and career expectations from the early stages of employment, along with a higher intention to leave the auditing profession. Ma (2022) further supports this by showing that, on average, female CEOs exhibit significantly higher turnover sensitivity than their male counterparts, as they tend to be evaluated more unfavorably when performance declines.
Regarding predictors of performance evaluation, as also discussed by He et al. (2024), Yaakobi and Weisberg (2018) found that manager gender moderates the weight assigned to different performance predictors. They attribute this to the differing emphasis each gender places on available resources: women tend to prioritize social resources, whereas men emphasize material ones, both at the group and organizational levels.
As already highlighted by Castilla (2011), the dyadic interaction between manager and employee plays a fundamental role in performance evaluation. Johansson and Wennblom (2017) conducted a vignette experiment in which only the name of the supervisor was changed to a typically male or female name in a 2x2 design involving subordinates (i.e. the interviewees). They found that female subordinates expressed more negative attitudes toward the fairness of the evaluation process. Moreover, male subordinates with a female supervisor appeared to place more trust in management than either males with a male supervisor or females with a female supervisor.
Nonetheless, as pointed out by Inesi and Cable (2015), when a male evaluator—typically with a high social dominance orientation—perceives a threat to the gender hierarchy, such as the potential career advancement (or promotion) of a female subordinate, he tends to underrate her performance. Reactions to unfair or inaccurate evaluations also appear to differ by gender: according to findings by Marchegiani et al. (2018), when sustained effort is not rewarded, men reduce engagement, whereas women persist.
Ciancetta and Roch (2021) also highlight gender differences in the language used in performance evaluations. Their findings reveal that women, across all hierarchical levels within the organization, receive a greater number of negative terms in the narratives used to describe their performance. Differences also emerge in how men and women respond to critical performance feedback: men report a significantly stronger negative effect in response to critical feedback than women, which may indicate a defensive adaptation developed in response to negative workplace experiences. These findings are consistent with earlier work by Correll et al. (2020) and Wynn and Carian (2023), who, in analysing the language used in performance evaluations, found that gender influences feedback—and consequently numeric ratings—in subtle yet significant ways. Such differences lead to important gendered patterns in the traits and behaviors that managers tend to reward with higher scores. For instance, being described as truly exceptional, “a visionary”, results in higher ratings for men but not for women. Similarly, forward-looking evaluations, such as comments suggesting that employees need to improve their technical skills to advance to the next level, lead to lower ratings for women but not for men. Evaluator traits and dyadic (dis)similarities show how stereotypes are activated in concrete interactions—particularly under uncertainty—shaping both ratings and narrative feedback.
4.5 Care responsibilities
Care responsibilities consistently emerge as a central determinant of performance evaluations. Chauhan et al. (2022), studying the Indian IT sector, show that family responsibilities, perceived support and mentoring shape women's perceptions of career success, though the narrow sample limits generalizability.
Other studies emphasize that family-friendly corporate policies can mitigate backlash effects—defined as negative career outcomes triggered by family demands (Crowley, 2013; Boyar et al., 2016; Kelley et al., 2024; Montanye and Livingston, 2024).
According to Snir (2019), who analyses requests for changes in working hours (increases or reductions) following the birth of a child, such desired shifts may already be occurring: findings show that high-performing parents—especially mothers—are perceived as both better workers and parents. However, it is important to note that despite the study's sizable sample (844 respondents), the research design is based on a vignette experiment combined with responses to an online survey conducted on an Israeli platform. Consequently, three elements limit the generalizability of Snir's (2019) findings: the highly experimental nature of the study, the absence of implementation in a real organizational setting and the fact that only a small portion of the sample may have had actual experience in evaluating employees.
Findings that appear more consistent with social role theory and the improvement proposals are the ones by Crowley (2013), Boyar et al. (2016), Kelley et al. (2024) and Montanye and Livingston (2024), which identify the continued persistence in many organizations to rely on the model of “ideal worker”—a standard that parents are not always able, nor should be expected, to meet. For instance, Steiner et al. (2022) document a backlash effect targeting employees with high family-related demands, regardless of gender. These individuals reported being systematically judged and treated less favourably by observers and supervisors than employees who prioritize work, particularly in terms of promotability and the assignment of bonuses and rewards.
Similar dynamics have been observed in the accounting sector: Sasmaz and Fogarty (2023), through an experimental design, show that supervisor support for employees' work–life balance significantly improves perceptions of promotability. Their findings reinforce the idea that family-related demands, when not adequately supported by organizational practices, systematically reduce advancement opportunities—particularly for women and parents.
Finally, the pioneering analysis by Waumsley and Houston, already in 2009, raised concerns about the impact of flexible work on performance evaluation, noting that it was perceived as detrimental to job performance and career progression compared to long or regular working hours. The fact that women make up the majority of those who utilize such work-life balance measures results in an indirect form of discrimination against mothers—mirroring what Crowley (2013) also observed in relation to the use of part-time employment.
Care responsibilities act as role-expectation signals that lower perceived commitment and potential; penalties intensify when organizations prioritize continuous availability over task outcomes.
5. Mother-specific dynamics in performance evaluation
As theorized by Collins (1998), motherhood represents not merely a personal life event but a structural position embedded in interlocking systems of gender, family roles and paid work. In this sense, the motherhood penalty cannot be reduced to a sub-dimension of gender bias but constitutes a distinctive mechanism of inequality. This theoretical standpoint justifies the decision to treat motherhood-related dynamics as an autonomous section within our findings, in line with our research question and highlights how organizational performance evaluation systems reproduce inequities through the specific lens of maternal status.
In fact, beyond general gendered mechanisms, evidence converges on motherhood-specific dynamics that systematically penalize mothers compared to other women: (1) Ideal-worker time norms: studies in healthcare, professional services and other high-demand settings show that norms of constant availability (24/7 responsiveness, long hours, uninterrupted presence) are tacitly embedded in performance standards; mothers are more often perceived as violating these norms, lowering ratings and promotability even when performance is comparable (Williams, 2001; Budig and Hodges, 2010; Crowley, 2013). (2) Flexibility and policy usage stigma: the use (or anticipated use) of part-time, flexible schedules or family-friendly policies carries a stigma of lower commitment; several studies report that such usage is interpreted as a negative performance signal, and eligibility rules for bonuses or promotions can de facto exclude policy users—mechanically reducing evaluations and rewards (Glass, 2004; Crowley, 2013). (3) Stereotype-consistent interpretations of behavior and feedback: qualitative and experimental evidence shows that identical errors, feedback, or teamwork behaviors are filtered through gendered lenses; for mothers this often means harsher or more cautionary feedback and lower “potential” assessments (Correll et al., 2007; Bono et al., 2017; Bloomfield et al., 2021; He et al., 2024). Together, these pathways indicate that motherhood penalties arise not only from explicit prejudice but also from the interaction between organizational time regimes, ambiguous criteria and information-poor appraisal settings. Notably, the empirical base on pre-/post-maternity changes in ratings remains limited (few longitudinal designs), and several studies rely on vignettes or single-organization samples; therefore, while effect directions are consistent, causal strength and generalizability remain limited. This gap suggests the need for designs that track individual evaluations before and after maternity-related events and that isolate appraisal-system features (criteria, scales, feedback sources) as moderators.
These motherhood penalties thus go beyond general gender bias: while gender stereotypes affect all women, the mechanisms identified here—time norm violations, flexibility stigma and stereotype-consistent interpretations—create distinctive disadvantages that mothers face specifically because of their caregiving role.
6. Conclusions
Scholars have long examined the relationship between gender gaps and performance management across disciplines and methods. Although progress has been made, relatively few studies directly link performance evaluations—often framed as gender-neutral—to the (re)production of inequalities. Despite experimental contributions (e.g. Bloomfield et al., 2021; Snir, 2019), robust field-based evidence remains scarce and only a limited number of works rely on real-world organizational data (e.g. Chauhan et al., 2022; Castilla, 2015).
Geographical coverage is uneven: the United States, parts of Europe, Asia and Africa are represented (Grund and Soboll, 2024; Festing et al., 2015; Crowley, 2013; Chauhan et al., 2022; Inesi and Cable, 2015), but studies on the public sector and on female-dominated, lower-income occupations are largely absent. Expanding research into these contexts could yield crucial insights, as economic precarity and job insecurity may amplify disparities.
While this review did not adopt an intersectional lens, it is notable that no studies examined the combined effects of gender with axes such as age or disability in performance evaluations. Within this binary focus, the findings underscore the salience of motherhood-specific dynamics: mothers are consistently penalized through time norms, stigmatization of flexible work or policy use, and stereotype-driven feedback. These mechanisms confirm that motherhood penalties are not merely a subset of gender bias but a distinct structural pathway of disadvantage.
Taken together, the review shows that no consensus has yet emerged on how to effectively redress gender disparities in evaluation systems.
In line with our broadened research question, the review has highlighted that gendered mechanisms in performance evaluations manifest across systemic, organizational, methodological, interactional and care-related dimensions. Evidence suggests that motherhood constitutes a particularly penalizing axis of inequality, though the empirical base remains limited. More systematic research is needed, especially longitudinal designs comparing pre- and post-maternity evaluations or testing how specific appraisal features amplify or mitigate penalties (Collins, 1998).
Beyond mapping existing evidence, this review identifies implications for both research and practice. For researchers, the analysis highlights the need for more longitudinal and field-based studies that track individual evaluations over time. Future work should also address underexplored contexts—such as the public sector and female-dominated or lower-income occupations—where performance systems may operate differently and where the motherhood penalty may be compounded by economic precarity. Moreover, there is considerable scope for adopting intersectional perspectives that consider how gender interacts with age, disability, or migration status, dimensions that remain virtually absent in the current literature. Practitioners should prioritize transparency in criteria, conduct regular gender audits of evaluation outcomes, and invest in training evaluators to recognize and counteract bias, bearing in mind that performance evaluation systems are not neutral instruments but can inadvertently reproduce inequalities. At the organizational level, aligning performance standards with task-related outcomes rather than time-based availability and designing family-friendly policies that do not stigmatize users are practical steps to reduce disparities. Together, these implications underscore the potential for more equitable evaluation systems and more robust empirical research to inform them.
The supplementary material for this article can be found online.





