This study examines whether an evaluee's prior performance anchors professional evaluators' assessments of current performance and whether situational conditions alter this performance evaluation bias (PEB). Integrating theory on heuristic and systematic models of information processing, it tests whether evaluator fatigue strengthens anchoring-based PEB and whether different forms of decision stakes weaken it.
The study analyzes 1,452,180 ball-and-strike calls made by Major League Baseball umpires during 2015–2018. Computerized pitch-location data provide an objective benchmark for each subjective call. Generalized linear models test whether pitchers' prior-season Wins Above Replacement predicts gifted strikes on objectively out-of-zone pitches and whether inning, run expectancy, strike count and batter All-Star appearances moderate this association.
Pitchers with stronger prior-season performance receive more gifted strikes, consistent with anchoring-based PEB. The moderating effects are neither uniform nor always as predicted. PEB decreases rather than increases in later innings, increases with run expectancy, decreases when the batter has one or two strikes and does not vary significantly with batter All-Star status. Situational conditions can therefore attenuate or intensify evaluators' reliance on prior-performance anchors.
The study positions anchoring as a central cognitive mechanism underlying PEB and extends prior research on status and reputation by demonstrating prior-performance effects among highly trained professional evaluators whose judgments can be compared with an objective benchmark. It also advances a context-sensitive account of PEB by showing that fatigue and decision stakes do not have uniform effects and that distinct forms of stakes can have opposing associations with evaluative bias.
Employees expect performance evaluations to foster growth and offer constructive feedback, yet in practice, these evaluations may be subjective, biased and inconsistent across managers. Performance evaluation bias (PEB) occurs when information about evaluees that is separate from their current performance (e.g. traits, demographic characteristics or reputational signals) has a disproportionate influence on current assessments (Levy and Williams, 2004; Prendergast and Topel, 1993) with consequences for motivation, perceived fairness and career outcomes (Greenberg, 1986; Kuvaas, 2006).
A central mechanism through which such bias may arise is anchoring. Anchoring bias describes the tendency for decision-makers to place disproportionate weight on an initial reference point when forming subsequent judgments and to adjust insufficiently as new information becomes available (Teovanović, 2019; Thorsteinson et al., 2008; Tversky and Kahneman, 1974). In performance evaluation settings, the evaluee's past performance may serve as such an anchor by establishing historically-based impressions of their capabilities that shape how evaluators interpret current performance (Berger and Daumann, 2021). Consequently, individuals with strong performance records may receive more favorable assessments than objectively warranted, whereas those with weaker histories may be judged more harshly despite equivalent current performance. This anchoring-based PEB may be especially pronounced when evaluations must be made rapidly and repeatedly under uncertainty, time pressure, and cognitive fatigue. Because adjustment away from an initial bias requires attention and cognitive effort, these conditions may constrain evaluators' ability to correct for prior impressions and increase reliance on prior performance cues (Epley and Gilovich, 2006; Gilbert et al., 1988; Maule et al., 2000; van der Linden et al., 2003).
Research also suggests that PEB cannot be understood solely as a function of the evaluee's individual-level attributes such as prior performance; the context of the evaluation is also relevant (Baron and Hershey, 1988; Lefgren et al., 2015). Since performance evaluation is a social-psychological process (Ferris et al., 1994; Murphy and Cleveland, 1991) involving evaluators and evaluees (Kozlowski et al., 1998) situated in a specific context (Levy and Williams, 2004), characteristics of that context affect how performance information is processed (Judge and Ferris, 1993; Kim and King, 2014; Levy and Williams, 2004). However, despite evidence that context matters, how situational factors strengthen or weaken anchoring-based PEB remains unclear. Clarifying this association is important for understanding when biased performance evaluations are most likely to occur.
In this paper, we draw on insights from theory on systematic and heuristic modes of information processing (Eagly and Chaiken, 1993; Petty et al., 1998) to examine how anchoring and situational factors influence PEB. We build from the proposition that an evaluee's prior performance may serve as a heuristic anchor that biases assessments of current performance. We then examine two situational conditions that may alter this bias. First, we consider whether cognitive fatigue intensifies PEB by making evaluators' current assessments more susceptible to the influence of prior performance. Second, we consider whether high-stakes situations reduce PEB by increasing evaluators' motivation to process current-performance information systematically. We therefore test whether PEB becomes stronger as evaluator fatigue increases and weaker as the stakes of the evaluation increase.
We explore these ideas empirically using a dataset of all “ball and strike calls” made by Major League Baseball (MLB) umpires between 2015 and 2018. Calling balls and strikes requires umpires to make repeated performance evaluations throughout a game. Every time a pitcher throws a pitch toward the batter, an umpire is responsible for evaluating whether the pitch, as it crosses home plate, is a “strike” (i.e. whether the pitcher has thrown the ball in the “strike zone” where the batter can reasonably be expected to hit it) or “ball” (i.e. outside the strike zone). In terms of game play, strikes favor the pitcher, while balls favor the batter. Umpires make this evaluation of the pitcher's performance in milliseconds, providing a valuable context for studying repeated instances of swift performance evaluation under high-pressure conditions. Moreover, variation across innings and game situations allows us to examine whether evaluator fatigue and the stakes of the evaluation moderate PEB.
Our study makes three primary contributions. First, we add to the literature on PEB by providing evidence that even highly trained professional evaluators exhibit anchoring-based PEB: pitchers with stronger past performance records receive more “gifted strikes” (i.e. the umpire calls a pitch a strike when it is actually a ball) than their peers with lower prior performance (Kim and King, 2014; Waguespack and Salomon, 2016; Zuckerman, 1977). Second, we show that contextual factors do moderate this PEB. However, our findings suggest that this moderation is more complex than initially predicted. We find less PEB when the umpire's call becomes more consequential and, contrary to expectations, later in the game. However, we find more PEB when run expectancy is higher. Third, our results contribute to broader debates about when situational demands result in evaluators using more careful judgment and reducing bias, and when they instead constrain information processing and increase reliance on heuristics (Lerner and Tetlock, 1999; Maule et al., 2000). In this way, the paper speaks to both organizational scholars interested in performance appraisal and sports scholars examining officiating and decision-making in high-stakes competitive environments. Across these domains, our findings suggest that anchoring on historical performance is not merely present in routine performance evaluations but a context-sensitive bias that may either diminish or intensify as the conditions surrounding the evaluation change.
Theory and hypotheses
Individuals may process information and form judgments through systematic or heuristic modes of processing (Eagly and Chaiken, 1993; Petty et al., 1998). Systematic processing is comparatively deliberate and effortful, involving careful attention to relevant information, consideration of alternative interpretations and explicit weighing of evidence before reaching a judgment (Calabretta et al., 2017; Kinicki et al., 2020). When valid current information is available and carefully considered, this mode may produce evaluations that are more responsive to current performance. However, it requires cognitive resources and time that may be scarce in applied settings.
In contrast, heuristic processing relies on simplifying judgment rules and previously acquired schemas that enable decision-makers to reduce effort by substituting readily available information for more extensive analysis (Eagly and Chaiken, 1993; Tversky and Kahneman, 1974). Because it is faster and less cognitively demanding, heuristic processing may occur more frequently when evaluators are making their assessments under time pressure, high workload or other constraints on processing capacity (Payne et al., 1993). Heuristic processing is not inherently inaccurate and can produce efficient and effective judgments. In performance evaluation, however, it can increase susceptibility to bias when information that is imperfectly related – or unrelated – to current performance influences evaluators' judgments. Under these conditions, evaluations may systematically deviate from judgments based exclusively on current performance.
Anchoring and performance evaluation bias
Performance evaluations often occur under conditions of uncertainty and constrained attention, making evaluators susceptible to prior impressions, social categories, reputations and status. Several theoretically distinct literatures document such influences, including research on the halo effect, the Matthew effect, expectation states and status-based evaluation (Cooper, 1981; Correll and Ridgeway, 2003; Merton, 1968; Podolny, 1993; Ridgeway and Correll, 2006). Although these perspectives invoke different causal processes, they share the proposition that information established before or outside the focal performance can shape how evaluators interpret current performance. Consistent with this proposition, research demonstrates that race, ethnicity and status can influence evaluations even when relatively objective performance information is available (Green et al., 2007; Kim and King, 2014; Parsons et al., 2011).
Anchoring effects occur when an initial value or judgment serves as a reference point and exerts a disproportionate influence on a subsequent assessment (Chapman and Johnson, 2002; Tversky and Kahneman, 1974). Some anchoring effects arise through insufficient adjustment away from that reference point, although the mechanisms may vary with the nature and source of the anchor (Epley and Gilovich, 2006; Teovanović, 2019). We propose that in performance evaluations, the evaluee's prior performance may serve as an informational anchor, establishing an initial expectation against which evaluators interpret current behavior. Information about prior performance does not necessarily bias current assessments because it may contain valid information about the evaluee's abilities. PEB arises, however, when the evaluee's prior performance anchors the evaluator's initial judgment and the evaluator adjusts insufficiently when assessing current-period performance (Epley and Gilovich, 2006; Tversky and Kahneman, 1974).
We extend these arguments by identifying anchoring as a central cognitive mechanism underlying PEB. Viewed through this lens, established impressions, reputations and status expectations may exhibit anchoring-like effects when they continue to influence judgments after new evidence arrives. We do not argue that halo effects, Matthew effects or expectation states are reducible to anchoring. Rather, anchoring offers a common mechanism through which the prior information emphasized by these distinct traditions may contribute to bias in performance evaluations.
Empirical research across several domains demonstrates that prior success, reputation and status can exert persistent effects on subsequent evaluations, although this research has generally not tested anchoring as the specific underlying mechanism. In science, Zuckerman (1977) documented processes of stratification and cumulative advantage among elite scientists, illustrating how recognition and resources become concentrated among already prominent researchers. In the Olympic Games, Waguespack and Salomon (2016) found that past country-level performance predicted current outcomes more strongly in subjectively evaluated sports than in objectively measured sports, a pattern consistent with favorable treatment of reputationally privileged competitors. In the same empirical context used in the current study, Kim and King (2014) found that Major League Baseball umpires were more likely to judge ambiguous pitches favorably when they were thrown by high-status pitchers. Together, these studies demonstrate that prior reputations carry over into subsequent evaluations, particularly when performance is ambiguous or subjectively assessed. This pattern is consistent with our proposed mechanism whereby prior performance establishes a reference point that may bias subsequent assessments.
Drawing on these theoretical insights, we theorize that an evaluee's prior performance will bias professional evaluators' assessments of their current performance. Specifically, when two individuals deliver objectively equivalent current-period performance, we expect the one with stronger prior performance to be more likely to receive a favorable subjective evaluation.
Evaluators' assessments of current-period performance will be biased in the direction of the evaluee's prior performance, such that stronger prior performance will be associated with more favorable assessments than indicated by objective current-period performance measures.
Situational influences on PEB
Although we propose anchoring as a central mechanism underlying PEB, evaluative judgments do not occur in a vacuum. Situational conditions may shape whether evaluators process information heuristically or engage in more effortful, systematic processing (Hastie and Dawes, 2009; Payne et al., 1993). We therefore consider situational factors that may strengthen or weaken the influence of prior performance on current assessments.
We first consider evaluator fatigue. The influence of prior performance should depend in part on the cognitive resources available to the evaluator. While heuristic processing allows evaluators to conserve resources by using readily available information (Eagly and Chaiken, 1993; Payne et al., 1993), systematic processing requires cognitive effort. Therefore, as evaluators become fatigued, it may be harder to make the cognitive effort needed to process performance information systematically. Applying this logic to our current theorizing and conceptualizing an initial judgment based on prior performance as an anchor from which the evaluator must adjust when assessing current performance, bias arises when that adjustment is insufficient (Epley and Gilovich, 2006; Tversky and Kahneman, 1974). Since adjustment is an effortful, controlled process that depends on available cognitive resources, fatigue should limit the evaluator's ability to make that adjustment (Gilbert, 2002; Gilbert et al., 1988). As a result, assessments should become more strongly biased by prior performance as evaluator fatigue increases.
The influence of the evaluee's prior performance on assessments of current-period performance will become stronger as evaluator fatigue increases.
We next consider whether high-stakes situations reduce PEB. Unlike cognitive fatigue, which limits evaluators' capacity for effortful processing, higher stakes may increase their motivation to evaluate current performance carefully. Decision-makers adapt their information-processing strategies to situational demands (Hastie and Dawes, 2009; Payne et al., 1993), and theories of heuristic and systematic processing propose that greater motivation can increase the likelihood of effortful examination of relevant information (Chaiken et al., 1989; Eagly and Chaiken, 1993). Because judgment errors have more consequential outcomes in high-stakes situations, evaluators should be more motivated to attend closely to diagnostic information about current performance rather than rely on readily available prior-performance cues.
With respect to anchoring-based PEB, this increased motivation should encourage evaluators to adjust more fully away from the initial judgment suggested by an evaluee's prior performance. Adjustment from an anchor requires controlled cognitive effort (Epley and Gilovich, 2006; Gilbert et al., 1988; Tversky and Kahneman, 1974). Thus, as the consequences of an evaluation increase, evaluators should process current-performance information more systematically, reducing the influence of prior-performance anchors and, consequently, anchoring-based PEB.
The influence of the evaluee's prior performance on assessments of current-period performance will become weaker as the stakes of the evaluation increase.
Methodology
Research context
We explore these ideas using a dataset of umpiring data from the 1,452,180 umpire calls made in Major League Baseball (MLB) during the 2015–2018 regular seasons. Researchers have used MLB data to address a range of questions in the management literature at both the macro- and micro-levels (see Day et al., 2012; Fonti et al., 2023 for reviews). Researchers have drawn on baseball's rich empirical setting to study personnel, OB and HR topics, such as compensation and mobility (Brymer et al., 2024; Choudhury et al., 2026; Hill et al., 2017; Manroop et al., 2026; Moliterno and Wiersema, 2007; Werner and Mero, 1999), equity and expectancy theory (Harder, 1991; Lord and Hohenfeld, 1979; Singh and Ramdeo, 2023), managerial succession (Gamson and Scotch, 1964), assessments of employee ability and promotion (Black and Vance, 2021), leadership style (Sverdlik et al., 2022) and organizational citizenship behavior (Love and Kim, 2019).
While much of the research using data from MLB has focused on players (e.g. Crocker and Eckardt, 2014) or competitive outcomes (e.g. Humphrey et al., 2009), our study focuses on performance evaluations made by MLB umpires (Flannagan et al., 2024; Kim and King, 2014; Terry et al., 2025). The role of umpires and officials in sports, in general, is crucial as they serve as arbiters of fair play, making split-second decisions that can significantly impact the outcome of a game. In baseball, umpires are involved in officiating every play over the course of the game. That is, baseball umpires do more than identify rule infractions; they make calls that determine how each play is adjudicated. As a result, the accuracy of their evaluations – referred to as “calls” – determines whether a game is perceived as objective, fair and legitimate. As a result, MLB umpires face a challenging task: they must continually make split-second decisions that can significantly impact game outcomes, often under intense scrutiny from players, coaches and thousands of fans. Prior management scholarship has used this high-pressure environment in studies of decision-making (Kim and King, 2014; Terry et al., 2025), and we leverage this context to examine professional evaluators' ability to remain objective and unbiased when making performance assessments.
Central to the umpire's role is the evaluation of “pitches.” Each plate appearance (i.e. “at bat”) involves a player on one team (the “pitcher”) throwing the ball toward a player on the opposing team (the “batter”) who is standing 60 feet and 6 inches (18.39 m) away at “home plate.” This pitch provides the opportunity for the batter to “put the ball in play,” which is necessary for their team to score runs. The home plate umpire, standing behind the catcher who receives the pitch, determines whether a pitch that the batter does not swing at is a “ball” or a “called strike.” A pitch is called a strike if any part of the ball passes through the “strike zone” (Figure 1) as it crosses home plate; otherwise, it is called a ball. The upper boundary of the strike zone is the midpoint between the top of the batter's shoulders and the top of the uniform pants, and the lower boundary is the hollow beneath the batter's kneecap. Horizontally, the strike zone extends across the width of home plate (17 inches/43.18 cm) [1]. Strikes favor the pitcher: if the batter does not put the ball in play before three strikes are called, they are “called out” and their at-bat ends. Pitches that are called balls favor the batter, who is allowed to progress to “first base” after four balls [2]. In short, umpires evaluate pitchers' performance every time they throw the ball, and their evaluation of that performance – whether a taken pitch is a ball or a strike – is central to the fair and unbiased outcome of the game.
The scene is set in a baseball stadium. A batter is positioned at home plate, ready for the pitch from the pitcher. Behind the batter, a catcher in protective gear is crouched, holding a catcher's mitt as a target for the pitcher. An umpire stands behind the catcher prepared to call the pitch. A rectangular box in front of the catcher indicates the location and size of the strike zone.Strike zone in MLB*. Note(s): *Image generated using ChatGPT
The scene is set in a baseball stadium. A batter is positioned at home plate, ready for the pitch from the pitcher. Behind the batter, a catcher in protective gear is crouched, holding a catcher's mitt as a target for the pitcher. An umpire stands behind the catcher prepared to call the pitch. A rectangular box in front of the catcher indicates the location and size of the strike zone.Strike zone in MLB*. Note(s): *Image generated using ChatGPT
It is important to note that evaluating pitches as balls or strikes is an exceptionally difficult and cognitively demanding task. In a typical Major League Baseball (MLB) game, home plate umpires observe 250–300 pitches per game, although only pitches at which the batter does not swing require the ball-or-strike assessment examined in our study. MLB pitchers throw the ball, on average, 92–93 miles per hour (148–150 km/h) and routinely exceed 100 mph (161 km/h). At these speeds, the ball covers the distance between the pitcher and the batter in under 400 milliseconds, and umpires are expected to call the pitch a ball or strike instantly when the catcher receives it. In addition to the challenges associated with the speed at which the ball travels and the time it takes to reach the home plate, all pitchers have unique ways of throwing the ball, adding a spin that changes the ball's trajectory as it travels toward the batter, sometimes changing direction just as it reaches the strike zone. In addition, the physical demands of crouching behind home plate and wearing heavy protective equipment in hot summer weather, together with the need to sustain concentration during games that may last over three hours, contribute to physical and cognitive fatigue that may impair umpires' decision-making. Finally, because MLB games are generally played before large crowds, umpires face considerable scrutiny of every call, with spectators loudly expressing their dissatisfaction with calls they consider unfair or inaccurate. Not surprisingly, MLB umpires require years of experience to develop their professional competencies, and before working in MLB, umpires hone their skills by officiating games in the “minor leagues” (i.e. the professional baseball leagues where MLB teams develop future players; MLB.com, 2024).
Importantly, umpires do not evaluate each pitch without prior information about the pitcher. Over the course of a season – and particularly over their careers – MLB umpires encounter a broad range of pitchers, including repeated encounters with many of the same players. These prior interactions, together with pitchers' established and well-known performance records and reputations, provide ample opportunity for umpires to form expectations about pitcher quality before evaluating a particular pitch. Thus, MLB umpiring combines professional evaluators who have well-developed beliefs about evaluees with a setting in which current performance can be measured independently. As described below, this allows us to examine whether a pitcher's prior performance systematically biases an umpire's evaluation of objectively measured current performance.
In sum, the nature of umpires' work, together with the decision-making environment in which they operate, provides a unique and empirically rich context for studying performance-anchored bias among professional performance evaluators. While they generally maintain high levels of accuracy, umpires also make mistakes. The repeated incidence and time pressure of their judgments present a challenge for systematic processing and increase the likelihood of biases that attend heuristic shortcuts in cognitive pathways. Furthermore, the public nature of MLB games, with players, coaches and fans closely observing umpires' calls, allows examination of the potential watchdog effect.
Data
We conducted our analysis using a dataset constructed from multiple sources. Data on individual umpire ball-and-strike calls for 1,452,180 pitches during 9,671 games in our sampling frame were collected from MLB Gameday. These data included the date, venue, participating teams, final score, inning-by-inning scoring and game-specific batting and pitching outcomes such as hits, runs, home runs and strikeouts. The MLB Gameday records incorporated estimates produced by MLB's pitch-tracking systems of each pitch's trajectory and location as it crossed home plate. During our 2015–2018 sampling period, MLB transitioned from the camera-based PITCHf/x system to the radar-based TrackMan component of Statcast. Both systems tracked the trajectory of each pitch as it traveled toward and crossed home plate and estimated its horizontal and vertical location as it crossed the plate. We refer to the information produced by these systems collectively as pitch-tracking data. These data allow us to compare the location of each pitch with the estimated boundaries of the batter's strike zone and thereby construct an objective benchmark against which to evaluate the umpire's subjective call. Importantly, the pitch-tracking systems were not used to make, challenge or overturn ball-and-strike calls during our sampling period [3]. Although graphics based on pitch-tracking data were routinely displayed during television broadcasts, the umpire necessarily made each call without access to the pitch-location comparison that we subsequently use to evaluate its accuracy.
To create the dataset used for our analysis, pitch-level data were combined with additional data from FanGraphs and the Lahman Baseball Database. We downloaded player WAR data (our main independent variable, see below) from FanGraphs. The Lahman Baseball Database provided historical player- and team-level data. We used this database to gather both the player and team statistics. At the player level, the database allowed us to collect individual batting, pitching and fielding performance metrics across different seasons, as well as player-level demographic data. At the team level, these data included seasonal performance statistics such as win-loss records.
We supplemented the data obtained from the Lahman Baseball Database with data from Spotrac.com, a comprehensive sports contract and salary database. Specifically, we used Spotrac to identify each pitcher's primary role, distinguishing between starting pitchers and relief pitchers (i.e. pitchers who come in as “substitutes” for the starting pitcher). This distinction is crucial for our analysis because starting and relief pitchers have different roles, performance expectations and usage patterns, which can significantly impact their statistics and how they are evaluated. Starting pitchers, who begin the game, typically pitch for longer durations (often 5–7 innings or more) and have a more varied pitch repertoire, allowing them to face batters multiple times during the game.
Starting pitchers usually pitch every four to five days in a rotation and work to develop the endurance needed to perform consistently when throwing up to 100 pitches [4]. In contrast, relief pitchers enter the game after the starting pitcher, generally pitch for shorter durations (often 1–2 innings or less), may specialize in specific situations (e.g. closing games, facing same-handed batters), can pitch more frequently (sometimes on consecutive days) and often throw at maximum effort for short bursts. By incorporating this detailed role information from Spotrac, we were able to conduct a more nuanced analysis of pitcher performance and its implications within the context of our study.
In our analysis, we focused solely on the first nine “innings” of each game. A regulation MLB game is scheduled for nine innings, although the bottom of the ninth inning is not played when the home team is already leading. Games tied after nine innings have “extra innings,” and play continues until one team leads after a completed inning or the home team takes the lead during the bottom half of an inning. By restricting our analysis to the nine innings of a standard game, we aimed to minimize potential confounding factors that could introduce bias in decision-making and performance evaluation bias (PEB). For example, if the game extends beyond the standard nine innings, fatigue accumulates, impacting the cognitive and physical demands placed on umpires. Additionally, the strategies employed by teams during extra innings often differ from those in the standard nine innings, introducing new factors that could influence umpires' decision-making and further complicating the analysis of PEB.
Variables: dependent variable
We operationalize our performance-anchored PEB by measuring “gifted strikes” (Kim and King, 2014), or incidents of umpire “overrecognition” of a pitch. Overrecognition occurs when the home plate umpire calls a given pitch a strike and when PITCH/fx determines it was a ball (i.e. a Type I error). We focus on overrecognition because it directly captures the form of favorable evaluative bias central to our empirical test: the attribution of successful current performance when the objective benchmark does not support that assessment. The complementary error – calling an objectively true strike a ball – captures underrecognition and can occur only within a different, mutually exclusive set of pitches. Examining underrecognition therefore requires a separate dependent variable and opportunity sample. Our purpose is not to explain both directions of umpire error, but to test whether prior performance predicts favorable evaluations unsupported by objective current performance. We accordingly use overrecognition as the focal outcome. This choice does not imply that prior-performance bias can operate only in this direction.
We constructed the variable overrecognition by comparing the umpire's ball/strike call against PITCH/fx's assessment of the same pitch. Following Kim and King (2014), our baseline measure was “true balls,” as reported by the PITCH/fx data. Thus, overrecognition takes a value of 1 when the umpire called a pitch a strike when PITCH/fx determined that it was a ball. Constructed in this manner, overrecognition has values of 0 when an umpire's subjective assessment of a given pitch is accurate and 1 when the umpire calls a pitch that is objectively a ball a strike: that is, when they “gift” the pitcher a strike. Of course, umpires could – and might reasonably be expected to – make mistakes; this would result, on average, in umpires calling gifted strikes equally for all pitchers. However, umpires may systematically give more or fewer gifted strikes to particular pitchers, providing evidence of bias. Accordingly, we test whether nonzero values for this variable are associated with an individual pitcher's past performance.
Variables: independent variable
Our key independent variable, pitcher's past performance, is a pitcher's prior-season “Wins Above Replacement” (WAR), a composite metric that quantifies the total value a pitcher contributed relative to a hypothetical replacement-level player over a given season. Specifically, we use each pitcher's lagged WAR from the preceding season – obtained from FanGraphs and merged to our pitch-level data via the Lahman baseball database player identifier – to capture a pitcher's historical performance that umpires may anchor upon when evaluating pitches. We choose WAR over alternative performance metrics for three interconnected reasons: its comprehensiveness, its adjustment for exogenous factors and its demonstrated superiority as a measure of pitcher skill in the sabermetric and sports analytics literatures [5].
First, WAR is the most comprehensive single-number summary of pitcher value available. Unlike counting statistics such as wins, strikeouts or saves – which capture only one dimension of performance and are heavily influenced by team context – WAR integrates a pitcher's contributions across all dimensions of run prevention into a single, additive quantity expressed in units of wins relative to a replacement-level player (Baumer and Zimbalist, 2014; Tango et al., 2007). The replacement-level baseline is itself a theoretically grounded construct representing the freely available, league-minimum talent pool from which teams can draw at any time, making WAR a genuinely value-added measure rather than an absolute count (Slowinski, 2010). For our purposes, this comprehensiveness is important because umpires are not evaluating isolated components of a pitcher's performance record; they are evaluating a pitcher holistically, and WAR approximates the overall performance signal conveyed by a pitcher's reputation.
Second, WAR is designed to isolate pitcher skill from exogenous confounds that would otherwise contaminate a pure measure of prior-season performance. FanGraphs' pitcher WAR (fWAR) estimates a pitcher's value above replacement using a modified version of Fielding Independent Pitching (FIP), which emphasizes outcomes over which pitchers exert relatively direct control – strikeouts, walks, hit batters, home runs and infield fly balls – rather than beginning with runs allowed (Slowinski, 2012). FanGraphs adjusts this measure for park effects and differences in scoring across leagues and seasons, compares it with the relevant league average and accounts for innings pitched when converting performance into wins above replacement. Because WAR is designed to be less sensitive than runs-allowed measures to team defense and the sequencing of balls in play, it provides a useful measure of a pitcher's demonstrated performance for our study. This is particularly important in a study of evaluator (i.e. umpire) cognition: if our measure of prior performance substantially reflected favorable defensive or park circumstances, any observed bias could be attributable to those circumstances rather than to evaluators anchoring on the pitcher's established performance record. The use of FIP and park adjustments in the calculation of WAR reduces this concern. By contrast, ERA – the most cited traditional pitching performance measure – is heavily influenced by team defense, the sequencing of events, overall scoring conditions and park dimensions, making it a comparatively noisy measure of pitcher performance (Bradbury, 2007; Hakes and Sauer, 2006; McCracken, 2001).
Third, pitcher WAR has been empirically validated as a strong predictor of future pitcher performance relative to traditional alternatives. McCracken's (2001) foundational work on defense-independent pitching demonstrated that pitchers have limited control over the outcomes of balls put in play, implying that metrics relying on hits allowed (e.g. earned run average; “ERA”) overweight performance outcomes not directly controllable by the focal player. Subsequent work similarly argues that fielding-independent and context-neutral metrics are more valid measures of pitcher performance than commonly used measures such as ERA or won-loss records (Albert, 2006; Bradbury, 2007; Jensen et al., 2009). Indeed, the won-loss record has repeatedly been shown to be among the least reliable indicators of pitcher performance, as it depends critically on run support from the pitcher's own offense and on the quality of the bullpen (Baumer and Zimbalist, 2014; Tango et al., 2007). The validity of WAR as a predictive instrument is precisely why it has become the standard evaluation tool in front offices, salary arbitration proceedings and academic studies employing baseball performance data (Hakes and Sauer, 2006; Kim and King, 2014).
The measure pitcher's past performance is each pitcher's lagged WAR from the season immediately preceding the pitch observation (i.e. WAR in season t−1 for pitches thrown in season t). Lagging the performance measure serves two purposes. First, it reflects the temporal structure of anchoring: umpires' expectations about a pitcher's quality are formed based on accumulated prior evidence, not contemporaneous performance within the same season, which is unobserved at the start of each new season. Second, lagging eliminates simultaneity bias that would otherwise arise if a pitcher's current-season performance drove both the WAR measure and overrecognition within the same season. Pitchers without a prior-season FanGraphs WAR value were assigned a pitcher's past performance equal to 0 and flagged with the dummy variable rookie set to equal 1.
Variables: situational moderators
In our theoretical development, we propose two situational factors that moderate the association between a pitcher's prior performance and the umpire's PEB: evaluator cognitive fatigue (H2) and high-stakes evaluations (H3).
Several features of MLB umpiring suggest that cognitive fatigue should become a factor that impacts performance-anchored PEB as a game progresses. Home plate umpires can make hundreds of ball/strike calls per game, each within milliseconds, while crouching in heavy protective equipment under physically demanding conditions. Over the course of a game, umpires therefore experience both physical fatigue, which degrades perceptual acuity and motor responsiveness, and cognitive fatigue, which arises from sustained discriminative judgment under time pressure. Both should attenuate the umpire's ability to engage in systematic information processing and avoid or correct for heuristic anchoring on the pitcher's prior performance. Accordingly, we operationalize fatigue with the variable inning, measured as the current inning of the game in which the focal observation occurs.
We used multiple variables to examine how the stakes of the situation affect an umpire's PEB. The first is run expectancy, which measures the average number of runs expected to be scored in an inning, given the current base runner-and-out configuration (Kim and King, 2014). For instance, when there are runners on all three bases (first, second and third) and no outs have yet been recorded in the inning, this situation is referred to as “bases loaded, no outs” and typically has high run expectancy, indicating a high likelihood of runs being scored during that batter's at bat. By contrast, a situation with a runner on base and two outs has a much lower run expectancy. In this way, run expectancy measures the stakes of the umpire's current call relative to the batting team's run scoring potential.
Second, we measure the number of strikes against the batter at the time of the focal pitch. Within an at-bat, strikes become increasingly valuable to the pitcher as the count progresses, particularly when there are two strikes: with no strikes, each pitch was relatively less consequential. However, on a 2-strike count, the next pitch could potentially end the hitter's at-bat; therefore, the stakes are higher for each call with a 2-strike count. This increases the stakes of the umpire's call, because the next call may determine whether the at-bat continues or ends. To account for the escalating stakes of the strike count, we created three categorical variables based on the number of strikes. The variable no strikes takes values of 0 when the batter has no strikes against him; the variable one strike takes values of 1 when the batter has one strike and 0 otherwise; two strikes takes values of 1 when the batter has two strikes, and 0 otherwise. In the regression models, no strikes served as the reference category.
Finally, we explored the dyadic situational context of a high-performing pitcher facing a high-status batter. The batter's status was assessed by counting the number of times they had been selected for the All-Star Game (Kim and King, 2014) [6]. Stakes are higher, and spectator monitoring greater, in high-profile pitcher/batter matchups, where a renowned pitcher faces a prominent batter: increased media and fan attention on such matchups amplifies the visibility of every call, placing greater pressure on umpires to make accurate decisions to avoid criticism. Additionally, the stakes of the game are higher when the top players are involved, making each call more impactful on the game's outcome. We follow prior research (Kim and King, 2014) by operationalizing this status effect with the variable batter: all-star, measured as the count of the number of times the current batter has appeared in an MLB All-Star game.
Variables: controls
To account for observable player-level characteristics that may influence umpire judgment, we included several controls for both pitchers and batters. Race was included for both pitchers and batters to assess potential racial bias in umpire decisions. Following prior research suggesting that umpires may favor White players in predominantly White sports (Birnbaum, 2010; Parsons et al., 2011), we included binary variables pitcher(batter): White coded as 1 if the player is White/Caucasian and 0 otherwise. We also control for age-related declines in motor skills, which could shape umpires' perceptions of player ability and thus influence borderline calls. The variable pitcher(batter) age measures the pitcher's(batter's) age in years. Handedness – that is, the pitcher's throwing hand and the batter's standing side – was included because pitch trajectory and batting angle vary systematically depending on handedness. These perceptual differences may influence how umpires visually process the location of the pitch. Both variables were coded as binary indicators: pitcher(batter) (right-handed) takes values of 1 for right-handed pitchers and right-standing batters, respectively.
We included three additional pitcher-specific controls. To measure the pitcher's status (and for symmetry with the batter status effect we model as a moderation), we included pitcher: all-star, measured by the number of times the pitcher had appeared in an MLB All-Star game (Kim and King, 2014). Second, we control for whether a pitch is located at the corner of the strike zone in a “vague” boundary region. We operationalize the variable edge pitch, which takes values of 1 when the pitch is within approximately 1 cm of both the horizontal and vertical edges of the batting zone and zero otherwise. Since these corner pitches are especially difficult to call, including this dummy variable helps separate location-driven ambiguity in ball/strike decisions from the effects of pitcher prior performance and other situational factors. Finally, we control for whether the pitcher is on the home team because partisan crowd pressure may make umpires more likely to issue favorable calls to the home pitcher. The variable home team equals 1 when the pitcher is playing for the home team and 0 otherwise.
Since we are theoretically interested in how the situational dynamics may impact the umpire's PEB during a given at-bat, we controlled for several game-level effects. Since the score margin in the game may influence umpire bias, we operationalized a measure of the run differential between the pitching and batting teams. The dummy variable small score difference was calculated for each at-bat and coded with values of 1 when there were five or fewer runs separating teams and 0 otherwise. Since a larger number of spectators may increase the stakes and attendant monitoring of the umpire's calls, we operationalized the variable attendance, measured as the number of spectators in the stadium for the game. The variable bases loaded captures a high-stakes situation in a baseball game in which there is a runner on each of the three bases. This situation is significant because it presents an immediate scoring opportunity for the team at bat, as any hit or walk will result in at least one run. In such high-pressure circumstances, the outcome of an umpire's call becomes critical as it can significantly impact the game's momentum and outcome. Additionally, we included controls for two weather factors, temperature and wind speed, because weather could influence umpires' and players' fatigue and the accuracy of the performance evaluation.
Results
We estimated generalized linear models (GLMs) with a binomial family and logit link. All models include umpire-by-year fixed effects to absorb unobserved umpire tendencies within each season and season-specific baseline differences in strike calling. To allow for heteroscedasticity and within-game dependence, we clustered robust standard errors at the game level (Greene, 2003). Table 1 presents descriptive statistics and a correlation matrix for the data: quite a few variables were significantly correlated at the p = 0.05 level. To explore whether multicollinearity was a concern, we calculated variance inflation factors (VIFs) separately for each model, centering the variables for the interaction terms. Maximum VIFs ranged from 1.26 to 3.36 across the seven models, with a maximum of 3.36 in the full model containing all interactions. Following O'Brien (2007), we interpret these values in the context of the model specifications rather than applying a rigid threshold. The consistently low VIFs, including in the full model, indicate that multicollinearity is unlikely to affect our estimates materially.
Descriptive statistics and pairwise correlations*
| # | Variable | Mean | SD | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | 10 |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | Overrecognition | 0.09 | 0.29 | – | |||||||||
| 2 | Pitcher's past performance | 1.21 | 1.55 | 0.017 | – | ||||||||
| 3 | Run expectancy | 0.48 | 0.35 | −0.010 | −0.033 | – | |||||||
| 4 | One strike | 0.31 | 0.46 | 0.040 | 0.002 | −0.014 | – | ||||||
| 5 | Two strikes | 0.32 | 0.47 | 0.043 | 0.019 | −0.018 | −0.456 | – | |||||
| 6 | Batter All-Star | 0.38 | 1.19 | −0.012 | 0.005 | 0.006 | −0.002 | −0.010 | – | ||||
| 7 | Pitcher All-Star | 0.15 | 0.65 | 0.023 | 0.444 | −0.024 | −0.001 | 0.017 | −0.004 | – | |||
| 8 | Rookie | 0.14 | 0.35 | −0.008 | −0.322 | 0.009 | −0.000 | −0.009 | −0.005 | −0.089 | – | ||
| 9 | Pitcher (right-handed) | 0.73 | 0.45 | 0.001 | −0.057 | −0.002 | −0.002 | 0.003 | −0.002 | −0.095 | 0.024 | – | |
| 10 | Pitcher: White | 0.76 | 0.43 | −0.004 | 0.023 | 0.000 | 0.001 | 0.002 | −0.000 | 0.046 | −0.024 | −0.056 | – |
| 11 | Pitcher's age | 28.81 | 3.78 | 0.001 | 0.161 | −0.010 | 0.001 | 0.005 | −0.001 | 0.081 | −0.294 | 0.012 | 0.027 |
| 12 | Edge pitch | 0.00 | 0.04 | 0.004 | −0.000 | −0.001 | 0.001 | −0.004 | −0.001 | 0.000 | −0.000 | −0.001 | 0.000 |
| 13 | Batter: White | 0.55 | 0.50 | −0.003 | −0.002 | −0.001 | −0.001 | 0.002 | −0.011 | 0.001 | 0.001 | 0.004 | 0.001 |
| 14 | Batter age | 28.88 | 3.91 | −0.017 | 0.012 | 0.006 | 0.003 | −0.008 | 0.061 | 0.001 | −0.019 | −0.000 | −0.002 |
| 15 | Batter (right-handed) | 0.57 | 0.50 | 0.020 | 0.020 | 0.006 | 0.003 | 0.016 | 0.046 | 0.026 | −0.010 | −0.189 | 0.009 |
| 16 | Small score difference | 0.91 | 0.29 | 0.000 | 0.111 | −0.002 | −0.000 | −0.001 | 0.016 | 0.036 | −0.058 | 0.004 | 0.006 |
| 17 | Attendance | 29,770 | 10,221 | 0.002 | 0.069 | −0.003 | −0.000 | 0.003 | 0.047 | 0.055 | −0.028 | −0.031 | −0.011 |
| 18 | Inning | 4.89 | 2.54 | 0.024 | −0.225 | 0.001 | −0.010 | 0.001 | −0.033 | −0.014 | 0.011 | 0.030 | −0.027 |
| 19 | Bases loaded | 0.02 | 0.15 | 0.004 | −0.028 | 0.358 | −0.004 | −0.001 | −0.002 | −0.011 | 0.007 | −0.004 | −0.001 |
| 20 | Home team | 0.51 | 0.50 | 0.005 | −0.001 | −0.015 | −0.000 | 0.007 | −0.002 | 0.003 | −0.002 | 0.004 | 0.002 |
| 21 | Degrees | 73.59 | 10.66 | 0.004 | −0.036 | 0.000 | −0.002 | 0.000 | 0.001 | −0.004 | 0.070 | 0.002 | −0.012 |
| 22 | Windspeed | 7.53 | 5.06 | −0.006 | 0.009 | 0.010 | 0.000 | −0.002 | −0.009 | −0.009 | −0.013 | −0.006 | 0.010 |
| # | Variable | Mean | SD | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | 10 |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | Overrecognition | 0.09 | 0.29 | – | |||||||||
| 2 | Pitcher's past performance | 1.21 | 1.55 | 0.017 | – | ||||||||
| 3 | Run expectancy | 0.48 | 0.35 | −0.010 | −0.033 | – | |||||||
| 4 | One strike | 0.31 | 0.46 | 0.040 | 0.002 | −0.014 | – | ||||||
| 5 | Two strikes | 0.32 | 0.47 | 0.043 | 0.019 | −0.018 | −0.456 | – | |||||
| 6 | Batter All-Star | 0.38 | 1.19 | −0.012 | 0.005 | 0.006 | −0.002 | −0.010 | – | ||||
| 7 | Pitcher All-Star | 0.15 | 0.65 | 0.023 | 0.444 | −0.024 | −0.001 | 0.017 | −0.004 | – | |||
| 8 | Rookie | 0.14 | 0.35 | −0.008 | −0.322 | 0.009 | −0.000 | −0.009 | −0.005 | −0.089 | – | ||
| 9 | Pitcher (right-handed) | 0.73 | 0.45 | 0.001 | −0.057 | −0.002 | −0.002 | 0.003 | −0.002 | −0.095 | 0.024 | – | |
| 10 | Pitcher: White | 0.76 | 0.43 | −0.004 | 0.023 | 0.000 | 0.001 | 0.002 | −0.000 | 0.046 | −0.024 | −0.056 | – |
| 11 | Pitcher's age | 28.81 | 3.78 | 0.001 | 0.161 | −0.010 | 0.001 | 0.005 | −0.001 | 0.081 | −0.294 | 0.012 | 0.027 |
| 12 | Edge pitch | 0.00 | 0.04 | 0.004 | −0.000 | −0.001 | 0.001 | −0.004 | −0.001 | 0.000 | −0.000 | −0.001 | 0.000 |
| 13 | Batter: White | 0.55 | 0.50 | −0.003 | −0.002 | −0.001 | −0.001 | 0.002 | −0.011 | 0.001 | 0.001 | 0.004 | 0.001 |
| 14 | Batter age | 28.88 | 3.91 | −0.017 | 0.012 | 0.006 | 0.003 | −0.008 | 0.061 | 0.001 | −0.019 | −0.000 | −0.002 |
| 15 | Batter (right-handed) | 0.57 | 0.50 | 0.020 | 0.020 | 0.006 | 0.003 | 0.016 | 0.046 | 0.026 | −0.010 | −0.189 | 0.009 |
| 16 | Small score difference | 0.91 | 0.29 | 0.000 | 0.111 | −0.002 | −0.000 | −0.001 | 0.016 | 0.036 | −0.058 | 0.004 | 0.006 |
| 17 | Attendance | 29,770 | 10,221 | 0.002 | 0.069 | −0.003 | −0.000 | 0.003 | 0.047 | 0.055 | −0.028 | −0.031 | −0.011 |
| 18 | Inning | 4.89 | 2.54 | 0.024 | −0.225 | 0.001 | −0.010 | 0.001 | −0.033 | −0.014 | 0.011 | 0.030 | −0.027 |
| 19 | Bases loaded | 0.02 | 0.15 | 0.004 | −0.028 | 0.358 | −0.004 | −0.001 | −0.002 | −0.011 | 0.007 | −0.004 | −0.001 |
| 20 | Home team | 0.51 | 0.50 | 0.005 | −0.001 | −0.015 | −0.000 | 0.007 | −0.002 | 0.003 | −0.002 | 0.004 | 0.002 |
| 21 | Degrees | 73.59 | 10.66 | 0.004 | −0.036 | 0.000 | −0.002 | 0.000 | 0.001 | −0.004 | 0.070 | 0.002 | −0.012 |
| 22 | Windspeed | 7.53 | 5.06 | −0.006 | 0.009 | 0.010 | 0.000 | −0.002 | −0.009 | −0.009 | −0.013 | −0.006 | 0.010 |
| # | Variable | Mean | SD | 11 | 12 | 13 | 14 | 15 | 16 | 17 | 18 | 19 | 20 | 21 | 22 |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 11 | Pitcher age | 28.81 | 3.78 | – | |||||||||||
| 12 | Edge pitch | 0.00 | 0.04 | 0.000 | – | ||||||||||
| 13 | Batter is white/Caucasian | 0.55 | 0.50 | −0.003 | 0.001 | – | |||||||||
| 14 | Batter age | 28.88 | 3.91 | 0.013 | −0.001 | 0.019 | – | ||||||||
| 15 | Batter (right-handed) | 1.57 | 0.50 | 0.007 | −0.000 | −0.089 | −0.015 | – | |||||||
| 16 | Small score difference | 0.91 | 0.29 | 0.017 | 0.001 | −0.002 | 0.013 | −0.002 | – | ||||||
| 17 | Attendance | 29,770 | 10,221 | 0.057 | 0.000 | 0.024 | 0.026 | 0.003 | 0.005 | – | |||||
| 18 | Inning | 4.89 | 2.54 | 0.056 | −0.000 | 0.002 | −0.013 | −0.006 | −0.224 | 0.002 | – | ||||
| 19 | Bases loaded | 0.02 | 0.15 | −0.003 | −0.000 | 0.001 | −0.000 | 0.004 | −0.009 | −0.003 | 0.013 | – | |||
| 20 | Home team | 0.51 | 0.50 | 0.001 | 0.002 | −0.001 | 0.003 | −0.001 | 0.008 | −0.013 | 0.044 | −0.006 | – | ||
| 21 | Degrees | 73.59 | 10.66 | −0.037 | −0.001 | −0.007 | −0.031 | −0.005 | −0.014 | 0.053 | 0.002 | −0.003 | −0.000 | – | |
| 22 | Wind speed | 7.53 | 5.06 | 0.015 | 0.001 | −0.018 | 0.010 | −0.010 | −0.017 | 0.118 | −0.001 | 0.004 | −0.007 | −0.115 | – |
| # | Variable | Mean | SD | 11 | 12 | 13 | 14 | 15 | 16 | 17 | 18 | 19 | 20 | 21 | 22 |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 11 | Pitcher age | 28.81 | 3.78 | – | |||||||||||
| 12 | Edge pitch | 0.00 | 0.04 | 0.000 | – | ||||||||||
| 13 | Batter is white/Caucasian | 0.55 | 0.50 | −0.003 | 0.001 | – | |||||||||
| 14 | Batter age | 28.88 | 3.91 | 0.013 | −0.001 | 0.019 | – | ||||||||
| 15 | Batter (right-handed) | 1.57 | 0.50 | 0.007 | −0.000 | −0.089 | −0.015 | – | |||||||
| 16 | Small score difference | 0.91 | 0.29 | 0.017 | 0.001 | −0.002 | 0.013 | −0.002 | – | ||||||
| 17 | Attendance | 29,770 | 10,221 | 0.057 | 0.000 | 0.024 | 0.026 | 0.003 | 0.005 | – | |||||
| 18 | Inning | 4.89 | 2.54 | 0.056 | −0.000 | 0.002 | −0.013 | −0.006 | −0.224 | 0.002 | – | ||||
| 19 | Bases loaded | 0.02 | 0.15 | −0.003 | −0.000 | 0.001 | −0.000 | 0.004 | −0.009 | −0.003 | 0.013 | – | |||
| 20 | Home team | 0.51 | 0.50 | 0.001 | 0.002 | −0.001 | 0.003 | −0.001 | 0.008 | −0.013 | 0.044 | −0.006 | – | ||
| 21 | Degrees | 73.59 | 10.66 | −0.037 | −0.001 | −0.007 | −0.031 | −0.005 | −0.014 | 0.053 | 0.002 | −0.003 | −0.000 | – | |
| 22 | Wind speed | 7.53 | 5.06 | 0.015 | 0.001 | −0.018 | 0.010 | −0.010 | −0.017 | 0.118 | −0.001 | 0.004 | −0.007 | −0.115 | – |
Note(s): *All correlations with absolute value of 0.002 or greater are significant at the 0.05 level (two-tailed)
Table 2 reports the results of the GLM regressions predicting overrecognition. Model 1 presents the controls-only specification. With the exceptions of batter: White and attendance, the control variable coefficients are statistically significant at the p < 0.05 level. For example, umpires were more likely to gift strikes when the pitcher or batter was right-handed (b = 0.049, p < 0.001 and b = 0.140, p < 0.001, respectively), the pitcher had more All-Star appearances (b = 0.097, p < 0.001), when the score difference was small (b = 0.056, p < 0.001), the bases were loaded (b = 0.189, p < 0.001) or the pitcher was playing for the home team (b = 0.019, p = 0.003). Conversely, umpires were less likely to award gifted strikes to rookies (b = −0.073, p < 0.001) and to older pitchers and batters (b = −0.005, p < 0.001 and −0.014, p < 0.001, respectively). These coefficients identify baseline associations with overrecognition.
Generalized linear model results for performance evaluation bias*
| DV: Overrecognition | Model 1 | Model 2 | Model 3 | Model 4 | Model 5 | Model 6 | Model 7 | |||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| b(s.e.) | p-val. | b(s.e.) | p-val. | b(s.e.) | p-val. | b(s.e.) | p-val. | b(s.e.) | p-val. | b(s.e.) | p-val. | b(s.e.) | p-val. | |
| Pitcher's past performance | 0.034 | [0.000] | 0.052 | [0.000] | 0.028 | [0.000] | 0.042 | [0.000] | 0.034 | [0.000] | 0.053 | [0.000] | ||
| (0.003) | (0.005) | (0.004) | (0.004) | (0.003) | (0.006) | |||||||||
| Pitcher's past performance × Inning | −0.004 | [0.000] | −0.004 | [0.000] | ||||||||||
| (0.001) | (0.001) | |||||||||||||
| Pitcher's past performance × Run expectancy | 0.015 | [0.010] | 0.014 | [0.014] | ||||||||||
| (0.006) | (0.006) | |||||||||||||
| Pitcher's past performance × One strike | −0.012 | [0.009] | −0.012 | [0.009] | ||||||||||
| (0.005) | (0.005) | |||||||||||||
| Pitcher's past performance × Two strikes | −0.009 | [0.050] | −0.009 | [0.051] | ||||||||||
| (0.005) | (0.005) | |||||||||||||
| Pitcher's past performance × Batter: All-Star | 0.001 | [0.395] | 0.001 | [0.537] | ||||||||||
| (0.002) | (0.002) | |||||||||||||
| Inning | 0.034 | [0.000] | 0.039 | [0.000] | 0.044 | [0.000] | 0.039 | [0.000] | 0.039 | [0.000] | 0.039 | [0.000] | 0.044 | [0.000] |
| (0.001) | (0.001) | (0.002) | (0.001) | (0.001) | (0.001) | (0.002) | ||||||||
| Run expectancy | −0.097 | [0.000] | −0.095 | [0.000] | −0.095 | [0.000] | −0.112 | [0.000] | −0.095 | [0.000] | −0.095 | [0.000] | −0.112 | [0.000] |
| (0.009) | (0.009) | (0.009) | (0.012) | (0.009) | (0.009) | (0.012) | ||||||||
| One strike | 0.611 | [0.000] | 0.611 | [0.000] | 0.611 | [0.000] | 0.611 | [0.000] | 0.626 | [0.000] | 0.611 | [0.000] | 0.626 | [0.000] |
| (0.007) | (0.007) | (0.007) | (0.007) | (0.009) | (0.007) | (0.009) | ||||||||
| Two strikes | 0.613 | [0.000] | 0.611 | [0.000] | 0.611 | [0.000] | 0.611 | [0.000] | 0.623 | [0.000] | 0.611 | [0.000] | 0.622 | [0.000] |
| (0.007) | (0.007) | (0.007) | (0.007) | (0.009) | (0.007) | (0.009) | ||||||||
| Batter: All-Star | −0.033 | [0.000] | −0.033 | [0.000] | −0.033 | [0.000] | −0.033 | [0.000] | −0.033 | [0.000] | −0.035 | [0.000] | −0.035 | [0.000] |
| (0.003) | (0.003) | (0.003) | (0.003) | (0.003) | (0.004) | (0.004) | ||||||||
| Pitcher: All-Star | 0.097 | [0.000] | 0.064 | [0.000] | 0.065 | [0.000] | 0.064 | [0.000] | 0.064 | [0.000] | 0.064 | [0.000] | 0.065 | [0.000] |
| (0.004) | (0.005) | (0.005) | (0.005) | (0.005) | (0.005) | (0.005) | ||||||||
| Rookie | −0.073 | [0.000] | −0.033 | [0.001] | −0.034 | [0.001] | −0.033 | [0.001] | −0.033 | [0.001] | −0.033 | [0.001] | −0.034 | [0.001] |
| (0.010) | (0.010) | (0.010) | (0.010) | (0.010) | (0.010) | (0.010) | ||||||||
| Pitcher (right-handed) | 0.049 | [0.000] | 0.050 | [0.000] | 0.050 | [0.000] | 0.050 | [0.000] | 0.049 | [0.000] | 0.050 | [0.000] | 0.050 | [0.000] |
| (0.007) | (0.007) | (0.007) | (0.007) | (0.007) | (0.007) | (0.007) | ||||||||
| Pitcher: White | −0.036 | [0.000] | −0.035 | [0.000] | −0.035 | [0.000] | −0.035 | [0.000] | −0.035 | [0.000] | −0.035 | [0.000] | −0.035 | [0.000] |
| (0.008) | (0.008) | (0.008) | (0.008) | (0.008) | (0.008) | (0.008) | ||||||||
| Pitcher age | −0.005 | [0.000] | −0.005 | [0.000] | −0.006 | [0.000] | −0.005 | [0.000] | −0.005 | [0.000] | −0.005 | [0.000] | −0.006 | [0.000] |
| (0.001) | (0.001) | (0.001) | (0.001) | (0.001) | (0.001) | (0.001) | ||||||||
| Edge pitch | 0.392 | [0.000] | 0.392 | [0.000] | 0.392 | [0.000] | 0.392 | [0.000] | 0.392 | [0.000] | 0.392 | [0.000] | 0.392 | [0.000] |
| (0.068) | (0.068) | (0.068) | (0.068) | (0.068) | (0.068) | (0.068) | ||||||||
| Batter: White | −0.008 | [0.202] | −0.008 | [0.213] | −0.008 | [0.213] | −0.007 | [0.217] | −0.008 | [0.213] | −0.008 | [0.213] | −0.007 | [0.217] |
| (0.006) | (0.006) | (0.006) | (0.006) | (0.006) | (0.006) | (0.006) | ||||||||
| Batter age | −0.014 | [0.000] | −0.014 | [0.000] | −0.014 | [0.000] | −0.014 | [0.000] | −0.014 | [0.000] | −0.014 | [0.000] | −0.014 | [0.000] |
| (0.001) | (0.001) | (0.001) | (0.001) | (0.001) | (0.001) | (0.001) | ||||||||
| Batter (right-handed) | 0.140 | [0.000] | 0.139 | [0.000] | 0.139 | [0.000] | 0.139 | [0.000] | 0.139 | [0.000] | 0.139 | [0.000] | 0.139 | [0.000] |
| (0.007) | (0.007) | (0.007) | (0.007) | (0.007) | (0.007) | (0.007) | ||||||||
| Small score difference | 0.056 | [0.000] | 0.050 | [0.000] | 0.053 | [0.000] | 0.050 | [0.000] | 0.050 | [0.000] | 0.050 | [0.000] | 0.053 | [0.000] |
| (0.011) | (0.011) | (0.011) | (0.011) | (0.011) | (0.011) | (0.011) | ||||||||
| Attendance | 0.000 | [0.068] | 0.000 | [0.206] | 0.000 | [0.214] | 0.000 | [0.207] | 0.000 | [0.205] | 0.000 | [0.207] | 0.000 | [0.214] |
| (0.000) | (0.000) | (0.000) | (0.000) | (0.000) | (0.000) | (0.000) | ||||||||
| Bases loaded | 0.189 | [0.000] | 0.194 | [0.000] | 0.193 | [0.000] | 0.196 | [0.000] | 0.194 | [0.000] | 0.194 | [0.000] | 0.196 | [0.000] |
| (0.021) | (0.021) | (0.021) | (0.021) | (0.021) | (0.021) | (0.021) | ||||||||
| Home team | 0.019 | [0.003] | 0.019 | [0.004] | 0.018 | [0.005] | 0.019 | [0.004] | 0.019 | [0.004] | 0.019 | [0.004] | 0.018 | [0.005] |
| (0.007) | (0.007) | (0.007) | (0.007) | (0.007) | (0.007) | (0.007) | ||||||||
| Degrees | 0.001 | [0.000] | 0.001 | [0.000] | 0.001 | [0.000] | 0.001 | [0.000] | 0.001 | [0.000] | 0.001 | [0.000] | 0.001 | [0.000] |
| (0.000) | (0.000) | (0.000) | (0.000) | (0.000) | (0.000) | (0.000) | ||||||||
| Wind speed | −0.003 | [0.000] | −0.003 | [0.000] | −0.003 | [0.000] | −0.003 | [0.000] | −0.003 | [0.000] | −0.003 | [0.000] | −0.003 | [0.000] |
| (0.001) | (0.001) | (0.001) | (0.001) | (0.001) | (0.001) | (0.001) | ||||||||
| Constant | −2.536 | [0.000] | −2.583 | [0.000] | −2.604 | [0.000] | −2.575 | [0.000] | −2.592 | [0.000] | −2.582 | [0.000] | −2.605 | [0.000] |
| (0.066) | (0.065) | (0.065) | (0.065) | (0.065) | (0.065) | (0.065) | ||||||||
| Log-likelihood | −445737.58 | −445620.75 | −445607.42 | −445617.09 | −445617.20 | −445620.38 | −445600.24 | |||||||
| Umpire × year FE | Yes | Yes | Yes | Yes | Yes | Yes | Yes | |||||||
| Observations | 1,452,180 | 1,452,180 | 1,452,180 | 1,452,180 | 1,452,180 | 1,452,180 | 1,452,180 | |||||||
| DV: Overrecognition | Model 1 | Model 2 | Model 3 | Model 4 | Model 5 | Model 6 | Model 7 | |||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| b(s.e.) | p-val. | b(s.e.) | p-val. | b(s.e.) | p-val. | b(s.e.) | p-val. | b(s.e.) | p-val. | b(s.e.) | p-val. | b(s.e.) | p-val. | |
| Pitcher's past performance | 0.034 | [0.000] | 0.052 | [0.000] | 0.028 | [0.000] | 0.042 | [0.000] | 0.034 | [0.000] | 0.053 | [0.000] | ||
| (0.003) | (0.005) | (0.004) | (0.004) | (0.003) | (0.006) | |||||||||
| Pitcher's past performance × Inning | −0.004 | [0.000] | −0.004 | [0.000] | ||||||||||
| (0.001) | (0.001) | |||||||||||||
| Pitcher's past performance × Run expectancy | 0.015 | [0.010] | 0.014 | [0.014] | ||||||||||
| (0.006) | (0.006) | |||||||||||||
| Pitcher's past performance × One strike | −0.012 | [0.009] | −0.012 | [0.009] | ||||||||||
| (0.005) | (0.005) | |||||||||||||
| Pitcher's past performance × Two strikes | −0.009 | [0.050] | −0.009 | [0.051] | ||||||||||
| (0.005) | (0.005) | |||||||||||||
| Pitcher's past performance × Batter: All-Star | 0.001 | [0.395] | 0.001 | [0.537] | ||||||||||
| (0.002) | (0.002) | |||||||||||||
| Inning | 0.034 | [0.000] | 0.039 | [0.000] | 0.044 | [0.000] | 0.039 | [0.000] | 0.039 | [0.000] | 0.039 | [0.000] | 0.044 | [0.000] |
| (0.001) | (0.001) | (0.002) | (0.001) | (0.001) | (0.001) | (0.002) | ||||||||
| Run expectancy | −0.097 | [0.000] | −0.095 | [0.000] | −0.095 | [0.000] | −0.112 | [0.000] | −0.095 | [0.000] | −0.095 | [0.000] | −0.112 | [0.000] |
| (0.009) | (0.009) | (0.009) | (0.012) | (0.009) | (0.009) | (0.012) | ||||||||
| One strike | 0.611 | [0.000] | 0.611 | [0.000] | 0.611 | [0.000] | 0.611 | [0.000] | 0.626 | [0.000] | 0.611 | [0.000] | 0.626 | [0.000] |
| (0.007) | (0.007) | (0.007) | (0.007) | (0.009) | (0.007) | (0.009) | ||||||||
| Two strikes | 0.613 | [0.000] | 0.611 | [0.000] | 0.611 | [0.000] | 0.611 | [0.000] | 0.623 | [0.000] | 0.611 | [0.000] | 0.622 | [0.000] |
| (0.007) | (0.007) | (0.007) | (0.007) | (0.009) | (0.007) | (0.009) | ||||||||
| Batter: All-Star | −0.033 | [0.000] | −0.033 | [0.000] | −0.033 | [0.000] | −0.033 | [0.000] | −0.033 | [0.000] | −0.035 | [0.000] | −0.035 | [0.000] |
| (0.003) | (0.003) | (0.003) | (0.003) | (0.003) | (0.004) | (0.004) | ||||||||
| Pitcher: All-Star | 0.097 | [0.000] | 0.064 | [0.000] | 0.065 | [0.000] | 0.064 | [0.000] | 0.064 | [0.000] | 0.064 | [0.000] | 0.065 | [0.000] |
| (0.004) | (0.005) | (0.005) | (0.005) | (0.005) | (0.005) | (0.005) | ||||||||
| Rookie | −0.073 | [0.000] | −0.033 | [0.001] | −0.034 | [0.001] | −0.033 | [0.001] | −0.033 | [0.001] | −0.033 | [0.001] | −0.034 | [0.001] |
| (0.010) | (0.010) | (0.010) | (0.010) | (0.010) | (0.010) | (0.010) | ||||||||
| Pitcher (right-handed) | 0.049 | [0.000] | 0.050 | [0.000] | 0.050 | [0.000] | 0.050 | [0.000] | 0.049 | [0.000] | 0.050 | [0.000] | 0.050 | [0.000] |
| (0.007) | (0.007) | (0.007) | (0.007) | (0.007) | (0.007) | (0.007) | ||||||||
| Pitcher: White | −0.036 | [0.000] | −0.035 | [0.000] | −0.035 | [0.000] | −0.035 | [0.000] | −0.035 | [0.000] | −0.035 | [0.000] | −0.035 | [0.000] |
| (0.008) | (0.008) | (0.008) | (0.008) | (0.008) | (0.008) | (0.008) | ||||||||
| Pitcher age | −0.005 | [0.000] | −0.005 | [0.000] | −0.006 | [0.000] | −0.005 | [0.000] | −0.005 | [0.000] | −0.005 | [0.000] | −0.006 | [0.000] |
| (0.001) | (0.001) | (0.001) | (0.001) | (0.001) | (0.001) | (0.001) | ||||||||
| Edge pitch | 0.392 | [0.000] | 0.392 | [0.000] | 0.392 | [0.000] | 0.392 | [0.000] | 0.392 | [0.000] | 0.392 | [0.000] | 0.392 | [0.000] |
| (0.068) | (0.068) | (0.068) | (0.068) | (0.068) | (0.068) | (0.068) | ||||||||
| Batter: White | −0.008 | [0.202] | −0.008 | [0.213] | −0.008 | [0.213] | −0.007 | [0.217] | −0.008 | [0.213] | −0.008 | [0.213] | −0.007 | [0.217] |
| (0.006) | (0.006) | (0.006) | (0.006) | (0.006) | (0.006) | (0.006) | ||||||||
| Batter age | −0.014 | [0.000] | −0.014 | [0.000] | −0.014 | [0.000] | −0.014 | [0.000] | −0.014 | [0.000] | −0.014 | [0.000] | −0.014 | [0.000] |
| (0.001) | (0.001) | (0.001) | (0.001) | (0.001) | (0.001) | (0.001) | ||||||||
| Batter (right-handed) | 0.140 | [0.000] | 0.139 | [0.000] | 0.139 | [0.000] | 0.139 | [0.000] | 0.139 | [0.000] | 0.139 | [0.000] | 0.139 | [0.000] |
| (0.007) | (0.007) | (0.007) | (0.007) | (0.007) | (0.007) | (0.007) | ||||||||
| Small score difference | 0.056 | [0.000] | 0.050 | [0.000] | 0.053 | [0.000] | 0.050 | [0.000] | 0.050 | [0.000] | 0.050 | [0.000] | 0.053 | [0.000] |
| (0.011) | (0.011) | (0.011) | (0.011) | (0.011) | (0.011) | (0.011) | ||||||||
| Attendance | 0.000 | [0.068] | 0.000 | [0.206] | 0.000 | [0.214] | 0.000 | [0.207] | 0.000 | [0.205] | 0.000 | [0.207] | 0.000 | [0.214] |
| (0.000) | (0.000) | (0.000) | (0.000) | (0.000) | (0.000) | (0.000) | ||||||||
| Bases loaded | 0.189 | [0.000] | 0.194 | [0.000] | 0.193 | [0.000] | 0.196 | [0.000] | 0.194 | [0.000] | 0.194 | [0.000] | 0.196 | [0.000] |
| (0.021) | (0.021) | (0.021) | (0.021) | (0.021) | (0.021) | (0.021) | ||||||||
| Home team | 0.019 | [0.003] | 0.019 | [0.004] | 0.018 | [0.005] | 0.019 | [0.004] | 0.019 | [0.004] | 0.019 | [0.004] | 0.018 | [0.005] |
| (0.007) | (0.007) | (0.007) | (0.007) | (0.007) | (0.007) | (0.007) | ||||||||
| Degrees | 0.001 | [0.000] | 0.001 | [0.000] | 0.001 | [0.000] | 0.001 | [0.000] | 0.001 | [0.000] | 0.001 | [0.000] | 0.001 | [0.000] |
| (0.000) | (0.000) | (0.000) | (0.000) | (0.000) | (0.000) | (0.000) | ||||||||
| Wind speed | −0.003 | [0.000] | −0.003 | [0.000] | −0.003 | [0.000] | −0.003 | [0.000] | −0.003 | [0.000] | −0.003 | [0.000] | −0.003 | [0.000] |
| (0.001) | (0.001) | (0.001) | (0.001) | (0.001) | (0.001) | (0.001) | ||||||||
| Constant | −2.536 | [0.000] | −2.583 | [0.000] | −2.604 | [0.000] | −2.575 | [0.000] | −2.592 | [0.000] | −2.582 | [0.000] | −2.605 | [0.000] |
| (0.066) | (0.065) | (0.065) | (0.065) | (0.065) | (0.065) | (0.065) | ||||||||
| Log-likelihood | −445737.58 | −445620.75 | −445607.42 | −445617.09 | −445617.20 | −445620.38 | −445600.24 | |||||||
| Umpire × year FE | Yes | Yes | Yes | Yes | Yes | Yes | Yes | |||||||
| Observations | 1,452,180 | 1,452,180 | 1,452,180 | 1,452,180 | 1,452,180 | 1,452,180 | 1,452,180 | |||||||
Note(s): *Robust standard errors clustered at the game level are reported in parentheses; p-values are rounded to three decimal places and reported in brackets
Hypothesis 1 predicted that umpires are more likely to overrecognize a pitch when the pitcher has higher prior-season performance, a direct indicator of anchoring bias in performance evaluation. In Model 2, the coefficient for pitchers prior performance is positive and significant (b = 0.034, p < 0.001), indicating that each additional win above replacement in the prior season for a given pitcher increases the odds of the umpire awarding a gifted strike by approximately 3.5%. To convey the practical magnitude of this effect, consider a realistic comparison within the sample: moving from a replacement-level pitcher (WAR = 0) to an elite starter (WAR ≈ 6) is associated with approximately 23% higher odds of receiving a gifted strike. This is a substantively meaningful premium operating across every pitch thrown by high-performing pitchers throughout a season. The results strongly and consistently support H1 that evaluators' assessments of current-period performance will be biased in the direction of the evaluee's prior performance, such that stronger prior performance will be associated with more favorable assessments than indicated by objective current-period performance measures.
Hypothesis 2 predicted that anchoring bias would intensify in later innings as umpire fatigue accumulates. Model 3 reports the results of the test of this hypothesis. The coefficient on the interaction of pitcher's prior performance and inning is statistically significant and negative (b = −0.004, p < 0.001) indicating an association in the direction opposite our predictions. Figure 2 graphs this interaction, illustrating that while prior performance is positively associated with PEB in both earlier and later innings, this association is weaker in later innings.
The line graph presents the association between pitcher’s past performance and performance evaluation bias (PEB), moderated by inning number. The x-axis represents pitcher’s past performance, ranging from low to high, and the y-axis represents PEB, ranging from 0.080 to 0.110. Two lines are depicted: a solid line with filled circles representing low inning numbers and a dashed line with open triangles representing high inning numbers. Both lines slope upward, indicating that PEB increases as pitcher’s past performance increases. The flatter slope of the dashed line indicates that this association is weaker in later innings. All plotted values are approximate.Pitcher's past performance × inning
The line graph presents the association between pitcher’s past performance and performance evaluation bias (PEB), moderated by inning number. The x-axis represents pitcher’s past performance, ranging from low to high, and the y-axis represents PEB, ranging from 0.080 to 0.110. Two lines are depicted: a solid line with filled circles representing low inning numbers and a dashed line with open triangles representing high inning numbers. Both lines slope upward, indicating that PEB increases as pitcher’s past performance increases. The flatter slope of the dashed line indicates that this association is weaker in later innings. All plotted values are approximate.Pitcher's past performance × inning
These results confirm that umpires' pitch calls are systematically biased toward a pitcher's prior performance, but that performance-anchored PEB is attenuated – not amplified as predicted in Hypothesis 2 – in later innings. There are several possible explanations for this finding. One possibility is that umpires become less attuned to pitcher identity cues as the game advances, with accumulated cognitive and perceptual demands reducing the salience of prior performance anchors rather than increasing heuristic reliance on them. It is possible that the pitcher's accumulated current game performance makes it easier to systematically process their next pitch; that is, umpires may not rely on their anchor of prior performance as current performance information accumulates during the game. Yet another alternative explanation is structural: pitchers who remain in the game into late innings do so because they are pitching well that day, reducing the divergence between reputation and contemporaneous performance and thus leaving less room for performance bias to manifest. Finally, it is possible that the negative coefficient we observe here is a result of the mechanism that we focused on in Hypothesis 3: that the stakes become greater later in the game and this is triggering more systematic processing. Regardless of these possible alternative explanations, we do not find support for Hypothesis 2, that the influence of the evaluee's prior performance on assessments of current-period performance will become stronger as evaluator fatigue increases.
Hypothesis 3 predicted that anchoring bias would be weaker in high-stakes game situations. We test this prediction in three models interacting pitcher's prior performance with run expectancy, ball-strike count and batter All-Star appearances. Model 4 reports the results for the test using the run expectancy interaction. The coefficient is positive and significant (b = 0.015, p = 0.010). Contrary to predictions, this interaction indicates that the positive association between the pitcher's prior-season performance and PEB grows stronger as the game situation becomes more consequential: in high run-expectancy situations, umpires are increasingly biased by the pitcher's prior performance. Figure 3 illustrates the positive interaction between past performance and run expectancy suggesting that past performance exerts greater influence in situations with higher expected offensive output.
The line graph presents the association between pitcher’s past performance and performance evaluation bias (PEB), moderated by run expectancy. The x-axis represents pitcher’s past performance, ranging from low to high, and the y-axis represents PEB, ranging from 0.085 to 0.110. Two lines are depicted: a solid line with filled circles representing low run expectancy and a dashed line with open triangles representing high run expectancy. The low-run-expectancy line increases from approximately 0.095 to 0.105, while the high-run-expectancy line increases from approximately 0.090 to 0.100. Both lines slope upward, with higher PEB shown in low-run-expectancy situations. All plotted values are approximate.Pitcher's past performance × run expectancy
The line graph presents the association between pitcher’s past performance and performance evaluation bias (PEB), moderated by run expectancy. The x-axis represents pitcher’s past performance, ranging from low to high, and the y-axis represents PEB, ranging from 0.085 to 0.110. Two lines are depicted: a solid line with filled circles representing low run expectancy and a dashed line with open triangles representing high run expectancy. The low-run-expectancy line increases from approximately 0.095 to 0.105, while the high-run-expectancy line increases from approximately 0.090 to 0.100. Both lines slope upward, with higher PEB shown in low-run-expectancy situations. All plotted values are approximate.Pitcher's past performance × run expectancy
In comparison with the run expectancy interaction, the interaction of pitcher prior performance with the strike count variables provides conflicting evidence. In Model 5, the coefficients on the interactions of both the two-strikes and one-strike count interactions are negative and significant (b = −0.009, p = 0.050 and b = −0.012, p = 0.009, respectively). These findings indicate that in situations where the umpire's call is most consequential for the outcome of the at-bat, umpires appear to rely less on the prior performance as an anchor in heuristic processing when calling balls and strikes: perhaps they are exercising more restraint and switching to systematic processing when a mistaken strike call would directly end the at-bat. Figure 4 shows that past performance is positively associated with PEB at each strike count, but the association is weaker when the batter has one or two strikes than when the count contains no strikes, suggesting that strike pressure attenuates the influence of past performance on PEB.
The line graph presents the association between pitcher’s past performance and performance evaluation bias (PEB), moderated by strike count. The x-axis represents pitcher’s past performance, ranging from low to high, and the y-axis represents PEB, ranging from 0.060 to 0.130. Three lines are depicted, representing zero, one, and two strikes. The zero-strike line increases from approximately 0.065 to 0.075. The one-strike and two-strike lines each increase from approximately 0.110 to 0.125. All three lines slope upward, indicating that PEB increases as pitcher’s past performance increases, although the levels and slopes differ by strike count. All plotted values are approximate.Pitcher's past performance × strike count
The line graph presents the association between pitcher’s past performance and performance evaluation bias (PEB), moderated by strike count. The x-axis represents pitcher’s past performance, ranging from low to high, and the y-axis represents PEB, ranging from 0.060 to 0.130. Three lines are depicted, representing zero, one, and two strikes. The zero-strike line increases from approximately 0.065 to 0.075. The one-strike and two-strike lines each increase from approximately 0.110 to 0.125. All three lines slope upward, indicating that PEB increases as pitcher’s past performance increases, although the levels and slopes differ by strike count. All plotted values are approximate.Pitcher's past performance × strike count
Finally, Model 6 reports the model testing whether the batter's status affects the umpire's PEB. We do not observe a statistically significant association for the interaction of pitcher prior performance and batter: All-Star. Taken together, the three tests of how high-stakes game situations affect PEB provide inconclusive results for Hypothesis 3, that the influence of the evaluee's prior performance on assessments of current-period performance will become weaker as the stakes of the evaluation increase. While we find support for this prediction when measuring high stakes in terms of strike count, we do not when operationalizing batter's status. Moreover, we find contradictory results when considering how run expectancy affects the stakes of the umpire's call. These results suggest that the association between situational stakes and anchoring bias is nuanced and may hinge on what, exactly, is at stake in these situations. When the stakes are elevated at the game level (run expectancy), performance-anchored PEB intensifies; when the stakes are elevated at the pitch level (two-strike count), umpires exhibit a corrective tendency; and when stakes are operationalized using batter status, there is no statistically significant moderating effect on the relationship between pitcher prior performance and the umpire's PEB.
Discussion
This study examined whether an evaluee's prior performance biases professional evaluators' assessments of current performance and whether situational conditions strengthen or weaken that bias. Using objectively measured pitch location as a benchmark for MLB umpires' subjective calls, we find that pitchers with stronger prior-season performance receive more favorable evaluations than pitchers with weaker performance histories. This suggests evidence of prior performance-anchored PEB: an established performance record provides an initial reference point from which evaluators adjust insufficiently when assessing current evidence (Epley and Gilovich, 2006; Teovanović, 2019; Tversky and Kahneman, 1974). The finding extends research showing that prior success, reputation and status can carry over into subsequent subjective evaluations (Kim and King, 2014; Waguespack and Salomon, 2016; Zuckerman, 1977) by demonstrating that such effects persist among highly trained evaluators making repeated judgments within a familiar professional domain.
Importantly, our results suggest that anchoring-based PEB is context-sensitive, but not in a predictable and consistent way. We expected the evaluator's fatigue to strengthen heuristic reliance on prior performance anchors and higher decision stakes to weaken it by encouraging more systematic processing (Eagly and Chaiken, 1993; Epley and Gilovich, 2006; Lerner and Tetlock, 1999; Payne et al., 1993). Instead, we found interaction patterns that vary across situational conditions. Anchoring-based PEB decreases in later innings, increases as run expectancy rises, decreases when the count contains one or two strikes and is not significantly associated with the batter's All-Star status. Taken together, these results do not support our theoretical expectation that fatigue increases anchoring or that higher-stakes decisions reduce it.
The negative interaction with inning complicates the fatigue mechanism we discussed in our theoretical development. We predicted that accumulated cognitive fatigue would make umpires less able to shift away from heuristic processing (Epley and Gilovich, 2006; Gilbert et al., 1988; van der Linden et al., 2003). Instead, anchoring-based PEB declines as the game progresses. Several explanations are possible. Fatigue may reduce the salience of pitcher identity and prior reputation, later innings may include a selected group of pitchers whose current performance is more consistent with their established records, or umpires may become more attentive as the consequences of late-game calls increase. Future research could distinguish among these explanations by measuring evaluators' cognitive fatigue directly, tracking changes in attention across repeated judgments, or examining whether the evaluator's accuracy changes systematically over the course of an evaluation.
The contrasting findings for the moderating associations of run expectancy and strike count suggest that different forms of decision stakes may lead to different responses. Higher run expectancy increases the potential consequences of a call that impacts the broader game situation, and under these conditions umpires rely more heavily on the pitcher's prior performance. By contrast, one- and two-strike counts make the immediate consequence of the call more transparent because an additional strike can lead to the end of the batter's plate appearance; in these situations, reliance on prior-performance anchors declines. One possible interpretation is that diffuse, game-level consequences increase uncertainty and encourage reliance on established performance cues, whereas immediate and clearly attributable consequences prompt more systematic processing. Future research could provide greater insight into the mechanisms underlying these contrasting effects by examining the salience of different contextual outcomes associated with the evaluator's decision-making.
The nonsignificant interaction with batter All-Star status provides a further boundary to the situational argument. Although prominent pitcher–batter matchups may attract greater attention, batter status does not significantly alter the association between the pitcher's prior performance and PEB. This result suggests that the bias is not simply a response to the overall prominence of the matchup. Rather, the relevant prior information appears to be tied more closely to the player whose performance is being judged – the pitcher – than to the status of the opposing batter (Kim and King, 2014). Because a null interaction cannot establish that the anchor is exclusively pitcher-specific, future research could examine whether evaluators anchor on different information about the evaluee, the comparison or the broader contextual status cues.
Taken together, these findings expand – and challenge – our understanding of when prior-performance anchors shape evaluation and lead to PEB. The main effect shows that established performance records can bias judgments of current performance, while the moderation results show that this bias does not respond uniformly to fatigue or decision stakes. Instead, the influence of prior performance depends on the specific situational condition surrounding the judgment. This pattern suggests that theories of evaluation bias should take a careful and nuanced perspective on the impact of situational factors. Future research may identify which features of an evaluation context make prior-performance information more or less salient, and when evaluators shift from heuristic processing based on established impressions to systematic processing of current evidence.
Implications for practice
Our results demonstrate how performance evaluators can be influenced by performance-anchored PEB. These findings suggest that organizations should be cautious about assuming that experienced evaluators will naturally discount prior performance when assessing current results. Even highly trained professionals may rely on established performance histories when judgments must be made quickly and repeatedly. Reducing this bias may therefore require evaluation systems that direct attention explicitly toward current-performance evidence; for example, by using clearly defined criteria, separating prior records from current assessments where feasible, and requiring evaluators to justify judgments against contemporaneous benchmarks.
More broadly, the mixed moderation results suggest that increasing the stakes or pressure surrounding an evaluation is unlikely to improve objectivity consistently. Reducing bias may instead require redesigning the evaluation process. For example, organizations might clarify performance criteria, limit unnecessary discretion, reduce ambiguity and ensure that evaluators consider current evidence before reviewing prior performance records. In domains such as promotion decisions, academic reviewing, hiring, legal judgment and professional sports officiating, the design of the evaluation process may therefore matter more than the context surrounding the decision.
Limitations
As with all studies based on professional sports data, concerns may arise about the generalizability of the findings to other organizational settings. For instance, one might argue that comparing the umpiring of balls and strikes with typical performance evaluations in a company is fundamentally different. Umpires are required to make rapid decisions, whereas typical business performance evaluations are not bound by such strict time constraints. This discrepancy raises questions regarding the applicability of our findings to broader business contexts. However, there are instances in business settings where performance evaluators must make quick decisions under high-pressure conditions, including crisis management, hiring and other time-sensitive decisions. One prominent example is the job interview. In this situation, managers must evaluate a candidate's skills and future performance within a very limited timeframe, often relying on brief, structured interactions. Research in the human resources management literature documents the existence of cognitive bias in interviewing, where time is constrained and interviewers must deal with limited information (Barrick et al., 2009; Bragger et al., 2002; Kutcher and Bragger, 2004; Macan, 2009). Beyond interviews, routine workplace evaluations are also influenced by PEB. For instance, decisions about assigning stretch projects, providing real-time feedback and nominating individuals for leadership positions are often made spontaneously, based on recent and prominent performance cues. These evaluations, although less formalized, can have significant implications for career development and organizational performance. Thus, although the speed and repetition of umpire decisions distinguish our setting from many formal performance appraisals, the underlying tension between heuristic reliance on prior impressions and systematic processing of current evidence is relevant to a range of organizational evaluations.
Our setting may also provide insight into evaluations made under high-stakes conditions. In negotiations, crisis management, hiring and other consequential decisions, evaluators may rely on prior performance when assessing current evidence. However, our results show that higher stakes do not affect anchoring-based PEB uniformly: run expectancy strengthens the association between prior performance and favorable evaluation, whereas one- and two-strike counts weaken it. In order to fully inform practice, future research should examine whether similar differences emerge across organizational settings and identify when high-stakes conditions encourage heuristic reliance on established impressions rather than systematic processing of current evidence.
Another limitation of this study is that we examine only a limited set of situational conditions that may influence anchoring-based PEB. Individual differences among evaluators, relationships between evaluators and evaluees, organizational incentives and characteristics of the evaluation system may also affect whether prior performance shapes current judgments. Future research should develop a broader contingency approach that examines how these factors influence the relative use of heuristic processing based on established impressions and systematic processing of current-performance evidence.
Conclusion
In this study, we investigated the circumstances under which performance-anchored PEB is strengthened or weakened and demonstrated that situational factors play a significant role in shaping evaluators' PEB. Our aim was to gain insight into how these situational factors moderate evaluators' PEB through their influence on heuristic and systematic processing. Through this approach, we sought to re-examine PEB from a cognitive perspective, acknowledging the interplay of various contextual factors. The findings suggest that while past performance tends to increase PEB, situational factors do not uniformly attenuate performance-anchored PEB. Our study represents a step forward in understanding the complex dynamics of PEB and factors that can moderate it. By laying this groundwork, we hope to inspire future researchers to build on our findings and explore anchoring-based PEB across diverse contexts.
We thank Kwame Agyemang and two anonymous reviewers for their comments and suggestions, which significantly improved this paper. Rhett Brymer, Xu Huang, Rebecca Kehoe, and seminar participants at the Vrije Universiteit Amsterdam provided valuable feedback on earlier drafts. All remaining errors are our own. Earlier versions of the paper were presented at the 2020 Annual Meeting of the Academy of Management and nominated for the 2019 William H. Newman Award. This study was partially supported by the Korea University Business School Research Grant and the Insung Research Grant.
Notes
Given the fact that batters vary in their height, the size of the strike zone necessarily changes for each batter. Moreover, despite the computer-generated strike zone measurement we discuss below, each umpire ultimately determines the size of the strike zone themselves. As a result, there is not a single universally-sized strike zone for every batter in every game. Nonetheless, umpires are notably consistent in applying their respective strike zones, although some are known to be “pitcher-” or “batter-friendly,” applying a larger or smaller strike zone, respectively. Of course, batters are aware of the tendency of the home plate umpire for any given game and accommodate accordingly. See Huang and Hsu (2020) and Marchi and Albert (2013) for related empirical explorations.
This short summary simplifies the (sometimes complex) rules associated with pitching and batting in baseball. For example, a strike may also be called if the batter swings at the ball and does not hit it or hits it outside the field of play (i.e. “a foul”). The important and relevant point here is that pitchers try to get batters to swing at pitches that are balls (i.e. too far outside the strike zone to be hit well), and not swing at pitches that are strikes, which are more likely to be put into play and result in runs (i.e. points) for the batter's team. Batters, conversely, try to not swing at pitches that are balls, and swing only at those that are strikes and which they can put into play.
MLB first introduced limited instant-replay review in 2008 and expanded it to most types of on-field calls in 2014. Replay reviews can confirm or overturn an umpire's original call; however, ball-and-strike calls were not reviewable during our 2015–2018 sampling period. Separately, MLB used pitch-tracking data as part of its Zone Evaluation system to provide umpires with postgame feedback about the accuracy of their ball-and-strike calls. This retrospective evaluation did not alter the calls made during the game. Beginning with the 2026 season, MLB introduced the Automated Ball-Strike Challenge System, under which a batter, pitcher or catcher may challenge a ball-and-strike call. This challenge system was not available during our sampling period.
The number of pitches a starting pitcher throws in each game has been declining in recent years. From 1988 (when pitch count data started being collected) to 2019, the average was 94 per game (Potter, 2024).
While WAR is commonly used to measure pitchers' performance, we tested whether our models were robust to an alternative measure of prior performance, “Value Over Replacement Player.” VORP compares the focal player's performance to the performance of a hypothetical “replacement” player representing the average player's performance at the same position and league. The WAR models we report are almost identical to this alternative performance specification (results available from the authors).
In Major League Baseball (MLB), the All-Star Game is an annual exhibition game that features the best players from both the American League (AL) and the National League (NL), the two leagues that make up MLB. The players are selected based on fan, player and coach votes, as well as selections by the league's managers. The game typically occurs in July, during the middle of the MLB season, and serves as a showcase of the top talent in the league.

