The Commonplace
Home Three-study pilot Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Major League umpires favor high-performing pitchers: pitches outside the strike zone are more likely to be called strikes when the pitcher had stronger prior-season performance, and this reputational anchoring shifts with game context—rising when run expectancy is high but falling later in games and when batters already have strikes.

Past performance, present judgment and situational influences: performance evaluation bias in Major League Baseball
Yeongsu Anthony Kim, Jongsoo Jays Kim, Thomas P. Moliterno · September 16, 2026 · Sport Business and Management An International Journal
openalex correlational medium evidence 7/10 relevance Summary only summary available; pdf_status=paywall DOI Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

OpenAlex

Latest observation:

  1. Yeongsu Anthony Kim provider ID
  2. Jongsoo Jays Kim provider ID
  3. Thomas P. Moliterno provider ID

Semantic Scholar

Latest observation:

  1. Yeongsu Anthony Kim unresolved corpus identity
  2. Jongsoo Jays Kim unresolved corpus identity
  3. Thomas P. Moliterno unresolved corpus identity
Umpires are more likely to call objectively out-of-zone pitches strikes for pitchers with higher prior-season WAR, and the magnitude of this anchoring-based performance-evaluation bias varies with game context (e.g., run expectancy, inning, strike count).

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

Purpose This study examines whether an evaluee's prior performance anchors professional evaluators' assessments of current performance and whether situational conditions alter this performance evaluation bias (PEB). Integrating theory on heuristic and systematic models of information processing, it tests whether evaluator fatigue strengthens anchoring-based PEB and whether different forms of decision stakes weaken it. Design/methodology/approach The study analyzes 1,452,180 ball-and-strike calls made by Major League Baseball umpires during 2015–2018. Computerized pitch-location data provide an objective benchmark for each subjective call. Generalized linear models test whether pitchers' prior-season Wins Above Replacement predicts gifted strikes on objectively out-of-zone pitches and whether inning, run expectancy, strike count and batter All-Star appearances moderate this association. Findings Pitchers with stronger prior-season performance receive more gifted strikes, consistent with anchoring-based PEB. The moderating effects are neither uniform nor always as predicted. PEB decreases rather than increases in later innings, increases with run expectancy, decreases when the batter has one or two strikes and does not vary significantly with batter All-Star status. Situational conditions can therefore attenuate or intensify evaluators' reliance on prior-performance anchors. Originality/value The study positions anchoring as a central cognitive mechanism underlying PEB and extends prior research on status and reputation by demonstrating prior-performance effects among highly trained professional evaluators whose judgments can be compared with an objective benchmark. It also advances a context-sensitive account of PEB by showing that fatigue and decision stakes do not have uniform effects and that distinct forms of stakes can have opposing associations with evaluative bias.

Summary

Main Finding

Umpires' ball-and-strike calls show anchoring-based performance-evaluation bias (PEB): pitchers with stronger prior-season performance receive more "gifted" strikes (calls for strikes on pitches objectively outside the strike zone). Situational factors alter this bias—some attenuate it and others intensify it—so fatigue and decision stakes do not have uniform effects on evaluators' reliance on prior-performance anchors.

Key Points

  • Data: 1,452,180 ball-and-strike calls by Major League Baseball umpires (2015–2018) matched to objective pitch-location data.
  • Core result: Higher prior-season pitcher performance (measured by WAR) predicts a greater likelihood that umpires call out-of-zone pitches strikes (anchoring-based PEB).
  • Moderators tested:
    • Inning (proxy for fatigue): contrary to expectation, PEB decreases in later innings.
    • Run expectancy (proxy for stakes): PEB increases when run expectancy is higher.
    • Strike count (batter has 1 or 2 strikes): PEB decreases when the batter already has strikes.
    • Batter All-Star status (reputation/status): no significant moderating effect.
  • Interpretation: Anchoring on prior performance is a central cognitive mechanism driving evaluative bias among even highly trained professionals, but situational context (types of stakes, temporal factors) can either amplify or reduce reliance on that anchor.
  • Contribution: Extends status/reputation literature by demonstrating prior-performance effects against an objective benchmark and offering a context-sensitive account of PEB.

Data & Methods

  • Sample: 1,452,180 pitch calls from MLB games, 2015–2018.
  • Benchmark: Computerized pitch-location data determine whether a pitch was objectively inside or outside the strike zone.
  • Dependent variable: Gifted strike indicator — umpire called a strike on an objectively out-of-zone pitch.
  • Key independent variable: Pitcher prior-season performance (Wins Above Replacement, WAR).
  • Moderators/controls: Inning (time in game), run expectancy (game-state stakes), current pitch/at-bat strike count, batter All-Star appearances (status), and other standard controls for pitch/game context.
  • Statistical approach: Generalized linear models (logistic-type models) testing the association between prior-season performance and gifted strikes and interactions with situational moderators.
  • Robustness: Large-sample analysis with objective ground truth allows precise estimation of bias patterns among professional evaluators.

Implications for AI Economics

  • Human evaluators anchor on historical performance even when objective, moment-level data are available. In AI/economics contexts (e.g., model evaluation, forecasting, human-in-the-loop systems), anchoring on prior model or agent reputation can bias assessments of current outputs.
  • Context matters: Different forms of decision stakes (e.g., economic cost vs. time pressure) and temporal factors (fatigue) can have opposing effects on bias. Policy and system design should not assume uniform moderation by “stakes” or workload.
  • Design recommendations:
    • Use objective benchmarks and automated metrics where feasible to reduce reliance on reputational anchors.
    • Randomize or withhold prior-performance signals during evaluations when unbiased assessment of current outputs is needed.
    • Monitor and mitigate contextual drivers of bias—e.g., adjust oversight or review protocols under high run-expectancy/ high-stakes situations where anchoring may increase.
    • Incorporate debiasing training and decision aids for human evaluators to reduce anchoring (e.g., structured checklists, anonymized evaluations).
    • Model human-in-the-loop behavior: include anchoring priors and context-dependent moderator effects when simulating evaluator decisions or designing incentive schemes.
  • Research implications:
    • Replicate in other high-stakes evaluation settings (hiring, credit underwriting, peer review) to map generalizability.
    • Test interventions (signal removal, debiasing nudges, automated adjudication) and measure effects across different stake types and temporal loads.
    • Incorporate measured fatigue and richer stake metrics to better predict when anchoring will intensify or attenuate.
  • Economic modeling: When designing markets or platforms that rely on reputation signals, account for asymmetric and context-dependent evaluator biases; reputational advantages can persist or distort allocation of opportunities even with objective performance information available.

Assessment

Paper Typecorrelational Evidence Strengthmedium — High-quality objective outcome measurement (computerized pitch location) and very large sample give precise estimates of systematic associations, but causal interpretation is limited because prior-season performance is not exogenously assigned and may correlate with unobserved pitch- or batter-level features (pitch movement, speed, framing, pitcher/catcher strategies) that also influence umpire calls. Methods Rigormedium — Strong measurement and large-N logistic modelling with relevant moderators are strengths; however, the description does not document use of stronger identification techniques (e.g., instrumental variables, within-umpire or within-pitcher fixed effects that would isolate exogenous variation, event-study designs, or randomization of information). Potential confounders like pitch type, release point, catcher framing, umpire identity/time-in-sample, and matchups may be inadequately addressed or remain endogenous. Sample1,452,180 Major League Baseball pitch calls from 2015–2018 matched to computerized pitch-location data; dependent variable is indicator that an umpire called a strike when the pitch was objectively outside the strike zone; key independent variable is pitcher prior-season Wins Above Replacement (WAR); models include moderators (inning, run expectancy, current strike count, batter All-Star status) plus standard pitch/game-level controls. Themeshuman_ai_collab org_design IdentificationObservational association: uses large-sample pitch-level logistic models relating prior-season pitcher performance (WAR) to probability an objectively out-of-zone pitch is called a strike, controlling for pitch- and game-level covariates and testing interactions with situational moderators (inning, run expectancy, count, batter status). No exogenous variation or random assignment is reported. GeneralizabilitySports-specific setting: professional MLB umpires and the strike-call task may differ in important ways from other evaluative contexts (e.g., hiring, credit decisions, peer review)., Pitcher prior-season WAR may proxy for unobserved pitch characteristics (movement, velocity, pitch mix) that affect perception—limits external validity to contexts where prior status is plausibly exogenous to moment-level signal quality., Moderators (inning, run expectancy, strike count) are specific operationalizations of fatigue and stakes that may not map cleanly onto other domains’ workload or stakes., Findings are based on 2015–2018 data and MLB institutional details (e.g., rules, training, umpire evaluation systems) that may differ across leagues, levels, or time periods.

Claims (6)

ClaimDirectionOutcomeConfidence & EvidenceDetails
Pitchers with higher prior-season performance, measured by Wins Above Replacement (WAR), are more likely to receive gifted strikes: umpires call objectively out-of-zone pitches strikes more often for these pitchers. Decision Quality positive Probability that an objectively out-of-zone pitch is called a strike
Reading fidelity high
Study strength high
n=1452180
0.5
The anchoring-based performance-evaluation bias decreases in later innings. Decision Quality negative Strength of the association between prior-season pitcher performance and gifted-strike calls across innings
Reading fidelity high
Study strength medium
n=1452180
0.3
The anchoring-based performance-evaluation bias increases when run expectancy is higher. Decision Quality positive Strength of the association between prior-season pitcher performance and gifted-strike calls at different levels of run expectancy
Reading fidelity high
Study strength medium
n=1452180
0.3
The anchoring-based performance-evaluation bias decreases when the batter already has one or two strikes. Decision Quality negative Strength of the association between prior-season pitcher performance and gifted-strike calls by batter strike count
Reading fidelity high
Study strength medium
n=1452180
0.3
Batter All-Star status does not significantly moderate the relationship between pitcher prior-season performance and gifted-strike calls. Decision Quality null_result Moderation of gifted-strike calls by batter All-Star status
Reading fidelity high
Study strength medium
n=1452180
0.3
Situational factors do not uniformly moderate reliance on prior-performance anchors: later innings and existing batter strikes attenuate the bias, whereas higher run expectancy intensifies it. Decision Quality mixed Context-dependent variation in anchoring-based performance-evaluation bias
Reading fidelity high
Study strength medium
n=1452180
0.3

Notes