0 cumulative citations
View corpus contextAutomated umpiring in the KBO reveals where human judgment shifted the strike zone: marginal pitches were substantially more likely to be called strikes on 3–0 counts and far less likely on two-strike counts under human umpires, patterns that largely vanished after the league adopted the Automated Ball-Strike system.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
This paper uses the Korean Baseball Organization's adoption of the Automated Ball-Strike (ABS) system to audit long-standing claims about contextual bias in human ball-strike calls. Using pitch-level KBO data from 2021 through the available portion of the 2026 season, we model called-strike probability for taken pitches near the strike-zone boundary, with 2022-2023 as the primary human-umpire baseline and ABS seasons (2024 and onward) as a diagnostic benchmark. The strongest evidence concerns count pressure. Relative to 0--0 counts, human umpires called substantially fewer strikes in two-strike counts and more strikes in hitter-ahead three-ball counts. Specifically, in the main 0.25-ft boundary band, 0--2 was associated with a -17.17 percentage-point effect and 3--0 with a +6.61 percentage-point effect. Under ABS, the corresponding effects were close to zero and did not survive false-discovery-rate correction. Game progression shows a smaller but coherent pattern as human calls were less strike-prone in early innings and more strike-prone in innings 7--9+, especially in late-close situations, while complete ABS seasons were essentially flat. Other suspected biases are weaker or more localized. Salary-based reputation proxies provide suggestive but proxy-sensitive evidence, and catcher identity shows human-period residual heterogeneity that disappears under ABS. Home-context evidence is mostly null at the umpire level, with one FDR-significant human-period exception and an exploratory umpire-team gap best treated as an audit lead. Overall, the results do not show that human umpires were biased everywhere. Instead, they map where the human strike zone was most context-sensitive, where evidence was weaker, and where common suspicions received little support.
Summary
Main Finding
Human KBO umpires exhibited context-sensitive called-strike behavior concentrated near the rule-zone boundary, most clearly along count pressure: compared with 0–0, two-strike counts were substantially less likely to produce called strikes and 3–0 counts were substantially more likely. These count-dependent shifts largely disappear under the Automated Ball-Strike (ABS) system (2024+), consistent with those patterns being features of human judgment rather than pitch-selection or other mechanical factors. Other suspected biases (player reputation, game-progression, catcher effects, home advantage) show weaker, more localized, or proxy-sensitive evidence; some residual heterogeneity in the human period (e.g., by catcher) disappears under ABS.
Key Points
- Strongest and clearest result: count pressure.
- In the 0.25-ft boundary band (borderline pitches), human-period effects vs. 0–0:
- 0–2: −17.17 percentage points (pp)
- 3–0: +6.61 pp
- Under ABS (pooled 2024–2026), corresponding effects ≈ 0.3–0.4 pp and not FDR-significant.
- Interpreted as human umpires shifting the effective decision boundary to avoid walks (3–0) or to avoid ending plate appearances (two-strike).
- Game-progression effects (inning, late/close situations) exist but are smaller and structured: human calls became more strike-prone in later innings and especially in late-close situations; ABS seasons show a flat profile.
- Player-status (reputation) effects are suggestive but proxy-sensitive:
- Salary used as a reputation proxy; results depend on proxy validity.
- Example sensitivity: excluding a high-reputation but low-salary outlier (Shin-soo Choo) increased the negative batter slope (higher-salary batters received fewer strikes) to a statistically supported value.
- Pitcher top-vs-bottom salary contrasts in the human period showed positive effects (+1.62 to +2.30 pp at several cutoffs); attenuate under ABS.
- Catcher identity associated with residual called-strike heterogeneity in the human era; much of that heterogeneity disappears under ABS.
- Home-context (home-team advantage in calls) is mostly null at the umpire level; one human-period significant exception and an exploratory umpire–team gap were identified but treated as leads rather than robust findings.
- Overall interpretation: automation serves as a diagnostic benchmark. Patterns that attenuate under ABS are consistent with context-sensitive human judgment; persistent patterns may reflect pitch-selection, strategic change, measurement error, or model misspecification. The paper treats the ABS transition as diagnostic, not a randomized causal experiment.
Data & Methods
- Data:
- Pitch-level KBO data from Naver Sports, seasons 2021–partial 2026.
- Total rows: 1,216,246 pitches; called-pitch subset: 663,901; called strikes: 214,736.
- Primary human baseline: 2022–2023. ABS benchmark: 2024–2026 (pooled); some analyses use 2024–2025 as complete ABS seasons, 2026 partial as sensitivity.
- Focus sample: "boundary-band" taken pitches within 0.25 ft of the nearest rule-zone boundary (N reported per analysis).
- Outcome:
- CalledStrike indicator for taken pitches only (swings, fouls, balls in play, HBP, pitchouts excluded).
- Model:
- Logistic regression predicting called-strike probability with an explicit local strike-zone surface term:
- fzone is a parsimonious basis of location and geometry: [xi, yi, xi^2, yi^2, xiyi, di, d_i^2, ri, bi, hi] where
- xi, yi = normalized horizontal and vertical location,
- di = signed distance to nearest rule-zone boundary,
- ri = inside/outside indicator relative to rule-zone,
- bi, hi = batter-specific zone bottom and height.
- Context variables (Ci) tested include count state, salary-percentile proxies, inning/late-close flags, catcher/pitcher identity, home context.
- Controls (Xi): pitch-type group, standardized pitch speed, batter stance, pitcher hand; models include season fixed effects δs(i).
- Estimands reported as average marginal effects (percentage points) with confidence intervals and false-discovery-rate (FDR) adjusted q-values for families of tests.
- Robustness & sensitivity:
- FDR correction across figure-level test families.
- Salary-proxy sensitivity checks (e.g., excluding clear proxy mismatches like Shin-soo Choo).
- Multiple cutoffs for top-vs-bottom salary contrasts.
- ABS-period comparisons used diagnostically; authors emphasize non-randomized nature of the policy change.
Implications for AI Economics
- Automation as audit infrastructure: When algorithmic systems replace human decision-makers, they not only perform the task but create a natural benchmark to audit where human judgments were context-sensitive. Economists and policy-makers can exploit pre/post automation transitions to map human discretionary behavior without relying solely on cross-sectional human-call variation.
- Identifying where human discretion matters most: Focusing analysis on the “boundary” (cases where small contextual shifts matter most) is an efficient strategy for detecting human-versus-automated divergences. In AI economics, similar boundary-focused designs (near-decision-threshold observations) can isolate discretionary effects.
- Diagnostics vs causal claims: Automated attenuation of an effect is suggestive but not definitive proof the effect was caused by human bias—automation may alter strategic behavior or measurement. Economists should combine automation benchmarks with careful controls and sensitivity checks and be explicit when evidence is diagnostic rather than causal.
- Proxy validity is crucial: The salary–reputation exercise illustrates how measurement error in status proxies can flip inference. In AI auditing and economics, validate proxies, test for influential misclassifications, and present robustness exercises (exclude identifiable proxy-mismatch cases).
- Multiple-testing and structured inference: The paper’s use of FDR correction across related tests is a practical approach that AI-economics researchers should emulate when auditing many potential contextual effects.
- Policy and implementation lessons:
- Automation can reduce some context-driven discretionary biases (here, count pressure and catcher-associated heterogeneity).
- Policymakers considering automation for rule-bound decisions should expect concentrated gains where human discretion is most consequential (near boundaries), and should plan audits that exploit the automated benchmark to monitor residual biases post-adoption.
- Methodological template: The combination of:
- high-frequency microdata,
- geometry-aware local modeling of decision surfaces,
- pre/post automation comparisons,
- explicit sensitivity to proxies and multiple testing, offers a replicable framework for auditing human vs automated decision-making in other domains (credit underwriting thresholds, policing stops near legal thresholds, clinical triage near treatment cutoffs).
Limitations to note for economists translating the approach: - Non-randomized adoption: ABS rollout is observational; strategic behavior and league-wide changes can confound interpretations. - Partial data and season heterogeneity: 2026 was partial; vertical zone geometry differs by season and was controlled for but may introduce complexity. - Proxy measurement error: Reputation is latent; salary is an imperfect proxy with identifiable mismatches. - External validity: KBO institutional details (e.g., umpire training, cultural practices) may limit direct generalization to other contexts.
Overall, the paper demonstrates how automation can serve as a practical audit instrument to reveal where human discretionary judgments created systematic, context-dependent deviations from a rule-bound benchmark — a lesson with direct relevance for AI economics research and policy.
Assessment
Claims (10)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| In human-umpire games, borderline pitches in 3–0 counts were 6.61 percentage points more likely to be called strikes than pitches in 0–0 counts. Decision Quality | positive | Called-strike probability |
Reading fidelity
high
Study strength
high
|
n=75973
6.61 percentage points
|
| In human-umpire games, borderline pitches in two-strike counts were substantially less likely to be called strikes than pitches in 0–0 counts; the 0–2 effect was −17.17 percentage points. Decision Quality | negative | Called-strike probability |
Reading fidelity
high
Study strength
high
|
n=75973
-17.17 percentage points for 0–2 counts
|
| The human-period count pattern is largely absent under ABS: in pooled 2024–2026 ABS data, the 3–0 and 0–2 effects were only +0.37 and +0.33 percentage points, respectively, and neither survived false-discovery-rate correction. Decision Quality | null_result | Count-associated change in called-strike probability |
Reading fidelity
high
Study strength
high
|
n=94054
+0.37 percentage points for 3–0; +0.33 percentage points for 0–2
|
| The paper finds strong support for count balancing by human umpires: borderline pitches were more likely to become strikes when the alternative was ball four and less likely to become strikes when the call could end the plate appearance. Decision Quality | mixed | Context-dependent called-strike probability |
Reading fidelity
high
Study strength
high
|
n=75973
3–0: +6.61 percentage points; 0–2: −17.17 percentage points
|
| After excluding Shin-soo Choo, the human-period salary slope for batters was −3.82 percentage points, consistent with higher-salary batters receiving fewer called strikes on comparable borderline pitches. Decision Quality | negative | Residual called-strike effect for batters |
Reading fidelity
high
Study strength
medium
|
-3.82 percentage points, 95% CI -6.25 to -1.39
|
| Human-period top-versus-bottom salary contrasts for pitchers were positive across several salary cutoffs, with FDR-supported effects at the 20%, 30%, 35%, and 40% thresholds ranging from +1.62 to +2.30 percentage points. Decision Quality | positive | Called-strike probability for pitchers' pitches |
Reading fidelity
high
Study strength
medium
|
+1.62 to +2.30 percentage points
|
| The human-period batter salary association is not statistically supported when all players are included: the estimated slope is −1.50 percentage points with p = 0.496. Decision Quality | null_result | Residual called-strike effect associated with batter salary |
Reading fidelity
high
Study strength
medium
|
-1.50 percentage points, p = 0.496
|
| Game progression showed a smaller but coherent human-period pattern: calls were less strike-prone in early innings and more strike-prone in innings 7–9 or later, especially in late-close situations; complete ABS seasons were essentially flat. Decision Quality | mixed | Called-strike probability by inning and game situation |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Catcher identity was associated with residual called-strike variation during the human-umpire period, but this variation largely disappeared under ABS. Decision Quality | null_result | Catcher-associated residual variation in called-strike probability |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Average home-team advantage in ball-strike calls was mostly null at the umpire level, although the paper reports one FDR-significant human-period exception and an exploratory umpire–team gap as an audit lead. Decision Quality | null_result | Home-context effect on called-strike probability |
Reading fidelity
high
Study strength
low
|
not reported
|