The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Automated umpiring in the KBO reveals where human judgment shifted the strike zone: marginal pitches were substantially more likely to be called strikes on 3–0 counts and far less likely on two-strike counts under human umpires, patterns that largely vanished after the league adopted the Automated Ball-Strike system.

Auditing Contextual Bias in Human Ball-Strike Calls Using KBO's Automated Umpiring Transition
Kichang Lee, JeongGil Ko · September 03, 2026
arxiv quasi_experimental high evidence 7/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

Arxiv

Latest observation:

  1. Kichang Lee unresolved corpus identity
  2. JeongGil Ko unresolved corpus identity
Using the KBO's move to automated ball-strike adjudication as a diagnostic benchmark, the paper shows that human umpires exhibited strong count-dependent biases on borderline taken pitches (more strikes on 3–0, far fewer on two-strike counts), and those patterns largely attenuated under ABS.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

This paper uses the Korean Baseball Organization's adoption of the Automated Ball-Strike (ABS) system to audit long-standing claims about contextual bias in human ball-strike calls. Using pitch-level KBO data from 2021 through the available portion of the 2026 season, we model called-strike probability for taken pitches near the strike-zone boundary, with 2022-2023 as the primary human-umpire baseline and ABS seasons (2024 and onward) as a diagnostic benchmark. The strongest evidence concerns count pressure. Relative to 0--0 counts, human umpires called substantially fewer strikes in two-strike counts and more strikes in hitter-ahead three-ball counts. Specifically, in the main 0.25-ft boundary band, 0--2 was associated with a -17.17 percentage-point effect and 3--0 with a +6.61 percentage-point effect. Under ABS, the corresponding effects were close to zero and did not survive false-discovery-rate correction. Game progression shows a smaller but coherent pattern as human calls were less strike-prone in early innings and more strike-prone in innings 7--9+, especially in late-close situations, while complete ABS seasons were essentially flat. Other suspected biases are weaker or more localized. Salary-based reputation proxies provide suggestive but proxy-sensitive evidence, and catcher identity shows human-period residual heterogeneity that disappears under ABS. Home-context evidence is mostly null at the umpire level, with one FDR-significant human-period exception and an exploratory umpire-team gap best treated as an audit lead. Overall, the results do not show that human umpires were biased everywhere. Instead, they map where the human strike zone was most context-sensitive, where evidence was weaker, and where common suspicions received little support.

Summary

Main Finding

Human KBO umpires exhibited context-sensitive called-strike behavior concentrated near the rule-zone boundary, most clearly along count pressure: compared with 0–0, two-strike counts were substantially less likely to produce called strikes and 3–0 counts were substantially more likely. These count-dependent shifts largely disappear under the Automated Ball-Strike (ABS) system (2024+), consistent with those patterns being features of human judgment rather than pitch-selection or other mechanical factors. Other suspected biases (player reputation, game-progression, catcher effects, home advantage) show weaker, more localized, or proxy-sensitive evidence; some residual heterogeneity in the human period (e.g., by catcher) disappears under ABS.

Key Points

  • Strongest and clearest result: count pressure.
    • In the 0.25-ft boundary band (borderline pitches), human-period effects vs. 0–0:
    • 0–2: −17.17 percentage points (pp)
    • 3–0: +6.61 pp
    • Under ABS (pooled 2024–2026), corresponding effects ≈ 0.3–0.4 pp and not FDR-significant.
    • Interpreted as human umpires shifting the effective decision boundary to avoid walks (3–0) or to avoid ending plate appearances (two-strike).
  • Game-progression effects (inning, late/close situations) exist but are smaller and structured: human calls became more strike-prone in later innings and especially in late-close situations; ABS seasons show a flat profile.
  • Player-status (reputation) effects are suggestive but proxy-sensitive:
    • Salary used as a reputation proxy; results depend on proxy validity.
    • Example sensitivity: excluding a high-reputation but low-salary outlier (Shin-soo Choo) increased the negative batter slope (higher-salary batters received fewer strikes) to a statistically supported value.
    • Pitcher top-vs-bottom salary contrasts in the human period showed positive effects (+1.62 to +2.30 pp at several cutoffs); attenuate under ABS.
  • Catcher identity associated with residual called-strike heterogeneity in the human era; much of that heterogeneity disappears under ABS.
  • Home-context (home-team advantage in calls) is mostly null at the umpire level; one human-period significant exception and an exploratory umpire–team gap were identified but treated as leads rather than robust findings.
  • Overall interpretation: automation serves as a diagnostic benchmark. Patterns that attenuate under ABS are consistent with context-sensitive human judgment; persistent patterns may reflect pitch-selection, strategic change, measurement error, or model misspecification. The paper treats the ABS transition as diagnostic, not a randomized causal experiment.

Data & Methods

  • Data:
    • Pitch-level KBO data from Naver Sports, seasons 2021–partial 2026.
    • Total rows: 1,216,246 pitches; called-pitch subset: 663,901; called strikes: 214,736.
    • Primary human baseline: 2022–2023. ABS benchmark: 2024–2026 (pooled); some analyses use 2024–2025 as complete ABS seasons, 2026 partial as sensitivity.
    • Focus sample: "boundary-band" taken pitches within 0.25 ft of the nearest rule-zone boundary (N reported per analysis).
  • Outcome:
    • CalledStrike indicator for taken pitches only (swings, fouls, balls in play, HBP, pitchouts excluded).
  • Model:
    • Logistic regression predicting called-strike probability with an explicit local strike-zone surface term:
    • fzone is a parsimonious basis of location and geometry: [xi, yi, xi^2, yi^2, xiyi, di, d_i^2, ri, bi, hi] where
    • xi, yi = normalized horizontal and vertical location,
    • di = signed distance to nearest rule-zone boundary,
    • ri = inside/outside indicator relative to rule-zone,
    • bi, hi = batter-specific zone bottom and height.
    • Context variables (Ci) tested include count state, salary-percentile proxies, inning/late-close flags, catcher/pitcher identity, home context.
    • Controls (Xi): pitch-type group, standardized pitch speed, batter stance, pitcher hand; models include season fixed effects δs(i).
    • Estimands reported as average marginal effects (percentage points) with confidence intervals and false-discovery-rate (FDR) adjusted q-values for families of tests.
  • Robustness & sensitivity:
    • FDR correction across figure-level test families.
    • Salary-proxy sensitivity checks (e.g., excluding clear proxy mismatches like Shin-soo Choo).
    • Multiple cutoffs for top-vs-bottom salary contrasts.
    • ABS-period comparisons used diagnostically; authors emphasize non-randomized nature of the policy change.

Implications for AI Economics

  • Automation as audit infrastructure: When algorithmic systems replace human decision-makers, they not only perform the task but create a natural benchmark to audit where human judgments were context-sensitive. Economists and policy-makers can exploit pre/post automation transitions to map human discretionary behavior without relying solely on cross-sectional human-call variation.
  • Identifying where human discretion matters most: Focusing analysis on the “boundary” (cases where small contextual shifts matter most) is an efficient strategy for detecting human-versus-automated divergences. In AI economics, similar boundary-focused designs (near-decision-threshold observations) can isolate discretionary effects.
  • Diagnostics vs causal claims: Automated attenuation of an effect is suggestive but not definitive proof the effect was caused by human bias—automation may alter strategic behavior or measurement. Economists should combine automation benchmarks with careful controls and sensitivity checks and be explicit when evidence is diagnostic rather than causal.
  • Proxy validity is crucial: The salary–reputation exercise illustrates how measurement error in status proxies can flip inference. In AI auditing and economics, validate proxies, test for influential misclassifications, and present robustness exercises (exclude identifiable proxy-mismatch cases).
  • Multiple-testing and structured inference: The paper’s use of FDR correction across related tests is a practical approach that AI-economics researchers should emulate when auditing many potential contextual effects.
  • Policy and implementation lessons:
    • Automation can reduce some context-driven discretionary biases (here, count pressure and catcher-associated heterogeneity).
    • Policymakers considering automation for rule-bound decisions should expect concentrated gains where human discretion is most consequential (near boundaries), and should plan audits that exploit the automated benchmark to monitor residual biases post-adoption.
  • Methodological template: The combination of:
    • high-frequency microdata,
    • geometry-aware local modeling of decision surfaces,
    • pre/post automation comparisons,
    • explicit sensitivity to proxies and multiple testing, offers a replicable framework for auditing human vs automated decision-making in other domains (credit underwriting thresholds, policing stops near legal thresholds, clinical triage near treatment cutoffs).

Limitations to note for economists translating the approach: - Non-randomized adoption: ABS rollout is observational; strategic behavior and league-wide changes can confound interpretations. - Partial data and season heterogeneity: 2026 was partial; vertical zone geometry differs by season and was controlled for but may introduce complexity. - Proxy measurement error: Reputation is latent; salary is an imperfect proxy with identifiable mismatches. - External validity: KBO institutional details (e.g., umpire training, cultural practices) may limit direct generalization to other contexts.

Overall, the paper demonstrates how automation can serve as a practical audit instrument to reveal where human discretionary judgments created systematic, context-dependent deviations from a rule-bound benchmark — a lesson with direct relevance for AI economics research and policy.

Assessment

Paper Typequasi_experimental Evidence Strengthhigh — Large pitch-level sample (~1.2M pitches; ~210k called strikes), focused boundary-band where human/ABS differences concentrate, careful geometric controls (signed distance to rule boundary, batter-specific zone), rich covariates, season fixed effects, multiple robustness checks (player-level sensitivity, ABS pooled benchmark, FDR correction); main count-pressure finding is consistent, sizable, and attenuates under ABS, making the diagnostic contrast convincing for the documented biases despite non-random adoption. Methods Rigorhigh — Appropriate local modeling of the strike surface (polynomial basis + signed distance to boundary), models condition on fine-grained location and pitch/matchup covariates, pre/post benchmark using ABS, multiple hypothesis correction, sample restrictions (taken pitches, boundary band) that target the estimand; authors acknowledge non-random transition and proxy measurement limits (salary for reputation) and run sensitivity analyses. SampleKBO pitch-level data from publicly available Naver Sports records covering 2021–partial 2026 (1,216,246 pitch rows; 663,901 called pitches; 214,736 called strikes; 3,987 games). Primary human baseline: 2022–2023; ABS benchmark: pooled 2024–2026 (with 2024–2025 as complete-ABS primary for some analyses). Main analytical sample: taken pitches within 0.25 ft of nearest rule-zone boundary after strike-zone validation; excludes swings/fouls/balls in play and other non-taken adjudications. Themeshuman_ai_collab adoption IdentificationBefore–after comparison using the KBO's transition from human umpires (2022–2023 baseline) to Automated Ball-Strike (ABS) adjudication (2024–2026) as a diagnostic benchmark; focuses on taken pitches within 0.25 ft of the rule-zone boundary and estimates logistic models that control for pitch geometry (signed distance to boundary, batter-specific zone height/bottom), pitch/matchup covariates, season fixed effects, and uses false-discovery-rate correction to evaluate multiple tests—attenuation of human-period effects under ABS is interpreted as evidence that human calls were context-sensitive. GeneralizabilityFindings are specific to the KBO and its implementation of ABS; umpire behavior and league culture may differ in MLB/NPB/other leagues., ABS adoption was not randomized and could coincide with other league changes (rule/technology/strategic shifts) that affect pitch selection or batter/pitcher behavior., Analyses focus on taken borderline pitches (0.25 ft band) and may not generalize to far-inside or far-outside pitches or to swinging decisions., Reputation/status inference relies on salary as an imperfect proxy, creating measurement error for player-status effects., Partial 2026 season data and pooled ABS years may mask time-varying responses to automation during rollout.

Claims (10)

ClaimDirectionOutcomeConfidence & EvidenceDetails
In human-umpire games, borderline pitches in 3–0 counts were 6.61 percentage points more likely to be called strikes than pitches in 0–0 counts. Decision Quality positive Called-strike probability
Reading fidelity high
Study strength high
n=75973
6.61 percentage points
0.8
In human-umpire games, borderline pitches in two-strike counts were substantially less likely to be called strikes than pitches in 0–0 counts; the 0–2 effect was −17.17 percentage points. Decision Quality negative Called-strike probability
Reading fidelity high
Study strength high
n=75973
-17.17 percentage points for 0–2 counts
0.8
The human-period count pattern is largely absent under ABS: in pooled 2024–2026 ABS data, the 3–0 and 0–2 effects were only +0.37 and +0.33 percentage points, respectively, and neither survived false-discovery-rate correction. Decision Quality null_result Count-associated change in called-strike probability
Reading fidelity high
Study strength high
n=94054
+0.37 percentage points for 3–0; +0.33 percentage points for 0–2
0.8
The paper finds strong support for count balancing by human umpires: borderline pitches were more likely to become strikes when the alternative was ball four and less likely to become strikes when the call could end the plate appearance. Decision Quality mixed Context-dependent called-strike probability
Reading fidelity high
Study strength high
n=75973
3–0: +6.61 percentage points; 0–2: −17.17 percentage points
0.8
After excluding Shin-soo Choo, the human-period salary slope for batters was −3.82 percentage points, consistent with higher-salary batters receiving fewer called strikes on comparable borderline pitches. Decision Quality negative Residual called-strike effect for batters
Reading fidelity high
Study strength medium
-3.82 percentage points, 95% CI -6.25 to -1.39
0.48
Human-period top-versus-bottom salary contrasts for pitchers were positive across several salary cutoffs, with FDR-supported effects at the 20%, 30%, 35%, and 40% thresholds ranging from +1.62 to +2.30 percentage points. Decision Quality positive Called-strike probability for pitchers' pitches
Reading fidelity high
Study strength medium
+1.62 to +2.30 percentage points
0.48
The human-period batter salary association is not statistically supported when all players are included: the estimated slope is −1.50 percentage points with p = 0.496. Decision Quality null_result Residual called-strike effect associated with batter salary
Reading fidelity high
Study strength medium
-1.50 percentage points, p = 0.496
0.48
Game progression showed a smaller but coherent human-period pattern: calls were less strike-prone in early innings and more strike-prone in innings 7–9 or later, especially in late-close situations; complete ABS seasons were essentially flat. Decision Quality mixed Called-strike probability by inning and game situation
Reading fidelity high
Study strength medium
not reported
0.48
Catcher identity was associated with residual called-strike variation during the human-umpire period, but this variation largely disappeared under ABS. Decision Quality null_result Catcher-associated residual variation in called-strike probability
Reading fidelity high
Study strength medium
not reported
0.48
Average home-team advantage in ball-strike calls was mostly null at the umpire level, although the paper reports one FDR-significant human-period exception and an exploratory umpire–team gap as an audit lead. Decision Quality null_result Home-context effect on called-strike probability
Reading fidelity high
Study strength low
not reported
0.24

Notes