The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Gender-equity tools that rely on names disproportionately reward women with Western-sounding, gender-legible names while bypassing women whose names lose gender cues in transliteration; observational analysis of citation-diversity statements and two preregistered experiments across cultures identify a persistent 'legibility gap' driven by linguistic legibility rather than evaluator familiarity.

The Legibility Gap: How Gender Equity Interventions Redistribute Recognition Across Cultures
Binglu Wang, Jose Cervantez, Jiahui Xue, Katherine L. Milkman, Dashun Wang · August 21, 2026
arxiv rct high evidence 7/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

Arxiv

Latest observation:

  1. Binglu Wang unresolved corpus identity
  2. Jose Cervantez unresolved corpus identity
  3. Jiahui Xue unresolved corpus identity
  4. Katherine L. Milkman unresolved corpus identity
  5. Dashun Wang unresolved corpus identity

Semantic Scholar

Latest observation:

  1. Binglu Wang provider ID
  2. Jose Cervantez provider ID
  3. Jiahui Xue provider ID
  4. Katherine L. Milkman provider ID
  5. Dashun Wang provider ID
Name-based gender-inference tools and citation-diversity interventions increase representation of women overall but disproportionately benefit women with Western, gender-legible names while leaving women whose names lose gender cues in transliteration (notably many East Asian names) behind, as shown in observational data and two preregistered experiments.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

Efforts to promote gender equity in science increasingly rely on name-based inference to quantify representation and guide policy and behavior. Yet linguistic cues that signal gender vary across cultures and are often obscured when names are transliterated into English. Here we identify a pattern we call the "legibility gap": when gender is inferred from names, equity interventions systematically benefit women whose names signal gender while bypassing those whose names lose such cues in translation. Using both observational and experimental evidence, we show how this gap reshapes recognition in science. Analyzing citation diversity statements-an emerging practice in which authors report the algorithmically estimated gender composition of their reference lists-we find that papers that include this practice cite women more frequently, but the gains accrue almost entirely to authors with gender-signaling Western names. By contrast, women whose names lose gender cues in English transliteration, predominantly those with East Asian names, receive fewer citations in these same papers. Two preregistered experiments (N = 2,250) corroborate this pattern and identify its mechanism: linguistic legibility, not cultural unfamiliarity, determines who is recognized as a woman and who benefits from policies designed to support women in science. Overall, these findings expose a previously unrecognized layer of inequity embedded in global equity infrastructures. As science becomes increasingly global and equity efforts increasingly algorithmic, the legibility gap reveals how uneven identity recognition reshapes fairness. In global systems of recognition, equity depends not only on whether policies are effective on average, but also on whether they are equitable across cultures.

Summary

Main Finding

When interventions and measurements rely on name-based gender inference, they systematically benefit people whose names retain clear gender cues in English (primarily Western names) while bypassing those whose names lose such cues in transliteration (notably many East Asian names). This “legibility gap” means policies that increase women’s representation on average can nonetheless redistribute recognition away from culturally or linguistically less legible women.

Key Points

  • Definition: The “legibility gap” is the unequal visibility and benefit created when gender is inferred from names that vary in how clearly they signal gender once transliterated into English.
  • Observational result (citation diversity statements, CDSs):
    • Papers that included CDSs cited women more (30.8% vs 23.8%; +7.0 percentage points).
    • The increase in citations to women disproportionately accrued to authors with English-origin/Western names; citations to Chinese, Japanese, Arab, and other non-Western name groups fell.
    • The apparent gain for women was offset by reductions in citations to gender-blind names (−8.6 percentage points), suggesting reallocation rather than uniform improvement.
  • Causal evidence (preregistered experiments):
    • Study 2 (US participants, N=750): Gender-feedback nudges raised selection of Western women significantly (10.8% → 21.1%; β = 0.1029, p < .001) but did not significantly increase selection of Eastern women (8.2% → 10.8%; p = .217).
    • Study 3 (participants from China, South Korea, Italy, Germany; final N ≈ 1,443): Replicated pattern—gender feedback increased selection of Western women (7.8% → 21.4%; β = 0.1360, p < .001) but not Eastern women (no significant change). The asymmetry persisted even when evaluators shared the cultural background of Eastern names.
    • Mechanism: results support linguistic legibility (names encoding gender cues surviving transliteration) rather than cultural unfamiliarity of evaluators.
  • Scope/limits: Focuses on name-based (binary) gender inference tools; does not address nonbinary identities; country-level West/East coarse-grained; potential downstream cumulative effects not measured here.

Data & Methods

  • Observational study (Study 1):
    • Dataset: Papers published 2020–2024 collected May 2024.
    • Sample: 216 papers with citation diversity statements (CDS) and 432 matched comparison papers (total N = 648 papers), together citing 40,719 references.
    • Matching: Coarsened Exact Matching (CEM) to obtain comparable papers by journal/characteristics.
    • Gender inference: gender-guesser library.
    • Cultural origin inference: Ethnea.
    • Analysis: Paper-level regressions controlling for publication year, field, region, team size, citation impact. Key coefficients show CDS association with increased woman-citation share and decreased share of gender-blind/non-Western names.
  • Experiments (Studies 2 & 3):
    • Pre-registered designs (links provided in paper).
    • Study 2: US sample N = 750; selection task for a real Facebook campaign; random assignment to control vs gender-feedback; measured final pick distribution by candidate type (Western/Eastern, man/woman).
    • Study 3: Cross-cultural sample (China, South Korea, Italy, Germany; ~1,500 recruited, ~1,443 completed); same design as Study 2 but with heterogeneous evaluator cultures; tested interaction between evaluator region and candidate origin.
    • Statistical tests: logistic/linear models, Wald tests for differences across groups; results robust to controls for participant demographics.
  • Robustness: Observational and causal evidence converge; experiments isolate treatment and mechanism (legibility vs familiarity).

Implications for AI Economics

  • Measurement and allocation are linked: When algorithmic systems or dashboard metrics use inferred demographics as inputs, measurement errors that are systematic by language/culture become allocation errors—resources, visibility, or opportunities (e.g., citations, invited talks, candidate shortlists) can be misdirected.
  • Fairness metrics can be misleading: Aggregate parity gains (e.g., more women cited) can mask heterogeneous effects across cultural or linguistic groups. Relying on a single inferred attribute without auditing subgroup effects risks reinforcing cross-cultural inequality.
  • Market and career impacts:
    • Citation counts and visibility feed hiring, promotions, funding, and collaboration markets. The legibility gap can exacerbate cumulative advantage for those already legible to measurement systems, amplifying inequality over time.
    • Global talent evaluation platforms and marketplaces that use automated demographic inference may undervalue whole segments of the global workforce.
  • Policy and design recommendations for AI economists and practitioners:
    • Do not use name-based inferred gender as an allocation signal without auditing for cross-cultural heterogeneity.
    • Where possible, collect self-reported demographic data (opt-in, privacy-protected) rather than inferring from names.
    • If inference is necessary, model uncertainty: propagate classifier confidence into downstream decisions and avoid hard thresholds that treat noisy labels as ground truth.
    • Develop multilingual, script-aware classifiers and incorporate native-script information (avoid relying only on English transliterations).
    • Disaggregate impact analyses by cultural/linguistic groups to detect legibility gaps; report subgroup effects, not only averages.
    • Adopt corrective weighting or targeted outreach for under-legible groups, and monitor long-term compounding effects.
    • Perform regular algorithmic audits that test for interpreter-general biases (i.e., whether a measurement approach favors signals that survive translation).
  • Research agenda:
    • Quantify downstream economic costs of misallocation due to legibility gaps (career trajectories, funding misdirection).
    • Extend analysis beyond binary gender to nonbinary identities and other axes (ethnicity, caste, religion) that may similarly lose signal in transliteration.
    • Design and evaluate scalable alternatives to name-based inference (e.g., federated self-identification, multilingual name registries, pronoun metadata in scholarly profiles).

In short: algorithmic, name-based gender inference can introduce culturally structured bias into measurement and allocation. AI economics work that evaluates interventions, designs equitable algorithms, or builds policy tools must account for legibility heterogeneity across languages and cultures to avoid inadvertent redistribution of opportunity.

Assessment

Paper Typerct Evidence Strengthhigh — The paper combines a large observational dataset (216 papers with CDSs, 432 matched controls, ~40,700 references) with two preregistered randomized experiments (total N≈2,250) that replicate the core asymmetry across samples and cultures, providing convergent correlational and causal evidence; limitations remain around measurement of gender/culture and ecological scope, but the causal claims are supported by RCTs. Methods Rigorhigh — Pre-registered experimental designs, large sample sizes, cross-cultural replication, and matched observational analyses with controls and CEM indicate strong design and internal validity; notable limitations include reliance on name-based gender/cultural inference algorithms (binary gender assumptions), possible misclassification of transliterated names, coarse cultural categories, and laboratory-style selection tasks that may differ from real-world academic decision processes. SampleObservational: Papers published 2020–2024 collected May 2024; 216 papers containing citation diversity statements matched (CEM) to 432 papers from the same journals, yielding ~40,719 cited references; gender inferred via 'gender-guesser' and cultural origin via Ethnea. Experiments: Study 2 — N=750 U.S. participants who selected scholars for a Facebook ad task (final N reported); Study 3 — recruited 1,500 participants (375 each from China, South Korea, Italy, Germany) with 1,443 completers; stimuli included candidate names from Western and Eastern countries; participants randomized to gender-feedback vs. control feedback; both experiments preregistered. Themesinequality governance IdentificationMixed strategy: observational quasi-experimental analysis using Coarsened Exact Matching and regression controls on papers with vs. without citation diversity statements; causal identification from two preregistered randomized experiments that randomly assigned gender-feedback interventions to participants (Study 2: US sample; Study 3: cross-cultural sample), isolating the causal effect of gender-feedback on selection outcomes. GeneralizabilityRelies on name-based gender inference tools with binary gender outputs; excludes nonbinary and culturally specific gender systems., Cultural origin labels (Western/Eastern, country categories) are coarse and may mask within-region variation and diasporic naming practices., Transliteration effects and algorithm misclassification may vary by language and are concentrated on East Asian names in this study; other language groups may differ., Experimental tasks (selecting candidates for an ad) are proximate measures of selection behavior and may not map perfectly onto high-stakes real-world outcomes like hiring, promotion, or long-term citation dynamics., Participant recruitment platforms and sample composition (online respondents) may limit external validity to broader academic decision-makers.

Claims (7)

ClaimDirectionOutcomeConfidence & EvidenceDetails
Papers that include citation diversity statements cite women more frequently than matched papers without citation diversity statements: 30.8% versus 23.8%, a 7.0 percentage-point difference. Inequality positive Share of cited authors identified as women
Reading fidelity high
Study strength medium
n=648
7.0 percentage points higher (30.8% vs. 23.8%); β=0.0709, p<0.001
0.6
Citation diversity statements are associated with a shift away from citations to gender-blind names, while citations to men’s names do not significantly differ between papers with and without the statements. Inequality mixed Citation shares for men’s names and gender-blind names
Reading fidelity high
Study strength medium
n=648
Men’s names: β=0.0108, p=0.297; gender-blind names: β=-0.0860, p<0.001
0.6
Papers with citation diversity statements cite more authors with English-origin names and fewer authors with Chinese, Japanese, and Arab names. Inequality mixed Cultural-origin composition of cited authors’ names
Reading fidelity high
Study strength medium
n=648
English-origin names: β=0.0606; Chinese names: β=-0.0316; Japanese names: β=-0.0101; Arab names: β=-0.0080; all p<0.001
0.6
Gender feedback causally increased selection of Western women for a promotional campaign, but did not significantly increase selection of Eastern women. Inequality mixed Whether the participant’s final selected scholar was a Western woman or an Eastern woman
Reading fidelity high
Study strength high
n=750
Western women: 21.1% vs. 10.8%, β=0.1029, p<0.001; Eastern women: 10.8% vs. 8.2%, β=0.0265, p=0.217
1.0
The gender-feedback intervention benefited Western women significantly more than Eastern women in the first preregistered experiment. Inequality positive Difference in treatment effects on selection of Western versus Eastern women
Reading fidelity high
Study strength high
n=750
χ²(1)=4.450, p=0.0349
1.0
In the second preregistered experiment, gender feedback substantially increased selection of Western women but did not significantly increase selection of Eastern women. Inequality mixed Whether the final selected scholar was a Western or Eastern woman
Reading fidelity high
Study strength high
n=1443
Western women: 7.8% to 21.4%, β=0.1360, p<0.001; Eastern women: β=0.0184, p=0.233
1.0
The unequal benefit of gender feedback persisted even when evaluators shared the broad cultural background of the women being evaluated, providing no evidence that cultural familiarity eliminates the legibility gap. Inequality null_result Selection of Western and Eastern women following gender feedback, by evaluator region
Reading fidelity high
Study strength high
n=1443
Western-woman interaction: β=0.0750, p=0.041; Eastern-woman interaction: β=0.0188, p=0.543; Wald test χ²(1)=1.204, p=0.272
1.0

Notes