The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

AI writing tools surged among college applicants in 2024, with the largest uptake among fee-waiver (lower-SES) students; yet greater estimated LLM use is linked to larger declines in admission likelihood for lower-SES applicants, raising equity concerns about essay-based selection.

The Digital Divide in Generative AI: Evidence from Large Language Model Use in College Admissions Essays
Jinsook Lee, Conrad Borchers, AJ Alvero, Thorsten Joachims, Rene F. Kizilcec · February 19, 2026
arxiv correlational medium evidence 7/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

Arxiv

Latest observation:

  1. Jinsook Lee unresolved corpus identity
  2. Conrad Borchers unresolved corpus identity
  3. AJ Alvero unresolved corpus identity
  4. Thorsten Joachims unresolved corpus identity
  5. Rene F. Kizilcec unresolved corpus identity

Semantic Scholar

Latest observation:

  1. Jinsook Lee provider ID
  2. Conrad Borchers provider ID
  3. A. Alvero provider ID
  4. Thorsten Joachims provider ID
  5. René F. Kizilcec provider ID
Analyzing 81,663 applications (2020–2024) with an LLM-use detector, the authors find LLM adoption rose sharply in 2024—particularly among fee-waived (lower-SES) applicants—but higher estimated LLM use is more strongly associated with declines in predicted or actual admission probability for lower-SES students than for higher-SES students.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

Large language models (LLMs) have become popular writing tools among students and may expand access to high-quality feedback for students with less access to traditional writing support. At the same time, LLMs may standardize student voice or invite overreliance. This study examines how adoption of LLM-assisted writing varies across socioeconomic groups and how it relates to outcomes in a high-stakes context: U.S. college admissions. We analyze a de-identified longitudinal dataset of applications to a selective university from 2020 to 2024 (N = 81,663). Estimating LLM use using a distribution-based detector trained on synthetic and historical essays, we tracked how student writing changed as LLM use proliferated, how adoption differed by socioeconomic status (SES), and whether potential benefits translated equitably into admissions outcomes. Using fee-waiver status as a proxy for SES, we observe post-2023 convergence in surface-level linguistic features, with the largest changes in fee-waived and rejected applicants. Estimated LLM use rose sharply in 2024 across all groups, with disproportionately larger increases among lower SES applicants, consistent with an access hypothesis in which LLMs substitute for scarce writing support. However, increased estimated LLM use was more strongly associated with declines in predicted admission probability for lower SES applicants than for higher SES applicants, even after controlling for academic credentials and stylometric features. These findings raise concerns about equity and the validity of essay-based evaluation in an era of AI-assisted writing and provide the first large-scale longitudinal evidence linking LLM adoption, linguistic change, and evaluative outcomes in college admissions.

Summary

Main Finding

Estimated use of large language models (LLMs) in college admissions essays rose sharply in the 2024 cycle and was disproportionately higher among lower‑SES (fee‑waiver) applicants, consistent with LLMs substituting for scarce writing support. However, increased LLM‑like signatures in essays are more strongly associated with declines in predicted admission probability for lower‑SES applicants than for higher‑SES applicants — even after adjusting for academic credentials and observable stylometric features. The result is a new form of digital‑divide: broad access but unequal translation of AI assistance into favorable evaluative outcomes.

Key Points

  • Dataset: N = 81,663 de‑identified Common Application essays to a selective U.S. university across five admission cycles (2019–20 through 2023–24). Post‑GPT (2024) subsample N ≈ 17,654.
  • SES proxy: fee‑waiver status (binary indicator for documented financial need).
  • Linguistic changes: post‑2023 convergence on surface linguistic features (lexical diversity, complexity metrics), with the largest shifts among fee‑waived and rejected applicants.
  • LLM adoption: estimated LLM‑signature (ˆα) rose sharply in 2024 across all SES groups and increased relatively more for lower‑SES applicants.
  • Outcome association: higher ˆα correlates with lower predicted admission probability more strongly for lower‑SES students than for higher‑SES students.
  • Robustness & caveats: authors validate the detector’s calibration on synthetic data and run DiD, stratified, interaction, and mediation (change‑in‑coefficient with stylometrics) analyses. They caution results are associative (DiD assumptions partially violated due to concurrent policy/time shocks such as test‑optional changes).
  • Estimation categories: essays classified into no (ˆα = 0), low (0 < ˆα ≤ 0.07), medium (0.07 < ˆα ≤ 0.13), high (ˆα > 0.13) LLM‑use terciles for post‑GPT analyses.

Data & Methods

  • Corpus: 29,232 pre‑GPT essays used to build human reference (after excluding essays <250 words); synthetic LLM essays generated with GPT‑4o to match prompt and demographic distributions of human training data.
  • LLM usage estimator: distributional mixture approach adapted to the essay level (extension of the GPT quantification method). Each essay is scored with a continuous ˆα ∈ [0,1] measuring similarity to the LLM reference distribution versus the human reference distribution.
    • Validation: calibration plots (Appendix) show strong alignment between predicted and ground‑truth α on synthetic tests; authors emphasize ˆα is a relative, aggregate measure not proof of individual usage.
  • Linguistic features: multiple lexical diversity and complexity metrics (TTR, Maas TTR, MTLD, HDD, Yule’s K, average word length, and 1 − Flesch Reading Ease as complexity).
  • Outcomes modeling:
    • Difference‑in‑differences (DiD) comparing pre‑ and post‑GPT admission probabilities by fee‑waiver status, controlling for GPA, SAT/ACT, honors, demographics, school type.
    • Post‑GPT stratified logistic regressions (separate models by SES) and interaction models (FeeWaiver × ˆα) to test differential associations.
    • Mediation-style analysis (change‑in‑coefficient) adding 11 stylometric controls to assess how much observable writing changes explain the SES differential.
  • Sensitivity checks: event‑study, placebo timing, rolling windows, donut‑hole analysis, COVID interaction tests, and covariate stability diagnostics.

Implications for AI Economics

  • Access vs. effective use: LLMs reduce a traditional access barrier (free, widely available writing assistance), but benefits depend on users’ ability to elicit, integrate, and strategically use AI outputs. Economic inequality may therefore shift from access to differential effective use / skill in prompt engineering and editing.
  • Substitution and complementarities: LLMs can act as low‑cost substitutes for paid essay coaching or school‑based writing support. For disadvantaged students, substitution may increase uptake but not necessarily improve evaluative outcomes due to downstream interpretation by human evaluators.
  • Signaling and market for credentials: Essays historically function as a signal of non‑cognitive traits and fit. Widespread LLM use that standardizes surface features undermines the signal value of essays; markets (colleges, employers) may adapt (new signals, verification, or premium on other observables), changing returns to investments in human versus AI‑mediated writing skills.
  • Evaluation and discrimination risk: Admissions readers may interpret highly polished or “AI‑like” essays differently depending on applicant background, potentially penalizing lower‑SES applicants who use LLMs. This creates a form of evaluative discrimination linked to perceptions of authenticity and fit, with implications for fairness and redistribution.
  • Policy and institutional responses:
    • Admissions: reconsider essay role, adopt clearer guidance on acceptable AI use, train readers to account for AI‑assisted writing, or redesign evaluation rubrics to emphasize verifiable signals.
    • Education policy: invest in equitable instruction on prompt engineering, critical engagement with AI tools, and writing pedagogy to equalize effective use.
    • Market responses: expect growth in prompt‑coaching services and new intermediaries; monitoring and regulation of such markets may be warranted to avoid exacerbating inequality.
  • Measurement & empirical research: researchers and economists must treat written artifacts as potentially AI‑mediated signals. New measurement tools (like the paper’s distributional estimator) are needed but have limits; econometric analyses of ability, effort, and returns must consider AI distortion of observable inputs.
  • Normative tradeoffs: while LLMs can democratize access to high‑quality drafting support (welfare gains), they can simultaneously exacerbate outcome inequalities if institutions interpret AI‑assisted outputs heterogeneously. Policy design should aim to capture net welfare effects and mitigate harms to disadvantaged groups.

Limitations to keep in mind: ˆα is a distributional proxy (not ground truth); DiD causal claims are limited by contemporaneous shocks (e.g., test‑optional policies); findings are from a single selective institution and may not generalize across contexts or applicant pools.

Assessment

Paper Typecorrelational Evidence Strengthmedium — Large, multi-year dataset (N=81,663) and detailed text features provide statistical power and rich covariates, and the longitudinal frame captures timing of LLM adoption; however, LLM-use is inferred (measurement error/false positives possible), SES relies on an imperfect proxy (fee-waiver), and the design is observational so associations may reflect confounding or selection rather than causal effects. Methods Rigormedium — The study applies careful stylometric analysis, a trained detector for LLM use, and controls for academic credentials and other covariates, which is methodologically appropriate; but key risks remain (detector validity and calibration, nonrandom selection into LLM use, possible changes in admissions evaluation rules or reviewer behavior, and limited robustness checks described), limiting causal claims. SampleDe-identified longitudinal dataset of 81,663 undergraduate applications to a selective U.S. university from 2020–2024, including applicant essays, fee-waiver status (used as an SES proxy), academic credentials, stylometric/surface linguistic features, predicted admission probabilities, and admissions outcomes; LLM-use is estimated via a distribution-based detector trained on synthetic and historical essays. Themesinequality adoption skills_training IdentificationNo experimental or quasi-experimental source of exogenous variation; LLM use is estimated with a distribution-based detector trained on synthetic and historical essays, and the authors exploit longitudinal variation (2020–2024, with a sharp rise in 2024) and multivariate controls (academic credentials, stylometric features, predicted admission probability) to estimate associations between estimated LLM use, linguistic change, and admissions outcomes; SES is proxied by fee-waiver status. GeneralizabilitySingle selective U.S. university — findings may not generalize to less selective institutions or non-U.S. contexts, Applicants to a selective college are a nonrepresentative, advantaged sample compared with the general student population, Fee-waiver status is an imperfect proxy for socioeconomic status (omits family wealth, parental education, school resources), LLM-use detector may have measurement error or biases that vary by applicant subgroup (e.g., language background), affecting estimates, Admissions processes, reviewer practices, and essay prompts may differ across institutions and over time

Claims (8)

ClaimDirectionOutcomeConfidence & EvidenceDetails
We analyze a de-identified longitudinal dataset of applications to a selective university from 2020 to 2024 (N = 81,663). Other positive dataset size and scope (applications 2020–2024)
Reading fidelity high
Study strength high
n=81663
0.5
Estimated LLM use was measured using a distribution-based detector trained on synthetic and historical essays. Adoption Rate positive estimated LLM use (detector output)
Reading fidelity high
Study strength high
n=81663
0.5
We observe post-2023 convergence in surface-level linguistic features, with the largest changes in fee-waived and rejected applicants. Output Quality mixed surface-level linguistic/stylometric features of essays
Reading fidelity high
Study strength medium
n=81663
0.3
Estimated LLM use rose sharply in 2024 across all groups. Adoption Rate positive estimated LLM adoption rate by year
Reading fidelity high
Study strength medium
n=81663
0.3
Estimated LLM use increased disproportionately more among lower socioeconomic status (SES) applicants (fee-waived) in 2024, consistent with an access hypothesis in which LLMs substitute for scarce writing support. Adoption Rate positive relative increase in estimated LLM adoption by SES (fee-waiver status)
Reading fidelity high
Study strength speculative
n=81663
0.05
Increased estimated LLM use was more strongly associated with declines in predicted admission probability for lower SES applicants than for higher SES applicants, even after controlling for academic credentials and stylometric features. Hiring negative predicted admission probability (admissions outcome)
Reading fidelity high
Study strength medium
n=81663
0.3
These findings raise concerns about equity and the validity of essay-based evaluation in an era of AI-assisted writing. Inequality negative equity of admissions and validity of essay evaluation
Reading fidelity high
Study strength speculative
n=81663
0.05
This study provides the first large-scale longitudinal evidence linking LLM adoption, linguistic change, and evaluative outcomes in college admissions. Research Productivity positive existence of large-scale longitudinal evidence connecting LLM adoption, writing changes, and admissions outcomes
Reading fidelity medium
Study strength speculative
n=81663
0.03

Notes