3 cumulative citations
View corpus contextAI writing tools surged among college applicants in 2024, with the largest uptake among fee-waiver (lower-SES) students; yet greater estimated LLM use is linked to larger declines in admission likelihood for lower-SES applicants, raising equity concerns about essay-based selection.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
Large language models (LLMs) have become popular writing tools among students and may expand access to high-quality feedback for students with less access to traditional writing support. At the same time, LLMs may standardize student voice or invite overreliance. This study examines how adoption of LLM-assisted writing varies across socioeconomic groups and how it relates to outcomes in a high-stakes context: U.S. college admissions. We analyze a de-identified longitudinal dataset of applications to a selective university from 2020 to 2024 (N = 81,663). Estimating LLM use using a distribution-based detector trained on synthetic and historical essays, we tracked how student writing changed as LLM use proliferated, how adoption differed by socioeconomic status (SES), and whether potential benefits translated equitably into admissions outcomes. Using fee-waiver status as a proxy for SES, we observe post-2023 convergence in surface-level linguistic features, with the largest changes in fee-waived and rejected applicants. Estimated LLM use rose sharply in 2024 across all groups, with disproportionately larger increases among lower SES applicants, consistent with an access hypothesis in which LLMs substitute for scarce writing support. However, increased estimated LLM use was more strongly associated with declines in predicted admission probability for lower SES applicants than for higher SES applicants, even after controlling for academic credentials and stylometric features. These findings raise concerns about equity and the validity of essay-based evaluation in an era of AI-assisted writing and provide the first large-scale longitudinal evidence linking LLM adoption, linguistic change, and evaluative outcomes in college admissions.
Summary
Main Finding
Estimated use of large language models (LLMs) in college admissions essays rose sharply in the 2024 cycle and was disproportionately higher among lower‑SES (fee‑waiver) applicants, consistent with LLMs substituting for scarce writing support. However, increased LLM‑like signatures in essays are more strongly associated with declines in predicted admission probability for lower‑SES applicants than for higher‑SES applicants — even after adjusting for academic credentials and observable stylometric features. The result is a new form of digital‑divide: broad access but unequal translation of AI assistance into favorable evaluative outcomes.
Key Points
- Dataset: N = 81,663 de‑identified Common Application essays to a selective U.S. university across five admission cycles (2019–20 through 2023–24). Post‑GPT (2024) subsample N ≈ 17,654.
- SES proxy: fee‑waiver status (binary indicator for documented financial need).
- Linguistic changes: post‑2023 convergence on surface linguistic features (lexical diversity, complexity metrics), with the largest shifts among fee‑waived and rejected applicants.
- LLM adoption: estimated LLM‑signature (ˆα) rose sharply in 2024 across all SES groups and increased relatively more for lower‑SES applicants.
- Outcome association: higher ˆα correlates with lower predicted admission probability more strongly for lower‑SES students than for higher‑SES students.
- Robustness & caveats: authors validate the detector’s calibration on synthetic data and run DiD, stratified, interaction, and mediation (change‑in‑coefficient with stylometrics) analyses. They caution results are associative (DiD assumptions partially violated due to concurrent policy/time shocks such as test‑optional changes).
- Estimation categories: essays classified into no (ˆα = 0), low (0 < ˆα ≤ 0.07), medium (0.07 < ˆα ≤ 0.13), high (ˆα > 0.13) LLM‑use terciles for post‑GPT analyses.
Data & Methods
- Corpus: 29,232 pre‑GPT essays used to build human reference (after excluding essays <250 words); synthetic LLM essays generated with GPT‑4o to match prompt and demographic distributions of human training data.
- LLM usage estimator: distributional mixture approach adapted to the essay level (extension of the GPT quantification method). Each essay is scored with a continuous ˆα ∈ [0,1] measuring similarity to the LLM reference distribution versus the human reference distribution.
- Validation: calibration plots (Appendix) show strong alignment between predicted and ground‑truth α on synthetic tests; authors emphasize ˆα is a relative, aggregate measure not proof of individual usage.
- Linguistic features: multiple lexical diversity and complexity metrics (TTR, Maas TTR, MTLD, HDD, Yule’s K, average word length, and 1 − Flesch Reading Ease as complexity).
- Outcomes modeling:
- Difference‑in‑differences (DiD) comparing pre‑ and post‑GPT admission probabilities by fee‑waiver status, controlling for GPA, SAT/ACT, honors, demographics, school type.
- Post‑GPT stratified logistic regressions (separate models by SES) and interaction models (FeeWaiver × ˆα) to test differential associations.
- Mediation-style analysis (change‑in‑coefficient) adding 11 stylometric controls to assess how much observable writing changes explain the SES differential.
- Sensitivity checks: event‑study, placebo timing, rolling windows, donut‑hole analysis, COVID interaction tests, and covariate stability diagnostics.
Implications for AI Economics
- Access vs. effective use: LLMs reduce a traditional access barrier (free, widely available writing assistance), but benefits depend on users’ ability to elicit, integrate, and strategically use AI outputs. Economic inequality may therefore shift from access to differential effective use / skill in prompt engineering and editing.
- Substitution and complementarities: LLMs can act as low‑cost substitutes for paid essay coaching or school‑based writing support. For disadvantaged students, substitution may increase uptake but not necessarily improve evaluative outcomes due to downstream interpretation by human evaluators.
- Signaling and market for credentials: Essays historically function as a signal of non‑cognitive traits and fit. Widespread LLM use that standardizes surface features undermines the signal value of essays; markets (colleges, employers) may adapt (new signals, verification, or premium on other observables), changing returns to investments in human versus AI‑mediated writing skills.
- Evaluation and discrimination risk: Admissions readers may interpret highly polished or “AI‑like” essays differently depending on applicant background, potentially penalizing lower‑SES applicants who use LLMs. This creates a form of evaluative discrimination linked to perceptions of authenticity and fit, with implications for fairness and redistribution.
- Policy and institutional responses:
- Admissions: reconsider essay role, adopt clearer guidance on acceptable AI use, train readers to account for AI‑assisted writing, or redesign evaluation rubrics to emphasize verifiable signals.
- Education policy: invest in equitable instruction on prompt engineering, critical engagement with AI tools, and writing pedagogy to equalize effective use.
- Market responses: expect growth in prompt‑coaching services and new intermediaries; monitoring and regulation of such markets may be warranted to avoid exacerbating inequality.
- Measurement & empirical research: researchers and economists must treat written artifacts as potentially AI‑mediated signals. New measurement tools (like the paper’s distributional estimator) are needed but have limits; econometric analyses of ability, effort, and returns must consider AI distortion of observable inputs.
- Normative tradeoffs: while LLMs can democratize access to high‑quality drafting support (welfare gains), they can simultaneously exacerbate outcome inequalities if institutions interpret AI‑assisted outputs heterogeneously. Policy design should aim to capture net welfare effects and mitigate harms to disadvantaged groups.
Limitations to keep in mind: ˆα is a distributional proxy (not ground truth); DiD causal claims are limited by contemporaneous shocks (e.g., test‑optional policies); findings are from a single selective institution and may not generalize across contexts or applicant pools.
Assessment
Claims (8)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| We analyze a de-identified longitudinal dataset of applications to a selective university from 2020 to 2024 (N = 81,663). Other | positive | dataset size and scope (applications 2020–2024) |
Reading fidelity
high
Study strength
high
|
n=81663
|
| Estimated LLM use was measured using a distribution-based detector trained on synthetic and historical essays. Adoption Rate | positive | estimated LLM use (detector output) |
Reading fidelity
high
Study strength
high
|
n=81663
|
| We observe post-2023 convergence in surface-level linguistic features, with the largest changes in fee-waived and rejected applicants. Output Quality | mixed | surface-level linguistic/stylometric features of essays |
Reading fidelity
high
Study strength
medium
|
n=81663
|
| Estimated LLM use rose sharply in 2024 across all groups. Adoption Rate | positive | estimated LLM adoption rate by year |
Reading fidelity
high
Study strength
medium
|
n=81663
|
| Estimated LLM use increased disproportionately more among lower socioeconomic status (SES) applicants (fee-waived) in 2024, consistent with an access hypothesis in which LLMs substitute for scarce writing support. Adoption Rate | positive | relative increase in estimated LLM adoption by SES (fee-waiver status) |
Reading fidelity
high
Study strength
speculative
|
n=81663
|
| Increased estimated LLM use was more strongly associated with declines in predicted admission probability for lower SES applicants than for higher SES applicants, even after controlling for academic credentials and stylometric features. Hiring | negative | predicted admission probability (admissions outcome) |
Reading fidelity
high
Study strength
medium
|
n=81663
|
| These findings raise concerns about equity and the validity of essay-based evaluation in an era of AI-assisted writing. Inequality | negative | equity of admissions and validity of essay evaluation |
Reading fidelity
high
Study strength
speculative
|
n=81663
|
| This study provides the first large-scale longitudinal evidence linking LLM adoption, linguistic change, and evaluative outcomes in college admissions. Research Productivity | positive | existence of large-scale longitudinal evidence connecting LLM adoption, writing changes, and admissions outcomes |
Reading fidelity
medium
Study strength
speculative
|
n=81663
|