The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

AI-generated STEM instruction favors privileged profiles: audits of four large language models reveal up to a 2.55-grade-level advantage for the most privileged over the most marginalized, with income, medium of instruction and disability driving the largest single effects and biases persisting even within elite institutions.

Compounding Disadvantage: Auditing Intersectional Bias in LLM-Generated Explanations Across Indian and American STEM Education
Amogh Gupta, Niharika Patil, Sourojit Ghosh, SnehalKumar, S Gaikwad · January 20, 2026
arxiv quasi_experimental medium evidence 7/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

Arxiv

Latest observation:

  1. Amogh Gupta unresolved corpus identity
  2. Niharika Patil unresolved corpus identity
  3. Sourojit Ghosh unresolved corpus identity
  4. SnehalKumar unresolved corpus identity
  5. S Gaikwad unresolved corpus identity

Semantic Scholar

Latest observation:

  1. Amogh Gupta provider ID
  2. N. Patil provider ID
  3. Sourojit Ghosh provider ID
  4. Snehalkumar S. Gaikwad provider ID
Systematic audits of four LLMs show that AI-generated STEM content gives consistently higher-quality outputs to privileged synthetic student profiles than to marginalized profiles—producing gaps up to 2.55 grade levels driven mainly by income, medium of instruction (India), and disability status, with intersectional effects that compound across dimensions.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

Large language models are increasingly deployed in STEM education for personalized instruction and feedback across institutions in high- and low-income countries. These systems are designed to adapt content to student needs, but whether they adapt based on demonstrated ability or demographic signals remains untested at scale. Here we establish that LLM-generated STEM content systematically disadvantages marginalized student profiles across two cultural contexts, with the gap between the most privileged and most marginalized profiles reaching 2.55 grade levels. We audited four LLMs (Qwen 2.5-32B-Instruct, GPT-4o, GPT-4o-mini, GPT-OSS 20B) using synthetic profiles crossing dimensions specific to Indian education (caste, medium of instruction, college tier) and American education (race, HBCU attendance, school type), alongside income, gender, and disability, across ranking and generation tasks with FDR-corrected significance testing and SHAP feature attribution. Income produces significant effects across every model and context, medium of instruction drives the largest single effect in the Indian context, and disability status triggers simpler explanations. Effects compound non-additively: marginalization across multiple dimensions produces gaps larger than any single dimension predicts, and biases persist within elite institutions. Bias is consistent across all four architectures and persists through model selection, making intersectional, cross-cultural auditing a structural requirement before deployment.

Summary

Main Finding

LLM-generated STEM instructional content systematically disadvantages marginalized student profiles across both Indian and U.S. educational contexts. Income is a consistent and significant predictor of simpler treatment in every model and context; in India, medium of instruction (English vs. Hindi/regional) produces the largest single effect. Biases compound non‑additively across intersecting identities (caste/race, income, location, disability, institution tier), producing gaps up to 2.55 grade levels between the most privileged and most marginalized synthetic profiles. These patterns hold across four different LLMs (Qwen 2.5-32B‑Instruct, GPT‑4o, GPT‑4o‑mini, GPT‑OSS‑20B) and across selection/role framing, implying structural origins in training/alignment rather than idiosyncratic model choices.

Key Points

  • Audited models: Qwen 2.5‑32B‑Instruct, GPT‑4o, GPT‑4o‑mini, GPT‑OSS‑20B (temperature = 0).
  • Two contexts: India (caste, medium of instruction, college tier, board, location, income, gender, disability) and U.S. (race/ethnicity, HBCU attendance, school type, college tier, location, income, gender, disability).
  • Two tasks:
    • Ranking task → Mean Choice Value (MCV): model’s selected difficulty level (1–5).
    • Generation task → Mean Grade Level (MGL): average of Flesch‑Kincaid, Gunning Fog, Coleman‑Liau readability scores.
  • Datasets: MATH‑50 (general math benchmark) and JEEBench (India engineering exam problems).
  • Sample design: stratified sampling of 100 synthetic intersectional profiles per context (coverage across 5,184 Indian and 4,860 U.S. possible combinations).
  • Bias metrics: Mean Absolute Bias (MAB) and Maximum Difference Bias (MDB); SHAP used for feature attribution; statistical tests with FDR correction and Cohen’s d reported.
  • Quantitative highlights:
    • Intersectional gap (most privileged vs most marginalized): up to 2.55 grade levels.
    • Income effect sizes across conditions: Cohen’s d ≈ 0.21 to 0.81.
    • In India, English‑medium profiles received the highest complexity outputs (authors report 100% of highest‑complexity outputs assigned to English‑medium profiles in their tests).
    • Disability cues trigger systematically simpler explanations; effect nearly doubles when the model is asked to adopt a student (rather than teacher) perspective.
    • Bias persists inside elite institutions (e.g., low‑income / caste‑oppressed IIT profiles ≈ 0.9 grade levels below privileged IIT peers; similar penalties found for low‑income rural students at Ivy tier).
  • Consistency: Direction and broad magnitude of biases are comparable across all four model architectures and both tasks; variation mainly in consistency of application.

Data & Methods

  • Profile construction:
    • India: 8 attributes (caste: General/OBC/SC/ST; college tier: IIT/NIT/State Gov/Private; location: Metro/Tier‑2/Rural; medium: English/Hindi/regional; board: CBSE/ICSE/State; gender; income: High/Mid/Low; disability).
    • U.S.: 7 attributes (race/ethnicity categories; college tier including Ivy/HBCU etc.; location: Urban/Suburban/Rural; school type: Public/Private/Charter; gender; income; disability).
    • Stratified sampling produced 100 profiles per context to balance representation.
  • Tasks and samples:
    • Ranking experiments: 100 profiles × 7 subjects × 2 role frames (teacher/student) → 1,400 ranking trials reported.
    • Generation experiments: MATH‑50: 100 profiles × 7 subjects × 3 problems → 2,100 generations. JEEBench: 50 problems sampled → 5,000 generations reported.
  • Complexity measures:
    • MCV (ranking) and MGL (generation, average of three readability indices).
  • Bias quantification:
    • Normalize MCV/MGL within subgroup and subject.
    • MAB = average absolute deviation from subgroup mean; MDB = max difference between subgroup extremes.
    • SHAP applied to decompose dimension contributions to complexity differences.
    • Statistical validation: t‑tests, Cohen’s d, KL divergence; p‑values FDR‑corrected.
  • Models and settings:
    • Two closed models (GPT‑4o, GPT‑4o‑mini) and two open‑weight models (Qwen 2.5‑32B‑Instruct, GPT‑OSS‑20B).
    • Deterministic decoding (temperature = 0); served with float16 for open models.

Implications for AI Economics

  • Human capital formation and inequality
    • Differential instructional complexity translates into differential learning inputs. If marginalized students systematically receive simpler content, their skill accumulation and human capital growth may be reduced—potentially lowering lifetime earnings and perpetuating intergenerational inequality.
    • Non‑additive intersectional penalties (up to 2.55 grade levels) imply compounding losses that standard single‑axis evaluations underestimate; these can widen returns‑to‑education gaps across population groups.
  • Labor market and signaling effects
    • Reduced exposure to complex explanations could impair readiness for advanced STEM training, reducing the supply of qualified candidates from marginalized groups and reinforcing signaling advantages of privileged students/institutions.
    • Employers and institutions relying on LLM‑assisted preparation may inadvertently favor candidates from groups that received richer LLM tutoring.
  • Aggregate productivity and growth
    • If widely deployed LLM tutors systematically under‑educate marginalized cohorts, aggregate productivity gains from AI in education may be uneven, lowering aggregate social returns and potentially increasing macroeconomic inequality.
    • Economic growth models that incorporate human capital accumulation should account for AI‑mediated, distributional learning effects; neglecting them risks overstating egalitarian gains from AI in education.
  • Market and regulatory responses
    • Market: demand for localized/fairness‑tuned models (language, caste/race‑aware training) will rise; institutions may procure models audited for intersectional equity as a competitive or compliance signal.
    • Policy: audits that are intersectional and cross‑cultural should be required pre‑deployment. Disclosure requirements around personalization behavior, language/region training coverage, and subgroup performance should be mandated.
    • Cost/benefit: regulators and funders should evaluate both efficiency gains from LLM assistance and distributional harms; remediation (fine‑tuning, localized data, targeted tutoring supplements) entails nontrivial costs but may be necessary to prevent widening inequality.
  • Design & intervention economics
    • Intervention options (and their economic tradeoffs):
      • Intersectional audits and monitoring (low marginal cost, early detection).
      • Targeted fine‑tuning / data augmentation for underrepresented languages and intersectional profiles (higher upfront cost; reduces biased priors).
      • Prompt engineering and guardrails to equalize complexity conditional on demonstrated ability rather than demographic cues (operationally cheaper but brittle).
      • Subsidized deployment of fairness‑tuned models to underserved institutions (public expenditure; direct redistribution of learning inputs).
    • Comparative-effectiveness evaluation recommended: run randomized trials to estimate causal impacts of bias‑mitigating interventions on learning outcomes and downstream earnings to guide resource allocation.
  • Measurement and modelling recommendations for economists
    • Treat LLM personalization as a channel that shifts effective instructional quality; parameterize differential learning rates by subgroup in human capital accumulation models.
    • Quantify welfare impacts using counterfactual scenarios: (i) baseline no‑LLM, (ii) LLM with bias, (iii) LLM with mitigation—estimate present value of lifetime earnings and distributional metrics (Gini, poverty headcount).
    • Incorporate institution‑level heterogeneity: model interactions between institutional prestige and LLM effects (since biases persist inside elites).
    • Track language composition and training source exposure as key determinants of distributional outcomes across countries.
  • Policy takeaways
    • Require intersectional, cross‑cultural audits before large‑scale educational deployment.
    • Fund and prioritize training data and alignment efforts for non‑dominant languages and marginalized contexts.
    • Mandate transparency and reporting on subgroup performance metrics (MCV/MGL analogues) for educational LLMs.
    • Support remediation (targeted tutoring, subsidies) where audits reveal systematic harms.

Caveats and limits (affect interpretation and economic projections) - Synthetic profiles approximate but do not equal real student behavior; downstream educational outcomes were not measured—readability and selected difficulty are proxies for instructional quality. - Datasets focus on STEM/engineering; generalization to other subjects or K–12 contexts requires further study. - Only four models tested; while architectures varied, broader model family coverage would strengthen claims about universality. - Deterministic decoding (temperature = 0) removes stochastic variation but may not reflect all real‑world deployments.

Overall, the paper demonstrates measurable, consistent, and economically consequential patterns of intersectional disadvantage in LLM‑generated STEM instruction. For economists and policymakers, these results argue for integrating fairness audits into cost‑benefit analyses of AI in education and for modeling the distributional (not just aggregate) impacts of AI‑mediated learning.

Assessment

Paper Typequasi_experimental Evidence Strengthmedium — Strong internal validity from controlled, factorial manipulation of profile attributes, multiple models, statistical correction, and model-agnostic attribution (SHAP) supports robust detection of systematic biases in LLM outputs; however external validity is limited because profiles are synthetic, real-student behavior and downstream educational outcomes are not measured, only four model architectures/versions are tested, and mechanisms (training data, fine-tuning) are not directly identified. Methods Rigorhigh — Rigorous design for an audit study: cross-cultural intersectional profile matrix, multiple tasks (ranking and generation), multiple LLM architectures, false-discovery-rate correction for multiple tests, and SHAP-based attribution to decompose feature effects indicate careful methodological implementation and robustness checks; potential weaknesses (prompt choice sensitivity, synthetic profile realism, and lack of longitudinal/real-world validation) are acknowledged but do not undercut the within-experiment rigor. SampleSynthetic audit dataset consisting of factorial combinations of demographic and educational attributes (India: caste, medium of instruction, college tier; U.S.: race, HBCU attendance, school type) plus income, gender, and disability; evaluations run on four LLMs (Qwen 2.5-32B-Instruct, GPT-4o, GPT-4o-mini, GPT-OSS 20B) across ranking and generation STEM tasks, with outcomes scored in grade-level-equivalent metrics and subjected to FDR-corrected statistical tests and SHAP analysis. Themesinequality skills_training IdentificationControlled audit using synthetic student profiles: the authors systematically vary demographic and background attributes in prompts (caste, medium of instruction, college tier for India; race, HBCU attendance, school type for the U.S.; plus income, gender, disability) and compare model outputs across these experimental profile conditions, using ranking and generation tasks, FDR-corrected hypothesis testing for statistical significance, and SHAP feature-attribution to quantify each attribute's contribution to output quality differences. GeneralizabilityUses synthetic profiles rather than real student interactions, which may not capture real-world behavior or cueing, Only four model architectures/versions tested — results may not generalize to other LLMs or future model updates, Limited to STEM content and specific task types (ranking, generation) — other pedagogical formats may differ, Cultural mappings of attributes (e.g., caste proxies, race categories, HBCU effects) may oversimplify on-the-ground heterogeneity, Does not measure downstream educational or labor-market outcomes, so implications for productivity and wages are indirect, Prompt engineering choices and evaluation rubric may influence measured gaps

Claims (10)

ClaimDirectionOutcomeConfidence & EvidenceDetails
LLM-generated STEM content systematically disadvantages marginalized student profiles across two cultural contexts. Output Quality negative model-generated content quality for student profiles (advantage/disadvantage)
Reading fidelity high
Study strength medium
not reported
0.48
The gap between the most privileged and most marginalized profiles reaches 2.55 grade levels. Output Quality negative grade-level equivalent of LLM-generated STEM content (performance/quality)
Reading fidelity high
Study strength medium
2.55 grade levels
0.48
Four LLMs were audited: Qwen 2.5-32B-Instruct, GPT-4o, GPT-4o-mini, and GPT-OSS 20B. Other null_result None
Reading fidelity high
Study strength high
n=4
0.8
Synthetic profiles crossed Indian-specific dimensions (caste, medium of instruction, college tier) and American-specific dimensions (race, HBCU attendance, school type), alongside income, gender, and disability. Other null_result None
Reading fidelity high
Study strength high
not reported
0.8
Income produces significant effects across every model and context. Output Quality negative model output quality as a function of income signal
Reading fidelity high
Study strength medium
not reported
0.48
Medium of instruction drives the largest single effect in the Indian context. Output Quality negative model output quality as a function of medium of instruction
Reading fidelity high
Study strength medium
not reported
0.48
Disability status triggers simpler explanations. Output Quality negative explanation complexity / depth produced by the model
Reading fidelity medium
Study strength medium
not reported
0.29
Effects compound non-additively: marginalization across multiple dimensions produces gaps larger than any single dimension predicts. Output Quality negative combined effect on model output quality from multiple demographic/marginalization signals
Reading fidelity high
Study strength medium
not reported
0.48
Biases persist within elite institutions. Output Quality negative model output quality conditional on elite-institution signals
Reading fidelity medium
Study strength medium
not reported
0.29
Bias is consistent across all four architectures and persists through model selection, making intersectional, cross-cultural auditing a structural requirement before deployment. Other negative presence/consistency of bias across model architectures and selection steps
Reading fidelity high
Study strength medium
not reported
0.48

Notes