1 cumulative citations
View corpus contextStandard educational tests do not map onto LLM abilities the same way they do for students: latent factor analyses on chemistry and quantitative-reasoning exams show systematic structural differences between human and LLM response patterns, calling into question claims that human-designed assessments directly certify model competencies.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
The rapid development and growing deployment of large language models (LLMs) have made it increasingly important to understand their capabilities. A common approach is to evaluate LLMs using assessment instruments originally designed to measure skills and competencies in humans, such as standardized exams, and to use performance on these instruments as evidence for generalizable claims about LLMs' underlying abilities on the same skills the assessments are intended to measure in humans. However, from a validity perspective, such inferences require that the relationship between observed performance and underlying constructs established for humans also holds for LLMs. In particular, a necessary condition for transferring score interpretations is similarity in the latent structure of responses to the assessment. In this study, we examine whether this condition holds in two educational contexts: high-school chemistry and a quantitative reasoning section of a university entrance exam. Using a case study design, we compare human response data with responses generated by six multimodal LLMs. Our analytical approach combines exploratory factor analysis, factor congruence, and resampling to assess latent structure similarity across human learners and LLMs. Across both instruments, we find systematic differences between human and LLM factor structures, showing evidence that the analyzed assessments may not measure the same constructs for humans and LLMs. These findings call into question the validity of evaluation practices that use educational assessments to make claims about AI capabilities.
Summary
Main Finding
Assessment instruments developed for humans do not necessarily measure the same latent constructs when administered to large language models (LLMs). In two case studies (a high‑school chemistry diagnostic and a quantitative reasoning section of a national university entrance exam), exploratory factor analysis and resampling-based factor-congruence tests showed systematic differences between human and pooled LLM factor structures, casting doubt on the validity of interpreting LLM scores as evidence of the same human-targeted abilities.
Key Points
- Scope: Two multimodal, multiple-choice instruments:
- High‑school chemistry diagnostic (22 items; 931 students; 7 chemistry items later removed for stability).
- Quantitative reasoning section (20 items; >4,800 examinees).
- LLM sample: six multimodal models (OpenAI GPT-4o, GPT-5.2; Google Gemini 1.5 Pro, Gemini 3 Pro; Anthropic Claude 3.5 Sonnet, Claude 4.5). For each model/instrument, 20 independent runs → pooled LLM group of 120 response sets.
- Prompting & scoring: Minimal zero‑shot prompt; full PDF uploaded; models requested to provide only final answer choices; outputs binarized (1 correct / 0 incorrect); skipped/invalid treated as incorrect.
- Analytical pipeline:
- Preprocessing: removed zero-variance/sparse items based on 2×2 contingency diagnostics; used tetrachoric correlations as input.
- Factor-retention: Kaiser criterion and Parallel Analysis applied separately to human and LLM correlation matrices.
- Factor extraction: Exploratory factor analysis (minres / oblimin rotation) fit separately for humans and LLMs.
- Structural comparison: repeated resampling (100 iterations) with matched sample size (120): two human subsamples (H1,H2) and one pooled LLM sample (B) per iteration; Tucker congruence (cosine similarity) computed between loading vectors; factors matched via Hungarian algorithm using absolute congruence; compared mean matched congruence for human–human (HH) vs LLM–human (LH) using one‑sided Wilcoxon tests.
- Results:
- Factor-retention disagreement: Kaiser often produced the same retained factor count across groups (Chemistry=4, Quant=5), while Parallel Analysis diverged (Chemistry: humans 5 vs LLMs 4; Quantitative: humans 7–8 vs LLMs 5).
- Factor loadings and patterns differed qualitatively between humans and LLMs; resampling tests showed HH congruence distributions were systematically higher than LH congruence (i.e., human–human factor structure is more self-consistent than LLM–human).
- Conclusion: Evidence that these instruments may not capture the same latent constructs in humans and LLMs; transfer of human‑validated score interpretations to LLMs is not guaranteed.
Data & Methods
- Data:
- Human data: real exam/diagnostic administration (authentic conditions); non-public to ensure instrument non-exposure and quality.
- LLM data: collected via online model interfaces with fresh temporary chats to reduce carryover; responses pooled across models to increase variance suitable for EFA.
- Psychometric processing:
- Binary response matrices → tetrachoric correlation matrices.
- Preprocessing to remove items causing unstable tetrachoric estimates (zero/very small contingency cell counts).
- Factor analysis specifics:
- Retention rules: Kaiser criterion and Parallel Analysis (30 simulated datasets; reported stability across runs).
- EFA: least squares/minres extraction, oblimin (oblique) rotation to allow correlated factors.
- Structural-similarity quantification:
- Tucker congruence between factor loading vectors.
- Hungarian algorithm for optimal one-to-one factor matching (accounting for sign indeterminacy).
- Resampling with matched sample size (120) to obtain distributions for HH and LH congruence; statistical comparison via Wilcoxon rank-sum.
- Important methodological choices (and caveats):
- Pooling LLM responses across models: aligns with treating “LLMs as a class” and increases variance for correlation estimation, but mixes heterogeneous model behaviors.
- Zero‑shot minimal prompting chosen to mimic realistic exam-taking and avoid prompt-induced artifacts; results may depend on prompting strategy and interface.
- Human data non-public: quality advantage, but limits independent replication.
- Analysis focuses on EFA-based latent-structure similarity — a necessary but not sufficient condition for full construct equivalence.
Implications for AI Economics
- Measurement and benchmarking:
- Economic analyses that rely on published LLM performance on human-designed standardized tests (e.g., to infer workforce skill substitution, productivity impacts, or credential equivalence) risk misinterpretation if latent-structure equivalence is not established.
- Benchmarks that report aggregate scores without examining measurement invariance (latent structure) can send misleading market signals to firms, educators, and policymakers.
- Policy, regulation, and procurement:
- Regulators or institutions using test-based claims (e.g., “model achieves X percentile on [human] exam”) should require evidence that the exam measures comparable constructs in models and humans before relying on scores for high-stakes decisions (certification, deployment limits, procurement).
- Investment and adoption decisions:
- Investors and firms using benchmark performance to forecast returns or make adoption choices should account for measurement uncertainty: a high score may not reflect the same capability as a human score, affecting expected productivity gains and risk assessments.
- Labor-market modeling:
- Models of automation risk or occupational exposure that translate LLM test performance into task-capability estimates need measurement-invariance checks; otherwise, projections (e.g., job displacement probabilities) may be biased.
- Recommended best practices for AI economists and evaluators:
- Require and report measurement-invariance analyses (e.g., factor congruence, item-level DIF, EFA/confirmatory factor analysis) when using human-targeted assessments to evaluate AI.
- Favor developing AI-oriented or hybrid assessment instruments with explicit construct definitions for models, or adapt human instruments and validate construct equivalence empirically before transferring interpretations.
- Demand transparency: publish item-level LLM response patterns and analysis code to enable independent evaluation of latent structure.
- Incorporate uncertainty from measurement non-equivalence into economic models and sensitivity analyses (e.g., scenario bounds for productivity or substitution effects).
- Broader research implications:
- Need for interdisciplinary work combining psychometrics, ML evaluation, and economics to develop valid, reliable measurement frameworks for AI capabilities that are informative for market and policy decisions.
Summary takeaway: Before using human-designed educational assessments as direct evidence of LLM abilities for economic or policy decisions, verify that the instrument measures the same latent constructs in models as in humans. Without such validation, scores can mislead stakeholders about the nature and magnitude of AI capabilities.
Assessment
Claims (9)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| The study finds systematic differences between human and LLM factor structures across both assessment instruments, suggesting that the analyzed assessments may not measure the same constructs for humans and LLMs. Other | negative | Similarity of latent factor structures and construct interpretations across human and LLM responses |
Reading fidelity
high
Study strength
medium
|
not reported
|
| The chemistry assessment contained 22 multiple-choice items and was completed by 931 Grade 11–12 students. Other | null_result | Human response data coverage for the chemistry assessment |
Reading fidelity
high
Study strength
high
|
n=931
|
| The quantitative reasoning assessment consisted of 20 multiple-choice items and was completed by more than 4,800 examinees. Other | null_result | Human response data coverage for the quantitative reasoning assessment |
Reading fidelity
high
Study strength
high
|
not reported
|
| The researchers collected 20 independent response sets from each of six multimodal LLMs, producing 120 LLM response sets per assessment instrument. Other | null_result | LLM response-set sample size used for latent-structure analysis |
Reading fidelity
high
Study strength
high
|
n=120
120 response sets per instrument
|
| Seven chemistry items were removed because they contributed to highly sparse pairwise response tables and could make tetrachoric correlation estimates unstable. Other | negative | Usable item set and stability of item-correlation estimates |
Reading fidelity
high
Study strength
high
|
n=7
7 items removed
|
| Parallel analysis produced different factor-retention results for humans and LLMs in both datasets: humans retained five factors in chemistry while LLMs most often retained four, and humans retained seven to eight factors in quantitative reasoning while LLMs consistently retained five. Other | negative | Number of latent factors retained by the assessment |
Reading fidelity
high
Study strength
medium
|
Chemistry: 5 human vs. 4 LLM factors; quantitative reasoning: 7–8 human vs. 5 LLM factors
|
| The evidence for differences in the number of retained factors depends on the retention method: the Kaiser criterion produced the same number of factors for humans and LLMs within each instrument, whereas parallel analysis produced different numbers. Other | mixed | Agreement between human and LLM factor-retention results |
Reading fidelity
high
Study strength
medium
|
not reported
|
| The study used repeated resampling with matched samples of 120 human respondents and 120 LLM responses to compare human–human and LLM–human factor-structure similarity. Other | null_result | Factor congruence between assessment response groups |
Reading fidelity
high
Study strength
high
|
n=120
100 repetitions per factor-number choice
|
| The paper argues that using human-designed educational assessments to make construct-level claims about LLM capabilities may be invalid when the latent response structures differ between humans and LLMs. Ai Safety And Ethics | negative | Validity of transferring assessment-score interpretations from humans to LLMs |
Reading fidelity
high
Study strength
medium
|
not reported
|