The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Standard educational tests do not map onto LLM abilities the same way they do for students: latent factor analyses on chemistry and quantitative-reasoning exams show systematic structural differences between human and LLM response patterns, calling into question claims that human-designed assessments directly certify model competencies.

Do Assessment Instruments Measure the Same Thing for Humans and LLMs? A Latent Structure Analysis
Alona Strugatski, Licol Zeinfeld, Giora Alexandron · August 16, 2026
arxiv descriptive medium evidence 7/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

Arxiv

Latest observation:

  1. Alona Strugatski unresolved corpus identity
  2. Licol Zeinfeld unresolved corpus identity
  3. Giora Alexandron unresolved corpus identity

Semantic Scholar

Latest observation:

  1. Alona Strugatski provider ID
  2. Licol Zeinfeld provider ID
  3. Giora Alexandron provider ID
Using exploratory factor analysis and resampling on two educational assessments, the authors find systematic differences in latent factor structures between human examinees and pooled responses from six multimodal LLMs, suggesting the instruments may not measure the same constructs across the two populations.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

The rapid development and growing deployment of large language models (LLMs) have made it increasingly important to understand their capabilities. A common approach is to evaluate LLMs using assessment instruments originally designed to measure skills and competencies in humans, such as standardized exams, and to use performance on these instruments as evidence for generalizable claims about LLMs' underlying abilities on the same skills the assessments are intended to measure in humans. However, from a validity perspective, such inferences require that the relationship between observed performance and underlying constructs established for humans also holds for LLMs. In particular, a necessary condition for transferring score interpretations is similarity in the latent structure of responses to the assessment. In this study, we examine whether this condition holds in two educational contexts: high-school chemistry and a quantitative reasoning section of a university entrance exam. Using a case study design, we compare human response data with responses generated by six multimodal LLMs. Our analytical approach combines exploratory factor analysis, factor congruence, and resampling to assess latent structure similarity across human learners and LLMs. Across both instruments, we find systematic differences between human and LLM factor structures, showing evidence that the analyzed assessments may not measure the same constructs for humans and LLMs. These findings call into question the validity of evaluation practices that use educational assessments to make claims about AI capabilities.

Summary

Main Finding

Assessment instruments developed for humans do not necessarily measure the same latent constructs when administered to large language models (LLMs). In two case studies (a high‑school chemistry diagnostic and a quantitative reasoning section of a national university entrance exam), exploratory factor analysis and resampling-based factor-congruence tests showed systematic differences between human and pooled LLM factor structures, casting doubt on the validity of interpreting LLM scores as evidence of the same human-targeted abilities.

Key Points

  • Scope: Two multimodal, multiple-choice instruments:
    • High‑school chemistry diagnostic (22 items; 931 students; 7 chemistry items later removed for stability).
    • Quantitative reasoning section (20 items; >4,800 examinees).
  • LLM sample: six multimodal models (OpenAI GPT-4o, GPT-5.2; Google Gemini 1.5 Pro, Gemini 3 Pro; Anthropic Claude 3.5 Sonnet, Claude 4.5). For each model/instrument, 20 independent runs → pooled LLM group of 120 response sets.
  • Prompting & scoring: Minimal zero‑shot prompt; full PDF uploaded; models requested to provide only final answer choices; outputs binarized (1 correct / 0 incorrect); skipped/invalid treated as incorrect.
  • Analytical pipeline:
    • Preprocessing: removed zero-variance/sparse items based on 2×2 contingency diagnostics; used tetrachoric correlations as input.
    • Factor-retention: Kaiser criterion and Parallel Analysis applied separately to human and LLM correlation matrices.
    • Factor extraction: Exploratory factor analysis (minres / oblimin rotation) fit separately for humans and LLMs.
    • Structural comparison: repeated resampling (100 iterations) with matched sample size (120): two human subsamples (H1,H2) and one pooled LLM sample (B) per iteration; Tucker congruence (cosine similarity) computed between loading vectors; factors matched via Hungarian algorithm using absolute congruence; compared mean matched congruence for human–human (HH) vs LLM–human (LH) using one‑sided Wilcoxon tests.
  • Results:
    • Factor-retention disagreement: Kaiser often produced the same retained factor count across groups (Chemistry=4, Quant=5), while Parallel Analysis diverged (Chemistry: humans 5 vs LLMs 4; Quantitative: humans 7–8 vs LLMs 5).
    • Factor loadings and patterns differed qualitatively between humans and LLMs; resampling tests showed HH congruence distributions were systematically higher than LH congruence (i.e., human–human factor structure is more self-consistent than LLM–human).
    • Conclusion: Evidence that these instruments may not capture the same latent constructs in humans and LLMs; transfer of human‑validated score interpretations to LLMs is not guaranteed.

Data & Methods

  • Data:
    • Human data: real exam/diagnostic administration (authentic conditions); non-public to ensure instrument non-exposure and quality.
    • LLM data: collected via online model interfaces with fresh temporary chats to reduce carryover; responses pooled across models to increase variance suitable for EFA.
  • Psychometric processing:
    • Binary response matrices → tetrachoric correlation matrices.
    • Preprocessing to remove items causing unstable tetrachoric estimates (zero/very small contingency cell counts).
  • Factor analysis specifics:
    • Retention rules: Kaiser criterion and Parallel Analysis (30 simulated datasets; reported stability across runs).
    • EFA: least squares/minres extraction, oblimin (oblique) rotation to allow correlated factors.
  • Structural-similarity quantification:
    • Tucker congruence between factor loading vectors.
    • Hungarian algorithm for optimal one-to-one factor matching (accounting for sign indeterminacy).
    • Resampling with matched sample size (120) to obtain distributions for HH and LH congruence; statistical comparison via Wilcoxon rank-sum.
  • Important methodological choices (and caveats):
    • Pooling LLM responses across models: aligns with treating “LLMs as a class” and increases variance for correlation estimation, but mixes heterogeneous model behaviors.
    • Zero‑shot minimal prompting chosen to mimic realistic exam-taking and avoid prompt-induced artifacts; results may depend on prompting strategy and interface.
    • Human data non-public: quality advantage, but limits independent replication.
    • Analysis focuses on EFA-based latent-structure similarity — a necessary but not sufficient condition for full construct equivalence.

Implications for AI Economics

  • Measurement and benchmarking:
    • Economic analyses that rely on published LLM performance on human-designed standardized tests (e.g., to infer workforce skill substitution, productivity impacts, or credential equivalence) risk misinterpretation if latent-structure equivalence is not established.
    • Benchmarks that report aggregate scores without examining measurement invariance (latent structure) can send misleading market signals to firms, educators, and policymakers.
  • Policy, regulation, and procurement:
    • Regulators or institutions using test-based claims (e.g., “model achieves X percentile on [human] exam”) should require evidence that the exam measures comparable constructs in models and humans before relying on scores for high-stakes decisions (certification, deployment limits, procurement).
  • Investment and adoption decisions:
    • Investors and firms using benchmark performance to forecast returns or make adoption choices should account for measurement uncertainty: a high score may not reflect the same capability as a human score, affecting expected productivity gains and risk assessments.
  • Labor-market modeling:
    • Models of automation risk or occupational exposure that translate LLM test performance into task-capability estimates need measurement-invariance checks; otherwise, projections (e.g., job displacement probabilities) may be biased.
  • Recommended best practices for AI economists and evaluators:
    • Require and report measurement-invariance analyses (e.g., factor congruence, item-level DIF, EFA/confirmatory factor analysis) when using human-targeted assessments to evaluate AI.
    • Favor developing AI-oriented or hybrid assessment instruments with explicit construct definitions for models, or adapt human instruments and validate construct equivalence empirically before transferring interpretations.
    • Demand transparency: publish item-level LLM response patterns and analysis code to enable independent evaluation of latent structure.
    • Incorporate uncertainty from measurement non-equivalence into economic models and sensitivity analyses (e.g., scenario bounds for productivity or substitution effects).
  • Broader research implications:
    • Need for interdisciplinary work combining psychometrics, ML evaluation, and economics to develop valid, reliable measurement frameworks for AI capabilities that are informative for market and policy decisions.

Summary takeaway: Before using human-designed educational assessments as direct evidence of LLM abilities for economic or policy decisions, verify that the instrument measures the same latent constructs in models as in humans. Without such validation, scores can mislead stakeholders about the nature and magnitude of AI capabilities.

Assessment

Paper Typedescriptive Evidence Strengthmedium — The paper uses appropriate psychometric methods (EFA, parallel analysis, factor congruence, resampling) and large human samples (931 and ~4,800) which support the claim that latent structures differ; however, inference is limited to two instruments, uses pooled LLM responses (120 runs) rather than fully independent model populations, and relies on a non-public human dataset, reducing reproducibility and external validation. Methods Rigormedium — Methodologically sound choices (tetrachoric correlations for binary data, multiple retention criteria, oblique rotation, Hungarian matching, resampling for human baseline) indicate careful design, but some design choices weaken rigor: pooling multiple LLMs into a single group obscures model-level heterogeneity, zero-shot single-prompt collection via web UIs may introduce uncontrolled variability, several chemistry items were removed due to sparsity, and human data are non-public which limits independent verification. SampleHuman data: (1) High-school chemistry diagnostic — 22 multiple-choice items (7 later removed), N=931 students (mean score 71.49/100, SD 16.95); (2) Quantitative reasoning section — 20 multiple-choice items, N>4,800 examinees (mean 12.45/20, SD 3.75). LLM data: responses from six multimodal LLMs (GPT-4o, GPT-5.2, Gemini 1.5 Pro, Gemini 3 Pro, Claude 3.5 Sonnet, Claude 4.5), 20 independent runs per model per instrument, pooled to 120 LLM response sets per instrument (chemistry pooled M=76.14, SD=15.41; quantitative pooled M=11.94, SD=4.52); responses scored dichotomously against answer keys. Themeshuman_ai_collab skills_training GeneralizabilityOnly two assessment instruments (high-school chemistry and a single quantitative-reasoning section) — results may not generalize to other domains or test types, LLM responses pooled across six models — obscures heterogeneity across individual model families and versions, Data are non-public, limiting external replication and verification, Items removed from chemistry due to sparsity may alter the instrument’s representativeness, Zero-shot, single-prompt web-interface protocol may not reflect other prompting styles or fine-tuned/evaluation-specific setups, Dichotomous scoring ignores partial reasoning traces or process-level outputs that could differ between humans and LLMs

Claims (9)

ClaimDirectionOutcomeConfidence & EvidenceDetails
The study finds systematic differences between human and LLM factor structures across both assessment instruments, suggesting that the analyzed assessments may not measure the same constructs for humans and LLMs. Other negative Similarity of latent factor structures and construct interpretations across human and LLM responses
Reading fidelity high
Study strength medium
not reported
0.18
The chemistry assessment contained 22 multiple-choice items and was completed by 931 Grade 11–12 students. Other null_result Human response data coverage for the chemistry assessment
Reading fidelity high
Study strength high
n=931
0.3
The quantitative reasoning assessment consisted of 20 multiple-choice items and was completed by more than 4,800 examinees. Other null_result Human response data coverage for the quantitative reasoning assessment
Reading fidelity high
Study strength high
not reported
0.3
The researchers collected 20 independent response sets from each of six multimodal LLMs, producing 120 LLM response sets per assessment instrument. Other null_result LLM response-set sample size used for latent-structure analysis
Reading fidelity high
Study strength high
n=120
120 response sets per instrument
0.3
Seven chemistry items were removed because they contributed to highly sparse pairwise response tables and could make tetrachoric correlation estimates unstable. Other negative Usable item set and stability of item-correlation estimates
Reading fidelity high
Study strength high
n=7
7 items removed
0.3
Parallel analysis produced different factor-retention results for humans and LLMs in both datasets: humans retained five factors in chemistry while LLMs most often retained four, and humans retained seven to eight factors in quantitative reasoning while LLMs consistently retained five. Other negative Number of latent factors retained by the assessment
Reading fidelity high
Study strength medium
Chemistry: 5 human vs. 4 LLM factors; quantitative reasoning: 7–8 human vs. 5 LLM factors
0.18
The evidence for differences in the number of retained factors depends on the retention method: the Kaiser criterion produced the same number of factors for humans and LLMs within each instrument, whereas parallel analysis produced different numbers. Other mixed Agreement between human and LLM factor-retention results
Reading fidelity high
Study strength medium
not reported
0.18
The study used repeated resampling with matched samples of 120 human respondents and 120 LLM responses to compare human–human and LLM–human factor-structure similarity. Other null_result Factor congruence between assessment response groups
Reading fidelity high
Study strength high
n=120
100 repetitions per factor-number choice
0.3
The paper argues that using human-designed educational assessments to make construct-level claims about LLM capabilities may be invalid when the latent response structures differ between humans and LLMs. Ai Safety And Ethics negative Validity of transferring assessment-score interpretations from humans to LLMs
Reading fidelity high
Study strength medium
not reported
0.18

Notes