The Commonplace
Home Three-study pilot Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

LLM labels of EU AI Act submissions are nearly perfectly reproducible but fail to match what stakeholders say in surveys; businesses tend to amplify AI-risk language in public consultations while regulators and some non-business groups do not.

Reproducibility is not construct validity: LLM measurement of institutionally situated communication
Batzdorfer, Veronika, Santagiustina, Carlo Romano Marcello Alessandro · September 17, 2026 · arXiv (Cornell University)
openalex correlational medium evidence 7/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

OpenAlex

Latest observation:

  1. Batzdorfer, Veronika provider ID
  2. Santagiustina, Carlo Romano Marcello Alessandro provider ID

Semantic Scholar

Latest observation:

  1. Veronika Batzdorfer provider ID
  2. C. R. M. A. Santagiustina provider ID
LLM annotations of EU AI Act consultation texts are extremely reproducible but correlate weakly with stakeholders' survey-reported AI concerns, with systematic divergences by stakeholder type (business associations amplify risk publicly compared with their survey responses).

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

High annotation reproducibility does not necessarily imply that an LLM-inferred measure captures the construct it is intended to measure. We test this distinction using a dataset from the European Commission's AI Act consultation, linking structured survey responses to free-text consultation submissions from the same stakeholders. LLM annotations of consultation submissions are highly reproducible (intraclass correlations > 0.99), yet show limited convergence with survey-reported measures of the nominal construct they were intended to approximate. Divergence between survey-and LLM-inferred text-based measures varies systematically across stakeholder groups: business associations express greater concern about AI risks in text-based consultations than in survey responses ({g} = +1.0), whereas public authorities and several nonbusiness groups show smaller or negative divergences. Divergences between scores suggest positive spatial autocorrelation across European countries (Moran's I = 0.347, p = 0.036), indicating that stakeholders from neighboring countries tend toward more similar text-based stances towards AI safety concerns. Despite divergence, survey-reported concerns remain strongly associated with support for explainability across all divergence levels. These results demonstrate that LLM annotation reproducibility can coexist with poor construct correspondence and motivate validation procedures that distinguish reproducibility, construct validity, and communication context variation when LLMs are used as measurement instruments.

Summary

Main Finding

LLM annotations of public consultation texts can be extremely reproducible while failing to validly recover the same latent constructs measured in surveys. In the European Commission AI Act consultation, Qwen3.5 annotations showed near-perfect inter-run reproducibility (ICC > 0.99) but weak convergence with stakeholders’ survey-reported AI concerns (Pearson r ≈ 0.03–0.18; Lin’s CCC < 0.08). Divergences are systematic across institutional and geographic contexts (e.g., business associations amplify AI-risk language in public submissions relative to their survey responses).

Key Points

  • Reproducibility ≠ construct validity: High LLM inter-run agreement (ICC(C,1) 0.994–0.996; ICC(C,k) > 0.999) did not imply that the LLM-derived scores measured the same latent constructs elicited by surveys.
  • Poor convergence with surveys:
    • Safety concern: r = 0.029 (p = 0.59)
    • Rights concern: r = 0.176 (p = 0.001)
    • Explainability: r = 0.097 (p = 0.072)
    • Lin’s concordance coefficients near zero (0.013–0.080).
  • Large systematic bias: Bland–Altman analysis showed limits of agreement spanning ~±4 points on a 0–5 scale; standardized mean differences (Cohen’s d) ≈ 1.0–1.18.
  • Systematic heterogeneity by stakeholder type:
    • Business associations: positive divergence (public texts show higher AI-risk emphasis than their surveys; mean gap ≈ +1.0).
    • Public authorities and several non-business groups: smaller or negative divergences.
  • Geographic patterning: positive spatial autocorrelation of divergence (Moran’s I = 0.347, p = 0.036), though the authors caution sensitivity to multiple-testing and sample limitations.
  • Downstream robustness: despite divergence, survey-reported concern remains strongly associated with support for explainability across divergence levels (tested with Double/Debiased Machine Learning), indicating distinct but still informative relationships.
  • Conceptual caution: public consultation texts are institutionally situated artifacts (audience targeting, reputation management, access channels). Treating text-derived proxies as error-free measures of private or survey-attested attitudes can induce collider or selection biases in downstream analyses (Text is a collider: T → Text ← Y).

Data & Methods

  • Data: Public submissions to the European Commission AI Regulation consultation (Better Regulation Portal), N = 857 submissions collected 2020-02-20 to 2021-04-27 across 3 rounds. Validation sample: N = 348 submissions with both free-text and matched structured survey items (0–5 scales for safety concern, rights concern, explainability).
  • LLM annotation:
    • Model: Qwen3.5-397b-A17b.
    • Task: annotate each document on three continuous 0–5 scales (safety_concern, rights_concern, explainability_trust); prompt enforced strict JSON and included scale anchors and binary flags (e.g., regulatory, biometric mentions).
    • Inter-run checks: five independent annotation runs; near-identical aggregate scores (mean vs median difference < 0.01).
  • Validation diagnostics:
    • Convergent/discriminant checks: Pearson correlation, Lin’s CCC, Bland–Altman limits, Cohen’s d, ICCs for cross-setting agreement.
    • Examination of systematic divergence across institutional groups and countries.
  • Spatial analysis:
    • Moran’s I with queen contiguity weights (shared land borders), permutation inference (B = 9,999), primary sample of 23 countries.
  • Causal/robustness check:
    • Double/Debiased Machine Learning (DML) to estimate association between high survey-reported safety concern and explainability support.
    • Implementations: LightGBM for nuisance estimation with 5-fold cross-fitting; Causal Forest DML parameters reported (Ntrees = 1000, minleaf = 10).
  • Conceptual framing: Directed acyclic graph showing Text as a downstream collider of multiple latent constructs (T and Y) and unobserved confounders U — conditioning on text-derived features can open spurious paths.

Implications for AI Economics

  • Measurement risks in text-based variables: LLM-derived textual measures may reflect public-facing, institutionally mediated communication (framing, signaling, audience targeting) rather than private or survey-attested preferences. Econometric analyses that treat such measures as direct proxies for latent attitudes risk mis-specification and biased estimates.
  • Collider and omitted-variable bias: Using text-derived features as covariates/controls without validating construct correspondence can induce collider stratification, opening spurious associations between explanatory variables and outcomes. Researchers should explicitly model the measurement process and possible downstream selection/collider structures.
  • Policy inference and stakeholder analysis: Public statements scraped from consultations or other regulatory submissions can systematically over- or understate concerns by stakeholder type (e.g., businesses amplify risk rhetorically). Policymakers and analysts should not equate public rhetorical intensity with private preferences or likelihood to support/comply with regulation.
  • Cross-country and institutional heterogeneity matter: Cultural, reputational, and access-channel differences shape public communication. Comparative AI-economics work that uses text measures must test for—and, where appropriate, model—heterogeneous measurement error across organizations and countries.
  • Recommended empirical practices:
    • Validate LLM text measures against external criteria (matched surveys, behavioral signals) before using them as proxies in causal models.
    • Distinguish measurement targets: if the interest is public communication/positioning, LLM scores can be valid indicators; if the interest is underlying private attitudes, additional validation and correction are required.
    • Report both reproducibility and construct validity metrics separately; high inter-run agreement alone is insufficient.
    • Conduct subgroup analyses (by stakeholder type, country) and sensitivity tests to detect systematic divergence.
    • Where text-derived measures are used as controls, consider causal identification strategies robust to measurement-induced collider bias (e.g., instrumenting, explicit measurement-error models, bounding approaches).
    • Combine sources (surveys, private channels, behavioral data) and exploit methods like DML while maintaining external validation of the target construct.
  • Practical takeaway for AI-economics researchers: LLMs are powerful scalers of text, but their outputs should be interpreted as indicators of communicative expression in context. Use them thoughtfully—validate, test heterogeneity, and avoid treating reproducible LLM annotations as automatically valid measures of latent economic or policy attitudes.

Assessment

Paper Typecorrelational Evidence Strengthmedium — The paper uses a direct within-subject linkage (text and survey from the same stakeholder), multiple complementary diagnostics, and robust reproducibility checks, which provide solid evidence that LLM annotations can be highly reproducible yet diverge from survey measures in this context; however, findings are based on a single consultation domain, one LLM/prompt formulation, and a modest validation sample (N=348), which limits wider generalization. Methods Rigorhigh — Multiple appropriate validation techniques (ICC, Pearson r, Lin's CCC, Bland–Altman), repeated LLM runs to assess reliability, subgroup analyses by stakeholder type, spatial autocorrelation testing with permutation inference, and DML with cross-fitting for downstream robustness together show careful and modern methodological practice; limitations include reliance on one model/prompt and sensitivity of some spatial results to multiple-testing correction. SampleWeb-crawled public consultation submissions to the European Commission's AI Act (N = 857 submissions from 20-02-2020 to 27-04-2021 across three rounds); validation sample includes N = 348 submissions that contain both free-text submissions and structured survey items measuring safety concern, rights concern, and explainability importance (0–5 scale); stakeholders span business associations, public authorities, NGOs and other groups across ~23 European countries (country-level divergence estimated where possible). Themesgovernance org_design IdentificationWithin-subject validation: link free-text public consultation submissions to structured survey responses from the same stakeholders in the EU AI Act consultation; generate continuous 0–5 annotations from Qwen3.5 (five independent runs) and assess reproducibility (ICCs) and construct correspondence via Pearson r, Lin's concordance, Bland–Altman limits, group-specific comparisons, spatial clustering (Moran's I), and Double/Debiased Machine Learning for downstream associations. GeneralizabilitySingle policy domain (EU AI Act) and institutional context (regulatory consultation) — may not generalize to other text genres (social media, interviews, private messages)., Validation sample limited to submissions that contained both text and survey items (selection/self-selection bias)., Annotations produced with a single LLM (Qwen3.5) and a specific prompt; results may vary with different models or prompt designs., Temporal limitation: documents collected 2020–2021; attitudes and public communication strategies may change over time., Geographic limitation: European stakeholders only; cultural and institutional patterns may differ outside Europe., Language and translation effects not fully discussed — multilingual submissions could affect LLM performance.

Claims (9)

ClaimDirectionOutcomeConfidence & EvidenceDetails
LLM annotations of consultation submissions were highly reproducible across independent annotation runs. Other positive Reproducibility of LLM-generated annotations
Reading fidelity high
Study strength high
n=348
ICC(C,1) = 0.994–0.996; ICC(C,k) > 0.999
0.5
LLM-derived text measures showed weak convergence with corresponding survey-reported measures of AI safety, rights, and explainability concerns. Ai Safety And Ethics negative Construct correspondence between LLM text scores and survey-reported concerns
Reading fidelity high
Study strength high
n=348
Pearson r = 0.029–0.176; Lin's CCC = 0.013–0.080
0.5
The LLM-derived safety-concern measure had essentially zero correlation with the corresponding survey measure. Ai Safety And Ethics null_result Agreement between LLM-inferred and survey-reported AI safety concern
Reading fidelity high
Study strength high
n=348
r = 0.029; p = 0.59; Lin's CCC = 0.013
0.5
The LLM-derived rights-concern measure had only weak positive correlation with the corresponding survey measure. Ai Safety And Ethics positive Agreement between LLM-inferred and survey-reported rights concern
Reading fidelity high
Study strength medium
n=348
r = 0.176; p = 0.001; Lin's CCC = 0.080
0.3
The LLM-derived scores exhibited large standardized mean differences from the survey measures, indicating systematic bias across constructs. Ai Safety And Ethics negative Systematic measurement bias between LLM text scores and survey responses
Reading fidelity high
Study strength high
n=348
Cohen's d = 1.05, 0.98, and 1.18
0.5
Business associations expressed greater AI-risk concern in public consultation texts than in their survey responses. Ai Safety And Ethics positive Divergence between public text-based and survey-reported AI-risk concern
Reading fidelity high
Study strength medium
n=348
mean divergence = +1.0
0.3
Public authorities and several non-business stakeholder groups showed smaller or negative divergences between public text-based and survey-reported AI concerns. Ai Safety And Ethics negative Stakeholder-group differences in text–survey divergence of AI concerns
Reading fidelity high
Study strength medium
n=348
smaller or negative divergences
0.3
Divergences between text-based and survey-based scores exhibited positive spatial autocorrelation across European countries. Ai Safety And Ethics positive Geographic clustering of text–survey divergence in AI-safety stances
Reading fidelity high
Study strength medium
n=23
Moran's I = 0.347; p = 0.036
0.3
Survey-reported AI concerns remained strongly associated with support for explainability across levels of text–survey divergence. Decision Quality positive Association between survey-reported AI safety concern and support for explainability
Reading fidelity high
Study strength low
n=348
0.15

Notes