0 cumulative citations
View corpus contextLLM labels of EU AI Act submissions are nearly perfectly reproducible but fail to match what stakeholders say in surveys; businesses tend to amplify AI-risk language in public consultations while regulators and some non-business groups do not.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
0 cumulative citations
View corpus contextHigh annotation reproducibility does not necessarily imply that an LLM-inferred measure captures the construct it is intended to measure. We test this distinction using a dataset from the European Commission's AI Act consultation, linking structured survey responses to free-text consultation submissions from the same stakeholders. LLM annotations of consultation submissions are highly reproducible (intraclass correlations > 0.99), yet show limited convergence with survey-reported measures of the nominal construct they were intended to approximate. Divergence between survey-and LLM-inferred text-based measures varies systematically across stakeholder groups: business associations express greater concern about AI risks in text-based consultations than in survey responses ({g} = +1.0), whereas public authorities and several nonbusiness groups show smaller or negative divergences. Divergences between scores suggest positive spatial autocorrelation across European countries (Moran's I = 0.347, p = 0.036), indicating that stakeholders from neighboring countries tend toward more similar text-based stances towards AI safety concerns. Despite divergence, survey-reported concerns remain strongly associated with support for explainability across all divergence levels. These results demonstrate that LLM annotation reproducibility can coexist with poor construct correspondence and motivate validation procedures that distinguish reproducibility, construct validity, and communication context variation when LLMs are used as measurement instruments.
Summary
Main Finding
LLM annotations of public consultation texts can be extremely reproducible while failing to validly recover the same latent constructs measured in surveys. In the European Commission AI Act consultation, Qwen3.5 annotations showed near-perfect inter-run reproducibility (ICC > 0.99) but weak convergence with stakeholders’ survey-reported AI concerns (Pearson r ≈ 0.03–0.18; Lin’s CCC < 0.08). Divergences are systematic across institutional and geographic contexts (e.g., business associations amplify AI-risk language in public submissions relative to their survey responses).
Key Points
- Reproducibility ≠ construct validity: High LLM inter-run agreement (ICC(C,1) 0.994–0.996; ICC(C,k) > 0.999) did not imply that the LLM-derived scores measured the same latent constructs elicited by surveys.
- Poor convergence with surveys:
- Safety concern: r = 0.029 (p = 0.59)
- Rights concern: r = 0.176 (p = 0.001)
- Explainability: r = 0.097 (p = 0.072)
- Lin’s concordance coefficients near zero (0.013–0.080).
- Large systematic bias: Bland–Altman analysis showed limits of agreement spanning ~±4 points on a 0–5 scale; standardized mean differences (Cohen’s d) ≈ 1.0–1.18.
- Systematic heterogeneity by stakeholder type:
- Business associations: positive divergence (public texts show higher AI-risk emphasis than their surveys; mean gap ≈ +1.0).
- Public authorities and several non-business groups: smaller or negative divergences.
- Geographic patterning: positive spatial autocorrelation of divergence (Moran’s I = 0.347, p = 0.036), though the authors caution sensitivity to multiple-testing and sample limitations.
- Downstream robustness: despite divergence, survey-reported concern remains strongly associated with support for explainability across divergence levels (tested with Double/Debiased Machine Learning), indicating distinct but still informative relationships.
- Conceptual caution: public consultation texts are institutionally situated artifacts (audience targeting, reputation management, access channels). Treating text-derived proxies as error-free measures of private or survey-attested attitudes can induce collider or selection biases in downstream analyses (Text is a collider: T → Text ← Y).
Data & Methods
- Data: Public submissions to the European Commission AI Regulation consultation (Better Regulation Portal), N = 857 submissions collected 2020-02-20 to 2021-04-27 across 3 rounds. Validation sample: N = 348 submissions with both free-text and matched structured survey items (0–5 scales for safety concern, rights concern, explainability).
- LLM annotation:
- Model: Qwen3.5-397b-A17b.
- Task: annotate each document on three continuous 0–5 scales (safety_concern, rights_concern, explainability_trust); prompt enforced strict JSON and included scale anchors and binary flags (e.g., regulatory, biometric mentions).
- Inter-run checks: five independent annotation runs; near-identical aggregate scores (mean vs median difference < 0.01).
- Validation diagnostics:
- Convergent/discriminant checks: Pearson correlation, Lin’s CCC, Bland–Altman limits, Cohen’s d, ICCs for cross-setting agreement.
- Examination of systematic divergence across institutional groups and countries.
- Spatial analysis:
- Moran’s I with queen contiguity weights (shared land borders), permutation inference (B = 9,999), primary sample of 23 countries.
- Causal/robustness check:
- Double/Debiased Machine Learning (DML) to estimate association between high survey-reported safety concern and explainability support.
- Implementations: LightGBM for nuisance estimation with 5-fold cross-fitting; Causal Forest DML parameters reported (Ntrees = 1000, minleaf = 10).
- Conceptual framing: Directed acyclic graph showing Text as a downstream collider of multiple latent constructs (T and Y) and unobserved confounders U — conditioning on text-derived features can open spurious paths.
Implications for AI Economics
- Measurement risks in text-based variables: LLM-derived textual measures may reflect public-facing, institutionally mediated communication (framing, signaling, audience targeting) rather than private or survey-attested preferences. Econometric analyses that treat such measures as direct proxies for latent attitudes risk mis-specification and biased estimates.
- Collider and omitted-variable bias: Using text-derived features as covariates/controls without validating construct correspondence can induce collider stratification, opening spurious associations between explanatory variables and outcomes. Researchers should explicitly model the measurement process and possible downstream selection/collider structures.
- Policy inference and stakeholder analysis: Public statements scraped from consultations or other regulatory submissions can systematically over- or understate concerns by stakeholder type (e.g., businesses amplify risk rhetorically). Policymakers and analysts should not equate public rhetorical intensity with private preferences or likelihood to support/comply with regulation.
- Cross-country and institutional heterogeneity matter: Cultural, reputational, and access-channel differences shape public communication. Comparative AI-economics work that uses text measures must test for—and, where appropriate, model—heterogeneous measurement error across organizations and countries.
- Recommended empirical practices:
- Validate LLM text measures against external criteria (matched surveys, behavioral signals) before using them as proxies in causal models.
- Distinguish measurement targets: if the interest is public communication/positioning, LLM scores can be valid indicators; if the interest is underlying private attitudes, additional validation and correction are required.
- Report both reproducibility and construct validity metrics separately; high inter-run agreement alone is insufficient.
- Conduct subgroup analyses (by stakeholder type, country) and sensitivity tests to detect systematic divergence.
- Where text-derived measures are used as controls, consider causal identification strategies robust to measurement-induced collider bias (e.g., instrumenting, explicit measurement-error models, bounding approaches).
- Combine sources (surveys, private channels, behavioral data) and exploit methods like DML while maintaining external validation of the target construct.
- Practical takeaway for AI-economics researchers: LLMs are powerful scalers of text, but their outputs should be interpreted as indicators of communicative expression in context. Use them thoughtfully—validate, test heterogeneity, and avoid treating reproducible LLM annotations as automatically valid measures of latent economic or policy attitudes.
Assessment
Claims (9)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| LLM annotations of consultation submissions were highly reproducible across independent annotation runs. Other | positive | Reproducibility of LLM-generated annotations |
Reading fidelity
high
Study strength
high
|
n=348
ICC(C,1) = 0.994–0.996; ICC(C,k) > 0.999
|
| LLM-derived text measures showed weak convergence with corresponding survey-reported measures of AI safety, rights, and explainability concerns. Ai Safety And Ethics | negative | Construct correspondence between LLM text scores and survey-reported concerns |
Reading fidelity
high
Study strength
high
|
n=348
Pearson r = 0.029–0.176; Lin's CCC = 0.013–0.080
|
| The LLM-derived safety-concern measure had essentially zero correlation with the corresponding survey measure. Ai Safety And Ethics | null_result | Agreement between LLM-inferred and survey-reported AI safety concern |
Reading fidelity
high
Study strength
high
|
n=348
r = 0.029; p = 0.59; Lin's CCC = 0.013
|
| The LLM-derived rights-concern measure had only weak positive correlation with the corresponding survey measure. Ai Safety And Ethics | positive | Agreement between LLM-inferred and survey-reported rights concern |
Reading fidelity
high
Study strength
medium
|
n=348
r = 0.176; p = 0.001; Lin's CCC = 0.080
|
| The LLM-derived scores exhibited large standardized mean differences from the survey measures, indicating systematic bias across constructs. Ai Safety And Ethics | negative | Systematic measurement bias between LLM text scores and survey responses |
Reading fidelity
high
Study strength
high
|
n=348
Cohen's d = 1.05, 0.98, and 1.18
|
| Business associations expressed greater AI-risk concern in public consultation texts than in their survey responses. Ai Safety And Ethics | positive | Divergence between public text-based and survey-reported AI-risk concern |
Reading fidelity
high
Study strength
medium
|
n=348
mean divergence = +1.0
|
| Public authorities and several non-business stakeholder groups showed smaller or negative divergences between public text-based and survey-reported AI concerns. Ai Safety And Ethics | negative | Stakeholder-group differences in text–survey divergence of AI concerns |
Reading fidelity
high
Study strength
medium
|
n=348
smaller or negative divergences
|
| Divergences between text-based and survey-based scores exhibited positive spatial autocorrelation across European countries. Ai Safety And Ethics | positive | Geographic clustering of text–survey divergence in AI-safety stances |
Reading fidelity
high
Study strength
medium
|
n=23
Moran's I = 0.347; p = 0.036
|
| Survey-reported AI concerns remained strongly associated with support for explainability across levels of text–survey divergence. Decision Quality | positive | Association between survey-reported AI safety concern and support for explainability |
Reading fidelity
high
Study strength
low
|
n=348
|