The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

AI-derived respondent-level measures can reproduce occupation-level and within-occupation associations for several job attributes in Chinese survey data when prompts use rich respondent and job features; but accuracy drops sharply with sparse information or when multiple related items are simultaneously unobserved.

AI Contextual Measurement for Recovering Individual and Group-Level Effects: Validation Against Survey Measures and an Occupational Application
Wenxin Jiang, Xuyang Wang, Yuxiao Wu · September 02, 2026
arxiv correlational medium evidence 7/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

Arxiv

Latest observation:

  1. Wenxin Jiang unresolved corpus identity
  2. Xuyang Wang unresolved corpus identity
  3. Yuxiao Wu unresolved corpus identity

Semantic Scholar

Latest observation:

  1. Wen-Xin Jiang unresolved corpus identity
  2. Xu-Yang Wang unresolved corpus identity
  3. Yu-Xiao Wu unresolved corpus identity
AICOME shows that respondent-level AI-derived measures, when fed rich respondent and job features, can recover both between- and within-occupation associations for several job-related constructs in CFPS, though performance falls with sparse information or multiple jointly missing constructs.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

Researchers increasingly use artificial intelligence to construct measures of social, organizational, and occupational characteristics that are absent from conventional surveys. We propose AICOME, AI COntextual MEasurement, a framework for evaluating whether AI-derived respondent-level measures can recover individual and group-level effects in contextual models. The key idea is that an AI measure constructed at the respondent level can be used to derive its group-level aggregate and its individual deviation, allowing researchers to estimate both between-group and within-group associations rather than treating AI measurement as response prediction alone. We validate the framework using the 2022 China Family Panel Studies (CFPS), where occupations provide the empirical grouping structure and several job-related survey variables provide validation benchmarks. For computer use, foreign-language use, weekly hours, and management responsibilities, we compare survey measures with AI-derived measures in response-level, model-level, contextual, and boundary-condition validations. The results show that AI contextual measurement can recover much of the contextual-model information contained in observed survey variables when rich respondent and job characteristics are available. Weekly hours provides the strongest validation case, with AI-derived measures reproducing the large negative between- and within-occupation associations with satisfaction observed in CFPS. The framework also identifies clear boundary conditions: performance deteriorates when information is restricted to occupation and basic demographics, and recovery is weaker when several related concepts are treated as simultaneously unobserved. The findings suggest that AICOME is most useful for recovering a limited number of theoretically important constructs from rich existing datasets.

Summary

Main Finding

Respondent-level AI-derived measures (AICOME) can recover much of the between-group and within-group (contextual) information contained in conventional survey items — but only under realistic limits. When rich respondent- and job-level features are available and the researcher targets a small number of well-specified constructs, AI-predicted respondent scores reproduce key contextual associations (direction, explanatory power, fitted-values) seen in survey data. Performance collapses when features are sparse (e.g., only occupation + basic demographics) or when many related constructs are jointly unobserved. AICOME is therefore a useful tool for recovering a few theoretically important missing constructs from rich datasets, not a replacement for direct survey measurement or for reconstructing entire batteries of missing items.

Key Points

  • Conceptual contribution: AICOME frames AI-derived respondent measures as alternative indicators of an underlying concept U and explicitly decomposes any respondent-level AI score Zki into a group (occupation) mean and an individual deviation:
    • Zki = Z̄k,g(i) + (Zki − Z̄k,g(i))
    • This allows estimation of both between-group (βG or βC) and within-group (βI) associations in contextual models.
  • Validation emphasis: The primary validation target is recovery of contextual inference (signs, magnitudes, incremental R2, fitted-value similarity, and whether between- and within-group conclusions match survey-based analyses), not simply item-level prediction or marginal distribution matching.
  • Empirical findings:
    • Using 2022 China Family Panel Studies (CFPS) and occupation as the grouping variable, AI-derived respondent measures for weekly hours, computer use, foreign-language use, and management responsibility were validated.
    • Weekly hours produced the strongest validation: AI measures (both "rich-prompt" and "survey-prompt") reproduced the large negative between- and within-occupation associations between weekly hours and job satisfaction observed in CFPS.
    • Occupation-level AI scores (one score per occupation) missed within-occupation variation and therefore failed to recover within-group associations.
    • Boundary conditions: recovery degrades when features are limited to occupation + basic demographics, and when several related constructs are simultaneously unobserved.
  • Theoretical framing: The paper uses a proxy-validity assumption and linear-projection identities to interpret how R2 and incremental R2 relate across candidates (survey W, AI Z, true U). This clarifies that better R2 signals greater recovery of the outcome-relevant component of the underlying concept, but does not require the AI score to equal the survey item.

Data & Methods

  • Data: 2022 China Family Panel Studies (CFPS). Grouping structure: occupation g(i). Outcome: job satisfaction Y. Validation items (survey indicators Wk): weekly hours, computer use, foreign-language use, management responsibilities. Later application also constructs latent dimensions (autonomy, people–things, creative–routine, technology) without direct CFPS validation.
  • AI measurement:
    • Respondent-level AI scores Zki constructed using a feature vector Fki drawn from the CFPS respondent and job attributes. Two prompting strategies:
    • Rich-prompt: uses a richer questionnaire-style description of the target concept.
    • Survey-prompt: concise prompt mirroring CFPS question wording.
    • Group means Z̄k,g computed from respondent-level Zki; models use either raw parametrization (Zki and Z̄k,g) or deviation parametrization (Zki − Z̄k,g and Z̄k,g).
  • Contextual models:
    • Raw parameterization: Yi = Σk βIk Zki + Σk βCk Z̄k,g(i) + Xi'γ + εi
    • Deviation parameterization: Yi = Σk βIk (Zki − Z̄k,g(i)) + Σk βGk Z̄k,g(i) + Xi'γ + εi
    • Relation: βG = βI + βC, coefficients standardized so the identity holds across parameterizations.
  • Controls Xi: age, gender, education, marital status, urban residence, agricultural hukou, public/state employment, outdoors job indicator, health, log income. (Weekly hours excluded from controls when it is the focal validated dimension.)
  • Estimation and inference: standardized coefficients reported; cluster-robust standard errors at the occupation level.
  • Validation procedure (four levels):
  • Response-level: correlations and agreement between Z and observed W.
  • Model-level: R2, incremental R2, fitted-value similarity, coefficient-vector correlations and RMSE relative to survey-based models.
  • Contextual validation: recovery of between-group and within-group (individual-level) coefficient signs and magnitudes, and whether substantive contextual conclusions match.
  • Boundary-condition analysis: test sensitivity to feature sparsity (Reduced-X) and to jointly missing multiple constructs (AI4 scenario).
  • Theoretical diagnostics: use linear-projection identities (e.g., ∆R2_M = ∆R2_U · corr^2(b_YU|X, b_YM|X | X)) to interpret how much contextual signal an indicator recovers.

Implications for AI Economics

  • New empirical leverage: AICOME expands the set of contextual questions that can be addressed in existing datasets by producing respondent-level AI measures that permit within-group heterogeneity analysis (e.g., workers within occupations). This complements group-level resources (e.g., O*NET) that are inherently limited to between-group variation.
  • Practical recommendations for applied researchers:
    • Use respondent-level AI scoring (not one-score-per-group) when the research question requires within-group variation.
    • Validate AI-derived measures against survey items where available, focusing on contextual validation (incremental R2, coefficient signs, fitted-value similarity) rather than only response-level accuracy.
    • Restrict AI recovery ambitions: target a small set of well-defined constructs and provide rich feature sets (demographic, job/task, prior survey items) to the model.
    • Report sensitivity/boundary checks: show how estimates change under reduced feature sets and when multiple related constructs are jointly unobserved.
    • Treat AI-derived measures as imperfect indicators — do not assume they replace direct measurement; account for possible measurement error in interpretation and inference.
  • Econometric cautions:
    • AICOME is conceptually closer to construct validation than to a generated-regressor correction: AI outputs need not converge to the survey item and both may be imperfect proxies for the true underlying concept U.
    • Reliable downstream inference benefits from validation data; without it, conclusions are weaker because fitted-value similarity or coefficient agreement can fail even when marginal distributions match.
  • Research opportunities:
    • Use AICOME to augment studies of return-to-skill heterogeneity, task exposure, AI exposure, and other job/firm-level constructs by recovering individual deviations that matter for outcomes (wages, satisfaction, turnover).
    • Combine AICOME with causal designs (e.g., IV, panel, difference-in-differences) to study whether between- vs within-group mechanisms drive causal effects, subject to validation checks.
    • Extend the framework to other hierarchical settings (students in schools, patients in hospitals, residents in neighborhoods) where respondent-level AI measures could reconstruct contextual variation.
  • Bottom line: AICOME is a useful, pragmatic method for reclaiming contextual information missing from surveys when rich auxiliary data are available, but it must be used with pre-registered validation diagnostics and awareness of its boundary conditions.

Assessment

Paper Typecorrelational Evidence Strengthmedium — The paper provides thorough within-sample validation against multiple CFPS survey items, demonstrating that respondent-level AI measures can reproduce between- and within-occupation associations for several constructs (especially weekly hours). However, evidence is limited to one survey (China CFPS 2022), is observational, depends on the richness of available features and prompt choices, and does not establish external or causal validity. Methods Rigormedium — The authors present a clear conceptual framework (AICOME), use sensible multilevel/deviation regressions, report standardized coefficients, cluster-robust SEs, and several validation diagnostics (response-level agreement, model-level R2/fitted-value similarity, contextual decomposition, and boundary-condition tests). Limitations include reliance on a single dataset/country, potential sensitivity to prompt/model choice and feature set (possible leakage or overfitting to available covariates), and limited discussion of inference adjustments for generated regressors or model uncertainty. SampleEmpirical validation uses the 2022 China Family Panel Studies (CFPS); occupations define groups. Outcome is job satisfaction; validated target constructs include computer use, foreign-language use, weekly hours, and management responsibilities. Controls include age, gender, education, marital status, urban residence, agricultural hukou, public/state employment indicator, outdoor work indicator, health, and log income. The text does not report the sample size in the excerpt. Themeshuman_ai_collab labor_markets IdentificationNo claim to causal identification; uses cross-sectional contextual (multilevel) regressions decomposing respondent-level AI-derived measures into group means and within-group deviations, with controls (age, gender, education, hukou, sector, job characteristics, income, health) and occupation-clustered standard errors; validation is achieved by comparing coefficients, R2, fitted values, and coefficient-vector similarity between AI-derived indicators and observed survey measures. GeneralizabilitySingle-country sample (China CFPS) — may not generalize to other countries or survey designs, Performance depends on richness of respondent- and job-level feature set (may deteriorate with sparse data), Depends on prompt protocol, LLM model/version, and access to features used in prompts (sensitivity to implementation), Validated for a small set of job-related constructs; not validated for many other latent dimensions or for simultaneous recovery of multiple related constructs, Observational design — cannot establish causal effects; intended for measurement/recovery, not causal identification

Claims (10)

ClaimDirectionOutcomeConfidence & EvidenceDetails
AICOME constructs respondent-level AI measures and decomposes each measure into a group mean and an individual deviation, allowing contextual models to estimate both between-group and within-group associations. Other positive Recovery of individual-level and group-level contextual associations
Reading fidelity high
Study strength medium
not reported
0.3
Direct group-level AI scores cannot identify within-group heterogeneity or within-group associations because all individuals in the same group receive the same score. Other negative Identification of within-group heterogeneity
Reading fidelity high
Study strength high
not reported
0.5
In the 2022 China Family Panel Studies application, AI contextual measurement recovered much of the contextual-model information contained in observed survey variables when rich respondent and job characteristics were available. Other positive Recovery of contextual-model information
Reading fidelity high
Study strength medium
much of the contextual-model information
0.3
Weekly hours was the strongest validation case: AI-derived measures reproduced the large negative between-occupation and within-occupation associations with job satisfaction observed in CFPS. Worker Satisfaction negative Job satisfaction
Reading fidelity high
Study strength medium
large negative between- and within-occupation associations
0.3
Occupation-level measures missed the individual-level signal that respondent-level AI measures recovered for weekly hours. Worker Satisfaction negative Individual-level association between weekly hours and job satisfaction
Reading fidelity high
Study strength medium
not reported
0.3
AICOME performance deteriorates when the available information is restricted to occupation and basic demographics. Other negative Accuracy of contextual inference recovery
Reading fidelity high
Study strength medium
not reported
0.3
Recovery is weaker when several related concepts are treated as simultaneously unobserved. Other negative Recovery of contextual-model information under joint missingness
Reading fidelity high
Study strength medium
not reported
0.3
AICOME is most useful for recovering a limited number of theoretically important constructs from rich existing datasets, rather than replacing direct survey measurement or entire batteries of jointly missing items. Other mixed Appropriate use and scope of AI-derived measurement
Reading fidelity high
Study strength medium
not reported
0.3
The paper's primary validation target is recovery of contextual inference, not merely response-level agreement between AI-derived and survey measures. Other positive Similarity of contextual inferences from AI-derived and survey measures
Reading fidelity high
Study strength medium
not reported
0.3
The occupation-level application to autonomy, people-things, creative-routine, and technology is an application of the validated framework rather than direct evidence that these latent dimensions are true measures. Other null_result Direct validation of latent occupational dimensions
Reading fidelity high
Study strength high
not reported
0.5

Notes