0 cumulative citations
View corpus contextAI-derived respondent-level measures can reproduce occupation-level and within-occupation associations for several job attributes in Chinese survey data when prompts use rich respondent and job features; but accuracy drops sharply with sparse information or when multiple related items are simultaneously unobserved.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
Researchers increasingly use artificial intelligence to construct measures of social, organizational, and occupational characteristics that are absent from conventional surveys. We propose AICOME, AI COntextual MEasurement, a framework for evaluating whether AI-derived respondent-level measures can recover individual and group-level effects in contextual models. The key idea is that an AI measure constructed at the respondent level can be used to derive its group-level aggregate and its individual deviation, allowing researchers to estimate both between-group and within-group associations rather than treating AI measurement as response prediction alone. We validate the framework using the 2022 China Family Panel Studies (CFPS), where occupations provide the empirical grouping structure and several job-related survey variables provide validation benchmarks. For computer use, foreign-language use, weekly hours, and management responsibilities, we compare survey measures with AI-derived measures in response-level, model-level, contextual, and boundary-condition validations. The results show that AI contextual measurement can recover much of the contextual-model information contained in observed survey variables when rich respondent and job characteristics are available. Weekly hours provides the strongest validation case, with AI-derived measures reproducing the large negative between- and within-occupation associations with satisfaction observed in CFPS. The framework also identifies clear boundary conditions: performance deteriorates when information is restricted to occupation and basic demographics, and recovery is weaker when several related concepts are treated as simultaneously unobserved. The findings suggest that AICOME is most useful for recovering a limited number of theoretically important constructs from rich existing datasets.
Summary
Main Finding
Respondent-level AI-derived measures (AICOME) can recover much of the between-group and within-group (contextual) information contained in conventional survey items — but only under realistic limits. When rich respondent- and job-level features are available and the researcher targets a small number of well-specified constructs, AI-predicted respondent scores reproduce key contextual associations (direction, explanatory power, fitted-values) seen in survey data. Performance collapses when features are sparse (e.g., only occupation + basic demographics) or when many related constructs are jointly unobserved. AICOME is therefore a useful tool for recovering a few theoretically important missing constructs from rich datasets, not a replacement for direct survey measurement or for reconstructing entire batteries of missing items.
Key Points
- Conceptual contribution: AICOME frames AI-derived respondent measures as alternative indicators of an underlying concept U and explicitly decomposes any respondent-level AI score Zki into a group (occupation) mean and an individual deviation:
- Zki = Z̄k,g(i) + (Zki − Z̄k,g(i))
- This allows estimation of both between-group (βG or βC) and within-group (βI) associations in contextual models.
- Validation emphasis: The primary validation target is recovery of contextual inference (signs, magnitudes, incremental R2, fitted-value similarity, and whether between- and within-group conclusions match survey-based analyses), not simply item-level prediction or marginal distribution matching.
- Empirical findings:
- Using 2022 China Family Panel Studies (CFPS) and occupation as the grouping variable, AI-derived respondent measures for weekly hours, computer use, foreign-language use, and management responsibility were validated.
- Weekly hours produced the strongest validation: AI measures (both "rich-prompt" and "survey-prompt") reproduced the large negative between- and within-occupation associations between weekly hours and job satisfaction observed in CFPS.
- Occupation-level AI scores (one score per occupation) missed within-occupation variation and therefore failed to recover within-group associations.
- Boundary conditions: recovery degrades when features are limited to occupation + basic demographics, and when several related constructs are simultaneously unobserved.
- Theoretical framing: The paper uses a proxy-validity assumption and linear-projection identities to interpret how R2 and incremental R2 relate across candidates (survey W, AI Z, true U). This clarifies that better R2 signals greater recovery of the outcome-relevant component of the underlying concept, but does not require the AI score to equal the survey item.
Data & Methods
- Data: 2022 China Family Panel Studies (CFPS). Grouping structure: occupation g(i). Outcome: job satisfaction Y. Validation items (survey indicators Wk): weekly hours, computer use, foreign-language use, management responsibilities. Later application also constructs latent dimensions (autonomy, people–things, creative–routine, technology) without direct CFPS validation.
- AI measurement:
- Respondent-level AI scores Zki constructed using a feature vector Fki drawn from the CFPS respondent and job attributes. Two prompting strategies:
- Rich-prompt: uses a richer questionnaire-style description of the target concept.
- Survey-prompt: concise prompt mirroring CFPS question wording.
- Group means Z̄k,g computed from respondent-level Zki; models use either raw parametrization (Zki and Z̄k,g) or deviation parametrization (Zki − Z̄k,g and Z̄k,g).
- Contextual models:
- Raw parameterization: Yi = Σk βIk Zki + Σk βCk Z̄k,g(i) + Xi'γ + εi
- Deviation parameterization: Yi = Σk βIk (Zki − Z̄k,g(i)) + Σk βGk Z̄k,g(i) + Xi'γ + εi
- Relation: βG = βI + βC, coefficients standardized so the identity holds across parameterizations.
- Controls Xi: age, gender, education, marital status, urban residence, agricultural hukou, public/state employment, outdoors job indicator, health, log income. (Weekly hours excluded from controls when it is the focal validated dimension.)
- Estimation and inference: standardized coefficients reported; cluster-robust standard errors at the occupation level.
- Validation procedure (four levels):
- Response-level: correlations and agreement between Z and observed W.
- Model-level: R2, incremental R2, fitted-value similarity, coefficient-vector correlations and RMSE relative to survey-based models.
- Contextual validation: recovery of between-group and within-group (individual-level) coefficient signs and magnitudes, and whether substantive contextual conclusions match.
- Boundary-condition analysis: test sensitivity to feature sparsity (Reduced-X) and to jointly missing multiple constructs (AI4 scenario).
- Theoretical diagnostics: use linear-projection identities (e.g., ∆R2_M = ∆R2_U · corr^2(b_YU|X, b_YM|X | X)) to interpret how much contextual signal an indicator recovers.
Implications for AI Economics
- New empirical leverage: AICOME expands the set of contextual questions that can be addressed in existing datasets by producing respondent-level AI measures that permit within-group heterogeneity analysis (e.g., workers within occupations). This complements group-level resources (e.g., O*NET) that are inherently limited to between-group variation.
- Practical recommendations for applied researchers:
- Use respondent-level AI scoring (not one-score-per-group) when the research question requires within-group variation.
- Validate AI-derived measures against survey items where available, focusing on contextual validation (incremental R2, coefficient signs, fitted-value similarity) rather than only response-level accuracy.
- Restrict AI recovery ambitions: target a small set of well-defined constructs and provide rich feature sets (demographic, job/task, prior survey items) to the model.
- Report sensitivity/boundary checks: show how estimates change under reduced feature sets and when multiple related constructs are jointly unobserved.
- Treat AI-derived measures as imperfect indicators — do not assume they replace direct measurement; account for possible measurement error in interpretation and inference.
- Econometric cautions:
- AICOME is conceptually closer to construct validation than to a generated-regressor correction: AI outputs need not converge to the survey item and both may be imperfect proxies for the true underlying concept U.
- Reliable downstream inference benefits from validation data; without it, conclusions are weaker because fitted-value similarity or coefficient agreement can fail even when marginal distributions match.
- Research opportunities:
- Use AICOME to augment studies of return-to-skill heterogeneity, task exposure, AI exposure, and other job/firm-level constructs by recovering individual deviations that matter for outcomes (wages, satisfaction, turnover).
- Combine AICOME with causal designs (e.g., IV, panel, difference-in-differences) to study whether between- vs within-group mechanisms drive causal effects, subject to validation checks.
- Extend the framework to other hierarchical settings (students in schools, patients in hospitals, residents in neighborhoods) where respondent-level AI measures could reconstruct contextual variation.
- Bottom line: AICOME is a useful, pragmatic method for reclaiming contextual information missing from surveys when rich auxiliary data are available, but it must be used with pre-registered validation diagnostics and awareness of its boundary conditions.
Assessment
Claims (10)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| AICOME constructs respondent-level AI measures and decomposes each measure into a group mean and an individual deviation, allowing contextual models to estimate both between-group and within-group associations. Other | positive | Recovery of individual-level and group-level contextual associations |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Direct group-level AI scores cannot identify within-group heterogeneity or within-group associations because all individuals in the same group receive the same score. Other | negative | Identification of within-group heterogeneity |
Reading fidelity
high
Study strength
high
|
not reported
|
| In the 2022 China Family Panel Studies application, AI contextual measurement recovered much of the contextual-model information contained in observed survey variables when rich respondent and job characteristics were available. Other | positive | Recovery of contextual-model information |
Reading fidelity
high
Study strength
medium
|
much of the contextual-model information
|
| Weekly hours was the strongest validation case: AI-derived measures reproduced the large negative between-occupation and within-occupation associations with job satisfaction observed in CFPS. Worker Satisfaction | negative | Job satisfaction |
Reading fidelity
high
Study strength
medium
|
large negative between- and within-occupation associations
|
| Occupation-level measures missed the individual-level signal that respondent-level AI measures recovered for weekly hours. Worker Satisfaction | negative | Individual-level association between weekly hours and job satisfaction |
Reading fidelity
high
Study strength
medium
|
not reported
|
| AICOME performance deteriorates when the available information is restricted to occupation and basic demographics. Other | negative | Accuracy of contextual inference recovery |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Recovery is weaker when several related concepts are treated as simultaneously unobserved. Other | negative | Recovery of contextual-model information under joint missingness |
Reading fidelity
high
Study strength
medium
|
not reported
|
| AICOME is most useful for recovering a limited number of theoretically important constructs from rich existing datasets, rather than replacing direct survey measurement or entire batteries of jointly missing items. Other | mixed | Appropriate use and scope of AI-derived measurement |
Reading fidelity
high
Study strength
medium
|
not reported
|
| The paper's primary validation target is recovery of contextual inference, not merely response-level agreement between AI-derived and survey measures. Other | positive | Similarity of contextual inferences from AI-derived and survey measures |
Reading fidelity
high
Study strength
medium
|
not reported
|
| The occupation-level application to autonomy, people-things, creative-routine, and technology is an application of the validated framework rather than direct evidence that these latent dimensions are true measures. Other | null_result | Direct validation of latent occupational dimensions |
Reading fidelity
high
Study strength
high
|
not reported
|