0 cumulative citations
View corpus contextAI turns measurement from a bottleneck into a choice problem: economists can cheaply generate many plausible variables from text and images, but reliable inference now depends on explicit construct definitions and careful validation (ideally randomized or well-designed validation samples).
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
Artificial intelligence (AI) is transforming measurement in economics. AI models convert unstructured data, such as text and images, into structured variables at low cost, making previously prohibitive measurement feasible at scale. This shifts the bottleneck from finding any scalable measure of a phenomenon to choosing among many plausible ones, which may support different empirical conclusions. This review provides guidance for navigating that shift. We describe three stages at which AI enters the measurement pipeline---discovery, construct definition, and observation---and what each demands of researchers. We argue that credible inference with AI-generated variables requires appropriately designed validation: anchoring measurement to explicit criteria, rather than informal claims that a proxy is reasonable. We then examine how validation samples support valid inference even when AI predictions are arbitrarily biased, and what can be done when a random validation sample is unavailable.
Summary
Main Finding
AI radically lowers the marginal cost of converting unstructured data into structured variables, shifting the empirical bottleneck from "finding any scalable measure" to "choosing among many plausible measures." Credible inference with AI-generated variables therefore depends less on informal claims that a proxy is reasonable and more on explicitly designed validation. Validation—anchoring measurement to observable criteria and learning prediction errors from a validation design—enables valid downstream inference even if model outputs are arbitrarily biased. A random validation sample is the gold standard; when unavailable, researchers must make explicit, testable assumptions about how models generalize and use partial-identification or sensitivity analyses.
Key Points
- Measurement as a three-stage pipeline:
- Discovery: which features to measure (theory-guided or data-driven hypothesis generation). AI can surface novel, interpretable features that researchers had not pre-specified.
- Construct definition: how to operationalize a concept. AI makes it cheap to generate multiple rubrics/operationalizations; domain theory and nomological networks are still essential to choose among them.
- Observation: implementing the operational definition at scale. AI (LLMs, vision models, etc.) provides high-throughput labeling but introduces systematic, opaque errors.
- The central challenge: abundance of plausible AI-generated measures and many researcher choices (rubrics, prompts, model, tuning data, thresholds) create risks of post-selection bias, irreproducibility, and “AI slop.”
- Validation is crucial and must be explicit: define observable criteria (labels) that operationalize the construct, and evaluate the AI-created variable against those criteria on held-out/validation data.
- Validation samples allow learning the mapping from AI output to true labels; if validation sampling is randomized relative to analysis sample, this supports corrected estimation and inference even when AI predictions are biased.
- If a random validation sample is unavailable, researchers must rely on assumptions about model behavior (e.g., covariate shift, label shift, invariance) that can sometimes be partially tested using model error patterns; otherwise use bounding/partial-identification and sensitivity analysis.
- Established frameworks are informative: common task framework for discovery/hypothesis generation; Cronbach & Meehl’s nomological network for construct validity; measurement-error and semiparametric inference literatures for observation-stage corrections.
- Practical research practices recommended: pre-specify constructs and validation plans, use held-out data for discovery and for validation, report alternative operationalizations and robustness checks, and document models/prompting/tuning and validation sampling to aid reproducibility.
Data & Methods
- Nature of the paper: conceptual review + technical synthesis focused on measurement and inference with AI-generated data (Dell & Rambachan, 2026).
- Illustrative applications discussed:
- LLMs coding open-ended surveys, historical documents, and web corpora.
- Computer vision extracting poverty, deforestation, pollution from satellite imagery.
- Hypothesis generation: discovering interpretable features in images, text, or waveforms that predict outcomes (examples: judicial decisions from mugshots; investigator performance from case notes; ECG waveform features).
- Methodological building blocks surveyed:
- Discovery tools: hypothesis generation procedures using LLMs, sparse autoencoders, morphing/latent traversals; evaluated with held-out data (common task framework).
- Construct-definition tools: rubrics, nomological networks (convergent/discriminant validity), theory-driven moment restrictions, and automated generation/comparison of alternative operationalizations.
- Observation-stage statistical methods:
- Validation-sample designs (random sampling as ideal).
- Error-correction methods that learn model error from validation labels to correct downstream estimators and standard errors.
- Approaches for nonrandom validation: explicit shift models (covariate/label shift), stability/invariance assumptions, partial identification and sensitivity analysis.
- Issues of post-selection inference when multiple measurements/choices are explored.
- Evidence and insights:
- Validation on held-out samples is the decisive standard for assessing discovery and observational validity.
- Model opacity implies researchers cannot rely on introspection into parameters; empirical validation is necessary.
- Reproducibility challenges when models change or are proprietary—documenting prompts, seeds, and validation data is essential.
Implications for AI Economics
- Research design:
- Budget for validation: allocate resources to create random/representative validation samples and document sampling protocols.
- Pre-specify construct definitions and validation criteria alongside analysis plans to limit researcher degrees of freedom and post-selection bias.
- When feasible, report multiple operationalizations and report how substantive conclusions vary across them.
- Inference and estimation:
- Use validation-based error-correction procedures to obtain unbiased or partially identified estimates and valid standard errors when incorporating AI-generated variables.
- If random validation is infeasible, state and defend explicit assumptions about model generalization; accompany estimates with robustness/sensitivity bounds.
- Institutional and policy use:
- Official statistics and policy research using AI-derived measures should require documented validation against explicit criteria and preferably independent validation samples.
- Maintain archives of validation labels, prompts/rubrics, and model versions where possible to improve reproducibility and auditability.
- Open research frontiers:
- Developing methods for inference under complex, nonrandom validation designs and for bounding effects under plausible model-shift scenarios.
- Better diagnostics and empirical characterizations of how modern models err across settings to inform credible assumptions.
- Standards for reporting AI-measurement pipelines (construct rubrics, validation sampling, model details) analogous to established standards in causal inference.
- Broader point: AI expands possibilities for empirical work but increases the need for disciplined measurement practice. The credibility revolution’s emphasis on design, transparency, and robustness applies upstream to measurement choices in the AI era.
Assessment
Claims (11)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| AI models can convert unstructured data, including text, images, audio, and video, into low-dimensional structured variables for economic analysis at very low marginal cost. Organizational Efficiency | positive | Cost and scalability of economic measurement |
Reading fidelity
high
Study strength
medium
|
not reported
|
| AI makes measurement endeavors that were previously prohibitively costly feasible at scale and enables economists to study questions that were previously out of reach. Research Productivity | positive | Feasibility and scale of economic research measurement |
Reading fidelity
high
Study strength
medium
|
not reported
|
| The availability of AI shifts the bottleneck in empirical measurement from finding any scalable measure of a phenomenon to choosing among many plausible measures. Decision Quality | mixed | Choice and credibility of measurement variables |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Different measurement functions applied to the same underlying concept can preserve different features of the data and therefore support different empirical conclusions. Decision Quality | mixed | Empirical conclusions produced by alternative measurements |
Reading fidelity
high
Study strength
medium
|
not reported
|
| The abundance of AI-generated measures creates risks of noisy, poorly validated, or difficult-to-interpret variables, referred to by the authors as “AI slop.” Error Rate | negative | Quality and interpretability of AI-generated variables |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Credible inference using AI-generated variables requires measurement validation anchored to explicit, observable criteria rather than informal claims that a proxy is reasonable. Decision Quality | positive | Credibility and validity of inference using AI-generated variables |
Reading fidelity
high
Study strength
high
|
not reported
|
| A validation sample can support valid inference even when AI predictions are arbitrarily biased, because prediction errors can be learned from the validation design rather than inferred from assumptions about how the model generates its outputs. Decision Quality | positive | Validity of statistical inference with biased AI predictions |
Reading fidelity
high
Study strength
high
|
not reported
|
| When a random validation sample is unavailable, the credibility of inference depends on assumptions about how AI models behave across settings rather than on the validation-sample design. Decision Quality | negative | Credibility of inference without random validation data |
Reading fidelity
high
Study strength
medium
|
not reported
|
| AI can improve the discovery stage by surfacing patterns that researchers did not specify in advance, but evidence for a generated hypothesis must come from held-out data. Innovation Output | positive | Discovery of predictive or economically meaningful patterns |
Reading fidelity
high
Study strength
medium
|
not reported
|
| AI lowers the cost of generating and comparing alternative operationalizations of a concept, but it cannot determine which operationalization answers the research question; that judgment requires theory and substantive expertise. Decision Quality | mixed | Quality and appropriateness of construct definitions |
Reading fidelity
high
Study strength
medium
|
not reported
|
| For the discovery stage, the quality of a generated hypothesis should be evaluated on held-out data by measuring the named construct and assessing how well it predicts the outcome of interest. Output Quality | positive | Out-of-sample predictive performance of generated hypotheses |
Reading fidelity
high
Study strength
high
|
not reported
|