The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

AI turns measurement from a bottleneck into a choice problem: economists can cheaply generate many plausible variables from text and images, but reliable inference now depends on explicit construct definitions and careful validation (ideally randomized or well-designed validation samples).

The Measurement Revolution? Credible Measurement and Inference in the Age of AI
Melissa Dell, Ashesh Rambachan · August 24, 2026
arxiv review_meta n/a evidence 8/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

Arxiv

Latest observation:

  1. Melissa Dell unresolved corpus identity
  2. Ashesh Rambachan unresolved corpus identity

Semantic Scholar

Latest observation:

  1. Melissa Dell provider ID
  2. Ashesh Rambachan unresolved corpus identity
AI greatly expands scalable measurement from unstructured data, shifting the empirical bottleneck from obtaining measures to selecting and validating among many plausible AI-generated measures, and credible inference requires explicit construct definitions and principled validation designs.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

Artificial intelligence (AI) is transforming measurement in economics. AI models convert unstructured data, such as text and images, into structured variables at low cost, making previously prohibitive measurement feasible at scale. This shifts the bottleneck from finding any scalable measure of a phenomenon to choosing among many plausible ones, which may support different empirical conclusions. This review provides guidance for navigating that shift. We describe three stages at which AI enters the measurement pipeline---discovery, construct definition, and observation---and what each demands of researchers. We argue that credible inference with AI-generated variables requires appropriately designed validation: anchoring measurement to explicit criteria, rather than informal claims that a proxy is reasonable. We then examine how validation samples support valid inference even when AI predictions are arbitrarily biased, and what can be done when a random validation sample is unavailable.

Summary

Main Finding

AI radically lowers the marginal cost of converting unstructured data into structured variables, shifting the empirical bottleneck from "finding any scalable measure" to "choosing among many plausible measures." Credible inference with AI-generated variables therefore depends less on informal claims that a proxy is reasonable and more on explicitly designed validation. Validation—anchoring measurement to observable criteria and learning prediction errors from a validation design—enables valid downstream inference even if model outputs are arbitrarily biased. A random validation sample is the gold standard; when unavailable, researchers must make explicit, testable assumptions about how models generalize and use partial-identification or sensitivity analyses.

Key Points

  • Measurement as a three-stage pipeline:
    • Discovery: which features to measure (theory-guided or data-driven hypothesis generation). AI can surface novel, interpretable features that researchers had not pre-specified.
    • Construct definition: how to operationalize a concept. AI makes it cheap to generate multiple rubrics/operationalizations; domain theory and nomological networks are still essential to choose among them.
    • Observation: implementing the operational definition at scale. AI (LLMs, vision models, etc.) provides high-throughput labeling but introduces systematic, opaque errors.
  • The central challenge: abundance of plausible AI-generated measures and many researcher choices (rubrics, prompts, model, tuning data, thresholds) create risks of post-selection bias, irreproducibility, and “AI slop.”
  • Validation is crucial and must be explicit: define observable criteria (labels) that operationalize the construct, and evaluate the AI-created variable against those criteria on held-out/validation data.
  • Validation samples allow learning the mapping from AI output to true labels; if validation sampling is randomized relative to analysis sample, this supports corrected estimation and inference even when AI predictions are biased.
  • If a random validation sample is unavailable, researchers must rely on assumptions about model behavior (e.g., covariate shift, label shift, invariance) that can sometimes be partially tested using model error patterns; otherwise use bounding/partial-identification and sensitivity analysis.
  • Established frameworks are informative: common task framework for discovery/hypothesis generation; Cronbach & Meehl’s nomological network for construct validity; measurement-error and semiparametric inference literatures for observation-stage corrections.
  • Practical research practices recommended: pre-specify constructs and validation plans, use held-out data for discovery and for validation, report alternative operationalizations and robustness checks, and document models/prompting/tuning and validation sampling to aid reproducibility.

Data & Methods

  • Nature of the paper: conceptual review + technical synthesis focused on measurement and inference with AI-generated data (Dell & Rambachan, 2026).
  • Illustrative applications discussed:
    • LLMs coding open-ended surveys, historical documents, and web corpora.
    • Computer vision extracting poverty, deforestation, pollution from satellite imagery.
    • Hypothesis generation: discovering interpretable features in images, text, or waveforms that predict outcomes (examples: judicial decisions from mugshots; investigator performance from case notes; ECG waveform features).
  • Methodological building blocks surveyed:
    • Discovery tools: hypothesis generation procedures using LLMs, sparse autoencoders, morphing/latent traversals; evaluated with held-out data (common task framework).
    • Construct-definition tools: rubrics, nomological networks (convergent/discriminant validity), theory-driven moment restrictions, and automated generation/comparison of alternative operationalizations.
    • Observation-stage statistical methods:
      • Validation-sample designs (random sampling as ideal).
      • Error-correction methods that learn model error from validation labels to correct downstream estimators and standard errors.
      • Approaches for nonrandom validation: explicit shift models (covariate/label shift), stability/invariance assumptions, partial identification and sensitivity analysis.
      • Issues of post-selection inference when multiple measurements/choices are explored.
  • Evidence and insights:
    • Validation on held-out samples is the decisive standard for assessing discovery and observational validity.
    • Model opacity implies researchers cannot rely on introspection into parameters; empirical validation is necessary.
    • Reproducibility challenges when models change or are proprietary—documenting prompts, seeds, and validation data is essential.

Implications for AI Economics

  • Research design:
    • Budget for validation: allocate resources to create random/representative validation samples and document sampling protocols.
    • Pre-specify construct definitions and validation criteria alongside analysis plans to limit researcher degrees of freedom and post-selection bias.
    • When feasible, report multiple operationalizations and report how substantive conclusions vary across them.
  • Inference and estimation:
    • Use validation-based error-correction procedures to obtain unbiased or partially identified estimates and valid standard errors when incorporating AI-generated variables.
    • If random validation is infeasible, state and defend explicit assumptions about model generalization; accompany estimates with robustness/sensitivity bounds.
  • Institutional and policy use:
    • Official statistics and policy research using AI-derived measures should require documented validation against explicit criteria and preferably independent validation samples.
    • Maintain archives of validation labels, prompts/rubrics, and model versions where possible to improve reproducibility and auditability.
  • Open research frontiers:
    • Developing methods for inference under complex, nonrandom validation designs and for bounding effects under plausible model-shift scenarios.
    • Better diagnostics and empirical characterizations of how modern models err across settings to inform credible assumptions.
    • Standards for reporting AI-measurement pipelines (construct rubrics, validation sampling, model details) analogous to established standards in causal inference.
  • Broader point: AI expands possibilities for empirical work but increases the need for disciplined measurement practice. The credibility revolution’s emphasis on design, transparency, and robustness applies upstream to measurement choices in the AI era.

Assessment

Paper Typereview_meta Evidence Strengthn/a — This is a methodological review and synthesis rather than a primary empirical or causal-identification paper; it collects and interprets existing literature and examples rather than producing new causal estimates requiring identification. Methods Rigorn/a — The paper is a high-quality conceptual and methodological synthesis (NBER Methods Lecture) rather than an empirical study with an applied identification strategy; it rigorously organizes literatures and lays out formal considerations but does not itself implement new empirical identification. SampleNo original sample or primary dataset; the paper synthesizes and reviews evidence from a broad set of literatures (econometrics, statistics, machine learning, psychometrics, computational social science) and illustrative empirical examples (e.g., LLM coding of text, computer-vision measures from satellite imagery, cited empirical papers). Themesinnovation adoption GeneralizabilityGuidance is conceptual and methodological rather than a one-size-fits-all recipe; applicability depends on context, data modality, and research question., Emphasis on the observation stage and statistical validation leaves some discovery/construct-definition problems less fully operationalized., Recommendations often assume some access to validation samples or plausibly testable model-behavior assumptions, which may not hold in settings with proprietary models or no labeled data., Does not provide modality-specific engineering details; practical performance may vary across text, image, audio, and video applications., Rapid model and platform turnover (proprietary model deprecation) may limit reproducibility of specific measurement pipelines discussed.

Claims (11)

ClaimDirectionOutcomeConfidence & EvidenceDetails
AI models can convert unstructured data, including text, images, audio, and video, into low-dimensional structured variables for economic analysis at very low marginal cost. Organizational Efficiency positive Cost and scalability of economic measurement
Reading fidelity high
Study strength medium
not reported
0.24
AI makes measurement endeavors that were previously prohibitively costly feasible at scale and enables economists to study questions that were previously out of reach. Research Productivity positive Feasibility and scale of economic research measurement
Reading fidelity high
Study strength medium
not reported
0.24
The availability of AI shifts the bottleneck in empirical measurement from finding any scalable measure of a phenomenon to choosing among many plausible measures. Decision Quality mixed Choice and credibility of measurement variables
Reading fidelity high
Study strength medium
not reported
0.24
Different measurement functions applied to the same underlying concept can preserve different features of the data and therefore support different empirical conclusions. Decision Quality mixed Empirical conclusions produced by alternative measurements
Reading fidelity high
Study strength medium
not reported
0.24
The abundance of AI-generated measures creates risks of noisy, poorly validated, or difficult-to-interpret variables, referred to by the authors as “AI slop.” Error Rate negative Quality and interpretability of AI-generated variables
Reading fidelity high
Study strength medium
not reported
0.24
Credible inference using AI-generated variables requires measurement validation anchored to explicit, observable criteria rather than informal claims that a proxy is reasonable. Decision Quality positive Credibility and validity of inference using AI-generated variables
Reading fidelity high
Study strength high
not reported
0.4
A validation sample can support valid inference even when AI predictions are arbitrarily biased, because prediction errors can be learned from the validation design rather than inferred from assumptions about how the model generates its outputs. Decision Quality positive Validity of statistical inference with biased AI predictions
Reading fidelity high
Study strength high
not reported
0.4
When a random validation sample is unavailable, the credibility of inference depends on assumptions about how AI models behave across settings rather than on the validation-sample design. Decision Quality negative Credibility of inference without random validation data
Reading fidelity high
Study strength medium
not reported
0.24
AI can improve the discovery stage by surfacing patterns that researchers did not specify in advance, but evidence for a generated hypothesis must come from held-out data. Innovation Output positive Discovery of predictive or economically meaningful patterns
Reading fidelity high
Study strength medium
not reported
0.24
AI lowers the cost of generating and comparing alternative operationalizations of a concept, but it cannot determine which operationalization answers the research question; that judgment requires theory and substantive expertise. Decision Quality mixed Quality and appropriateness of construct definitions
Reading fidelity high
Study strength medium
not reported
0.24
For the discovery stage, the quality of a generated hypothesis should be evaluated on held-out data by measuring the named construct and assessing how well it predicts the outcome of interest. Output Quality positive Out-of-sample predictive performance of generated hypotheses
Reading fidelity high
Study strength high
not reported
0.4

Notes