0 cumulative citations
View corpus contextAI is better at doing than judging: occupations concentrated in execution tasks have seen slower employment growth since 2012, a pattern that predates the recent AI surge; AI-capability scores rise after 2022, but the paper documents timing and measurement rather than proving an AI-driven causal effect.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
0 cumulative citations
View corpus contextArtificial intelligence automates execution more readily than evaluation: producing output is cheap, judging whether it is correct is not. Exposure measures rank tasks by whether AI can perform them, not by which function the human supplies. I score all $19{,}265$ O*NET task statements under fixed rubrics to build occupation-level execution and AI-capability shares. The execution share is reproducible across model coders and O*NET vintages and distinct from AI capability and routine-task intensity; it is a model-based measure, not human-validated ground truth, and adds only modest power beyond O*NET's evaluation activities. In a harmonized panel, employment growth is lower in execution-heavy white-collar occupations in every window since 2012, and equality of slopes cannot be rejected: the gradient is a secular trend rather than an AI-era event, largely between occupational families. The vintage-valid capability gradient steepens after 2022, a change that is dated but not causally attributable. The evidence establishes a measure and a chronology, not an AI-caused effect.
Summary
Main Finding
The paper builds a reproducible, task-level execution share (and a matched AI-capability share) by scoring all 19,265 O*NET task statements, and uses them to show (1) that occupations whose human work is concentrated in execution tasks exhibit lower employment growth in long windows (2012–2025), but (2) that this execution-gradient is a long-run, between-occupational-family association predating generative-AI diffusion rather than clear evidence of an AI-caused displacement. A vintage-valid capability measure steepens after 2022, but the paper emphasizes dating and descriptive chronology rather than causal attribution.
Key Points
- New measures
- Scored 19,265 O*NET task statements under two rubrics: execution (producing/performing vs evaluating/verifying) and AI capability (can a frontier LLM with tool access perform the task to professional standard).
- Aggregated to occupation-level shares using O*NET task weights.
- Reproducibility and robustness
- Execution share reproducible across coders (r = 0.924), across O*NET vintages (occupation-level corr. = 0.996 between 2021 and 2026 task universes), and across alternative task-weightings (corr. 0.98–0.99).
- Capability scores: 2026-vintage capability correlates 0.80 with early-2023 GPT-4 exposure (Eloundou et al., 2024); observed-usage index (Anthropic Economic Index) correlates ~0.60 (all occupations) / 0.48 (cognitive).
- Relation to O*NET evaluation activities
- Execution share is strongly negatively related to O*NET’s human-rated evaluation activity composite (construct validity), but the composite predicts 2019–2025 employment growth more strongly.
- Incremental predictive power: on the cognitive sample, adding execution share to the O*NET evaluation composite raises R2 by only 0.006; in the full controlled specification the execution share adds R2 = 0.011 vs 0.058 for the composite.
- Employment facts and chronology
- In harmonized OEWS panel analyses, execution-intensive white-collar occupations show lower employment growth in each window examined (2012–2016, 2016–2019, 2019–2025). Equality of execution slopes through 2025 cannot be rejected → a secular trend.
- Much of the execution-employment association is between occupational families: adding two-digit SOC major-group fixed effects reduces the execution coefficient from ≈ −0.70 to ≈ −0.14 and eliminates significance.
- Capability exposure (vintage-valid early-2023 measure) is already negatively associated with employment growth pre-2022 and steepens after 2022; the timing is dated but not claimed causal.
- Observed AI usage: employment contraction is concentrated in occupations that are both high in observed AI use and high in execution share; the high-high cell is negative even with family fixed effects, but continuous interaction and family-adjusted super-additivity tests are insignificant (descriptive concentration, not evidence of multiplicative treatment).
- Other outcomes
- Execution-intensive occupations have a positive wage coefficient (cautioning against reading employment declines purely as inward demand shifts).
- A CPS test for entry-age employment is null (but CPS misses hiring/entry margins).
- Measurement caveats
- The execution and capability shares are model-based/generated regressors (not a human-validated ground truth). Main regressions do not propagate scoring uncertainty in standard errors (Appendix B quantifies scoring uncertainty).
Data & Methods
- Task scoring
- Universe: 19,265 O*NET task statements (version 26.1; baseline), re-scored on 30.3 (2026) for robustness.
- Execution rubric: rates production/doing vs evaluation/verifying; explicitly separated from routineness.
- Capability rubric: whether a 2026-frontier language model with typical tool access could perform the task to professional standard.
- Occupation aggregation: weighted averages of task scores using O*NET task importance weights (core tasks double-weighted); alternative weighting schemes tested.
- Validation & comparisons
- Independent model-coder audit (r = 0.924), correlations across vintages and weightings reported.
- External anchors: Eloundou et al. (2024) early-2023 GPT-4 exposure and Anthropic Economic Index usage (post-2024 conversations).
- Outcomes and samples
- Employment and wages: OEWS (BLS) national cross-industry estimates. Six-year baseline outcome: Dj = ln(empj,2025) − ln(empj,2019). Longer windows include 2012–2016 and 2016–2019 harmonized panels.
- Principal sample: ten white-collar SOC major groups (400 occupations in task data; 337 cognitive occupations in 2019–2025 regressions, 284 in tighter samples).
- Regression specification: cross-occupation OLS with controls (routine intensity, interest-rate exposure, other covariates), interactions (execution × capability), occupational-family fixed effects, and robustness checks for generated-regressor issues in Appendix B.
- Theory linking measures to empirical tests
- CES-style execution production function blending human execution and AI input (equation (1)); human execution demand falls with capability frontier At (equation (2)).
- Oversight residual r: fraction of evaluation labor that remains human.
- Proposition 1: occupations with higher execution share ej experience larger conditional declines in human labor as capability rises.
- Proposition 2: if entry-level execution falls, a lagged decline in senior evaluators follows once affected cohorts reach seniority (cohort-lag channel); testable only with appropriate longitudinal hiring-to-seniority data.
Implications for AI Economics
- Measurement contribution
- Provides a portable, reproducible task-level axis (execution vs evaluation) that complements capability and usage measures. This permits separating "can AI do the task?" (capability) from "what human function remains?" (execution vs evaluation).
- Interpreting employment patterns
- The execution share helps interpret why some occupations are more sensitive to automation: when verification/oversight is the binding human input, execution automation need not translate into immediate occupational-level displacement.
- The observed long-run execution gradient (present before generative-AI diffusion) implies a secular reallocation across occupational families; attributing recent employment changes to generative AI requires care and causal identification.
- Dynamics and policy focus
- If evaluation skill is produced by execution practice (learning-by-doing), automation that reduces entry-level execution could produce delayed shortages of senior evaluators or reductions in evaluative quality — a policy-relevant cohort-lag risk that ordinary cross-sectional exposure measures miss.
- Policy levers should therefore attend to: preserving on-the-job execution practice for skill formation; certification/licensing and liability regimes that determine the oversight residual; and targeted training to maintain evaluator capacity.
- Research agenda
- For causal inference: combine the execution measure with firm-level hiring and vacancy flows, staggered adoption, or instrumented adoption to test multiplicative predictions (execution × capability × usage) and to separate contemporaneous usage from vintage-valid capability.
- For dynamics: link entry hiring, actual execution practice, and later senior stocks/assessment quality using administrative or panel data spanning the cohort lag L in the model.
- For measurement: propagate scoring uncertainty into inference and extend human validation exercises; apply the rubric to other task corpora beyond O*NET.
- Theoretical implications
- Supports models where verification/oversight is a bottleneck (verification costs decline slower than generation costs) and where automation effects depend on the human function (execution vs evaluation). Empirical patterns here suggest that simple capability-only exposure rankings can miss important heterogeneity in labour-market responses.
Summary takeaway: the paper supplies a carefully validated task-based execution measure and shows that execution intensity correlates with lower long-run employment growth across white-collar occupations — but that this gradient is largely a secular, between-family pattern predating generative-AI diffusion. The capability measure steepens post-2022, and observed usage concentrates contractions where execution and usage are both high, but causal claims about AI-driven displacement require more targeted, dynamic identification.
Assessment
Claims (11)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| Artificial intelligence automates execution more readily than evaluation: producing output is cheap, judging whether it is correct is not. Automation Exposure | positive | automation of execution vs. evaluation |
Reading fidelity
high
Study strength
speculative
|
not reported
|
| Exposure measures rank tasks by whether AI can perform them, not by which function the human supplies. Automation Exposure | negative | task ranking by AI performance vs. human function |
Reading fidelity
high
Study strength
medium
|
not reported
|
| I score all 19,265 O*NET task statements under fixed rubrics to build occupation-level execution and AI-capability shares. Automation Exposure | null_result | occupation-level execution share and AI-capability share (measurement construction) |
Reading fidelity
high
Study strength
medium
|
n=19265
|
| The execution share is reproducible across model coders and O*NET vintages. Automation Exposure | positive | reproducibility of execution share measure |
Reading fidelity
high
Study strength
medium
|
not reported
|
| The execution share is distinct from AI capability and routine-task intensity. Automation Exposure | null_result | distinctness/orthogonality of measures |
Reading fidelity
high
Study strength
medium
|
not reported
|
| The execution share is a model-based measure, not human-validated ground truth. Other | null_result | validity status of execution share |
Reading fidelity
high
Study strength
speculative
|
not reported
|
| The execution share adds only modest power beyond O*NET's evaluation activities. Other | null_result | incremental predictive power (model fit) for outcomes the paper examines |
Reading fidelity
medium
Study strength
medium
|
not reported
|
| In a harmonized panel, employment growth is lower in execution-heavy white-collar occupations in every window since 2012. Employment | negative | employment growth |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Equality of slopes cannot be rejected: the gradient is a secular trend rather than an AI-era event, largely between occupational families. Employment | null_result | time-invariance of execution-employment-growth gradient (slope equality) |
Reading fidelity
medium
Study strength
medium
|
not reported
|
| The vintage-valid capability gradient steepens after 2022, a change that is dated but not causally attributable. Employment | negative | employment-growth gradient with respect to AI-capability share (change after 2022) |
Reading fidelity
high
Study strength
medium
|
not reported
|
| The evidence establishes a measure and a chronology, not an AI-caused effect. Other | null_result | causal attribution of AI on employment/productivity |
Reading fidelity
high
Study strength
speculative
|
not reported
|