The Commonplace
Home Three-study pilot Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Graph rules that guarantee optimal covariate adjustment for the ATE do not generalise to the treated-population effect: there is no graph-only covariate set that is ATT-optimal under every compatible distribution, and adding outcome-side predictors can raise variance (especially when treatment is rare). The paper gives exact formulas, counterexamples, and practical guidance — including symmetric thresholds for overlap weights — plus simulations and a LaLonde illustration.

Optimal Covariate Adjustment beyond the Average Treatment Effect: Treated-Population and Overlap-Weighted Estimands
Shoki Okubo · September 10, 2026
arxiv theoretical n/a evidence 7/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

Arxiv

Latest observation:

  1. Shoki Okubo unresolved corpus identity
The paper shows that, unlike the ATE, no adjustment set computed from the causal graph is universally optimal for the ATT: adding covariates can sometimes increase asymptotic variance, and the paper gives exact variance-change identities, impossibility constructions, sufficiency conditions (no effect modification), and extensions to weighted estimands including overlap weights.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

Graphical causal inference supplies a complete theory of efficient covariate adjustment for the average treatment effect: one adjustment set, computable from the graph, is optimal under every compatible distribution. We show that this is a property of the average treatment effect's inverse-prevalence weights, not of causal estimands in general. For the average treatment effect on the treated we index the efficiency bound by the adjustment set and derive exact identities for its change under treatment-side and outcome-side extensions of a valid set. Covariates that predict only the treated-arm outcome are exactly efficiency-neutral, and covariates that predict the control-arm outcome can strictly increase the bound when the propensity is below one half -- a reversal of the supplementation lemma whose source is an arithmetic-geometric-mean inequality that holds for the average treatment effect and fails for the treated-population estimand. A construction with two faithful distributions on one graph proves that no graphical optimality criterion exists for the treated-population estimand; under no effect modification the ATE-optimal set is nonetheless optimal among the graphically valid sets, with an exact expression for its advantage. The results extend to weighted average treatment effects with propensity-dependent weights, yielding symmetric thresholds for overlap weights, an estimand-drift phenomenon under instrument adjustment, and a characterization of constant weights as the only smooth positive weights for which outcome-side supplementation never increases the bound. Simulations and the LaLonde data provide illustrations.

Summary

Main Finding

Graphical optimality of covariate adjustment (a single graph-computable adjustment set that is optimal for every compatible distribution) is a special property of the population average treatment effect (ATE) and its inverse-prevalence weights. It fails for the average treatment effect on the treated (ATT) and for general weighted average treatment effects (WATEs). For the ATT (and many WATEs) (i) whether supplementing a valid adjustment set with additional covariates raises or lowers the semiparametric variance depends on the data-generating law (not just the DAG), (ii) covariates that predict only the treated-arm outcome are neutral for ATT efficiency, while control-arm-only predictors can strictly harm ATT precision under low treatment prevalence, and (iii) there exist DAGs with two faithful laws for which the ATT-optimal valid set differs, so no graph-to-set rule can be ATT-optimal uniformly. The ATE-optimal set remains optimal for ATT among graphically valid sets under no effect modification. The paper extends these results to a parametric family of WATEs (including overlap weights) and gives thresholds and phenomena (e.g., estimand drift under instrument adjustment) that guide practice.

Key Points

  • Definitions and scope

    • Valid adjustment set S: suffices for unconfoundedness and positivity (standard definition). The paper contrasts sets valid at a law P versus graphically valid sets A(G).
    • Estimands considered: ATE (h ≡ 1), ATT (h(e)=e), ATC, ATO (overlap; h(e)=e(1−e)), and general WATEs with smooth h(e).
    • Efficiency comparisons are between adjustment estimators that do not impose additional restrictions (no propensity-exclusion restrictions).
  • Semiparametric bounds (closed forms)

    • ATT bound (for a valid set S): V_att(S) = (1/p^2) E[ e_S σ1^2(S) + e_S^2/(1−e_S) σ0^2(S) + e_S (τ_S − ψ)^2 ]. (Here e_S = P(A=1 | S), μ_a(S), σ_a^2(S), τ_S = μ1(S) − μ0(S), p = P(A=1), ψ = ATT.)
    • General WATE bound (normalizing D_S = E[h(e_S)]): V_h(S) = D_S^{−2} E[ h(e_S)^2/e_S σ1^2 + h(e_S)^2/(1−e_S) σ0^2 + (τ_S − ψ_h)^2 { h(e_S)^2 + h'(e_S)^2 e_S (1−e_S) } ].
    • The derivative term h'(e)^2 e(1−e) is the price of estimating propensity-dependent weights (vanishes for ATE).
  • Supplementation identities and consequences

    • Treatment-side supplementation (adding covariates unrelated to outcome given the current set) mirrors the ATE result: instruments are variance loads.
    • Outcome-side supplementation (adding covariates unrelated to treatment given the current set) behaves differently:
    • For ATE, adding a covariate that predicts only one arm cannot increase variance (supplementation lemma); arithmetic–geometric-mean inequality in ATE weights enforces this.
    • For ATT, adding a covariate that predicts only the treated-arm outcome is exactly efficiency-neutral.
    • For ATT, adding a covariate that predicts only the control-arm outcome can strictly increase the variance when treatment prevalence is low and effect modification is arm-asymmetric (reversal of the supplementation lemma).
    • These sign decisions are given in exact conditional-moment identities (the paper provides the closed-form difference of bounds when extending S by W subject to W ⫫ A | S or W ⫫ Y | (A,S)).
  • Graphical (im)possibility results

    • Theorem: There exists a DAG and two strictly positive distributions both Markov and faithful to the DAG such that their ATT-optimal graphically valid sets differ. Therefore, no rule mapping a DAG to a single adjustment set can be ATT-optimal across all compatible laws.
    • Corollary: The ATE-optimal set O(G) can be strictly suboptimal for ATT; the paper gives closed-form variance penalties.
    • Rescue: If there is no effect modification (so ATT = ATE numerically), the ATE-optimal set O(G) is ATT-optimal among graphically valid sets; an exact advantage expression is provided.
  • Extensions to WATE and overlap weights

    • The analysis extends to smooth h(e). For overlap weights h(e)=e(1−e) the paper derives symmetric pointwise thresholds:
    • A control-arm-only predictor harms ATO efficiency when e < 1/3.
    • A treated-arm-only predictor harms ATO efficiency when e > 2/3.
    • Constant weights (ATE) are characterized as the only smooth positive weights for which outcome-side supplementation never increases the bound.
    • Novel phenomenon: for overlap-type weights, adjusting for an instrument can change the estimand itself (estimand drift), so deleting instruments is not even a well-posed efficiency comparison for ATO without fixing a canonical weight.
  • Finite-sample and empirical illustrations

    • Simulations: in the designs studied, bound orderings predict finite-sample RMSE and influence-based SEs are well calibrated.
    • Random-structure scan: among 49 random four-covariate DAGs with arm-asymmetric modification, ATE- and ATT-optimal graphically valid sets differed in 10 cases; ATT penalties for using the ATE-optimal set ranged 2.9–25.9%.
    • LaLonde (Dehejia–Wahba subsample) example: rare-treatment regime (p ≈ 1.1%); leave-one-out deletions left estimates nearly unchanged; SE reductions up to 1.6% observed — the example illustrates the regime in which ATT-vs-ATE differences matter.
  • Practical takeaway emphasized in paper: substantive validity should dominate covariate choice for ATT; precision-based rules that rely only on the DAG can be misleading.

Data & Methods

  • Theoretical framework

    • Nonparametric structural equation model with a causal DAG; work distinguishes graphically valid sets A(G) and sets valid at a particular law A(P).
    • Semiparametric efficiency theory: pathwise differentiability, efficient influence functions, and nonparametric efficiency bounds for functionals χ_S (ATT) and ψ_h(S) (WATE).
    • Derived closed-form influence functions and bounds (see Lemmas 1–2 and Table 1 in paper).
    • Algebraic identities for how bounds change when adding covariates on treatment or outcome side; uses conditional variance decompositions and weight arithmetic (source of the ATE/ATT difference is an arithmetic–geometric-mean inequality that holds for ATE weights but not ATT weights).
  • Proof strategy and constructions

    • Constructive counterexample: a DAG and two faithful laws with differing ATT-optimal sets to prove nonexistence of universal graph-based optimality for ATT.
    • Rescue proofs use supplement-then-delete architecture (in parallel to Rotnitzky & Smucler 2020) but with ATT/WATE increments.
  • Estimators and attainment

    • Augmented inverse probability weighting (AIPW) with cross-fitted nuisance estimates and targeted maximum likelihood estimation (TMLE) are used to attain the derived bounds under standard regularity (overlap, bounded moments, L2 consistency of nuisances, product-rate conditions).
    • Estimators in simulations and application are cross-fitted AIPW/TMLE variants to align with the asymptotic theory.
  • Simulations and empirical data

    • Simulations calibrated to show finite-sample implications of bound orderings.
    • Random-structure scan across DAGs with four covariates.
    • Application to LaLonde training program (Dehejia–Wahba subsample): 185 treated vs 15,992 CPS controls; replication materials provided.
  • Reproducibility

    • Replication code and data: https://github.com/sokubo/paper-estimand-adjustment-replication
    • Proofs in Appendix A and numerical verification in Appendix B.

Implications for AI Economics

  • Estimand-aware covariate selection

    • When policy evaluation or program-effect targets focus on the treated (ATT) or other non-ATE WATEs, analysts cannot rely on a single DAG-derived adjustment set to be uniformly optimal. Variable selection and adjustment should be estimand-specific and may need distributional evidence (e.g., variance and arm-specific predictive power).
    • In practice, automated covariate-selection tools or ML pipelines that aim to minimize variance should be configured with the target estimand in mind (ATE vs ATT vs ATO) — the same covariate can help ATE but harm ATT.
  • Rare-treatment regimes

    • Many policy and pharmacoepidemiology applications have rare treatments. The paper shows that when p is small, control-arm-only predictors are particularly likely to harm ATT efficiency; therefore, in AI-economics workflows that use ML models for propensity or outcome modeling, careful assessment of control-arm predictors’ role is important.
  • Use of overlap weights and targeted weighting

    • Overlap-type weights (ATO) are attractive under limited overlap, but the paper warns that for these weights:
    • There are sharp symmetric thresholds (e < 1/3, e > 2/3) where single-arm predictors can harm efficiency.
    • Instruments (variables unrelated to the outcome but predictive of treatment) can change the estimand under overlap weights (estimand drift). Analysts should fix a canonical weighting scheme before treating instrument deletion as a variance-only decision.
  • Graphical analyses and ML pipelines

    • DAG-based adjustment rules remain valuable as conservative validity checks (they guarantee unconfoundedness under the graph), but practitioners should not assume they also deliver uniform efficiency for estimands beyond ATE.
    • For ATT-targeted analyses, complement graphical selection with distributional diagnostics (e.g., conditional variances, arm-specific predictive power), sensitivity analyses, and simulation-based checks.
  • Recommendations for applied AI-economics practice

    • Explicitly state the estimand (ATE vs ATT vs ATO) before variable-selection and modeling.
    • Use semiparametric estimators known to attain the bounds (AIPW/TMLE with cross-fitting) to avoid estimator suboptimality masking the bound-level effects.
    • Run sensitivity checks or small-scale simulations that vary the covariate set to see how estimated SEs and RMSE behave in finite samples; do not treat DAG-derived optimality as sufficient for ATT.
    • When using overlap weights or other propensity-dependent weights, be aware that including instruments can change what you’re estimating; fix the weight function before variable deletion.
    • When substantive arguments for inclusion/exclusion are weak, prefer reporting sensitivity to covariate choices rather than selecting a single “optimal” set.

In short: for ATE the graphical theory gives a distribution-free optimal adjustment set; for ATT and many WATEs it does not. Applied AI-economics workflows that use ML for causal estimation must be estimand-aware and combine DAG-based validity with distributional diagnostics and semiparametric estimators that attain the theoretical bounds.

Assessment

Paper Typetheoretical Evidence Strengthn/a — Paper is primarily a mathematical and semiparametric-theory contribution with proofs and formal derivations; empirical material is limited to simulations and an illustrative LaLonde reanalysis, so there is no substantive empirical claim about real-world AI effects to evaluate. Methods Rigorhigh — The paper derives exact semiparametric influence functions and efficiency bounds, proves impossibility and sufficiency theorems on DAGs with constructive counterexamples, extends results to a smooth class of propensity-dependent weights, and verifies findings with simulations and a canonical data illustration; proofs and connections to the literature are explicitly cited and sketched. SampleMain results are theoretical (nonparametric model for (S,A,Y) under overlap). Empirical checks: simulation designs (random-structure scan across four-covariate graphs and other simulated data; details in appendices) and an illustrative reanalysis of the Dehejia–Wahba subsample of the LaLonde data (185 treated trainees vs 15,992 CPS controls, mixing proportion ~1.1%). Replication code and data are provided in a public GitHub repository. Themeslabor_markets productivity IdentificationIdentification is via standard conditional ignorability / adjustment: (Y(0),Y(1)) ⟂ A | S together with positivity; valid adjustment sets are defined either directly at the law or via DAG-based graphical adjustment criteria (Shpitser et al., Perković et al.). The paper frames estimands (ATE, ATT, WATEs) as functionals of the observed law conditional on a valid S and derives pathwise derivatives / semiparametric influence functions and efficiency bounds in the nonparametric model for (S,A,Y). GeneralizabilityResults depend on unconfoundedness/valid-adjustment assumptions and positivity; if these fail (e.g., unobserved confounding) conclusions do not hold., Graphical impossibility results are about existence of a universally optimal set across all distributions compatible with a DAG; individual laws may admit different optimal sets not readable from the graph., Asymptotic semiparametric bounds may not fully predict finite-sample performance in all practical designs, though simulations suggest close alignment in studied designs., Extensions assume smoothness of weight functions (for WATE results); discrete or non-smooth weight schemes may require separate treatment., Theoretical results are for cross-sectional single-treatment settings; time-varying treatments or complex longitudinal data require additional work (though literature extensions exist).

Claims (11)

ClaimDirectionOutcomeConfidence & EvidenceDetails
For the ATT, adding a covariate that predicts only the treated-arm outcome is exactly efficiency-neutral. Organizational Efficiency null_result Semiparametric efficiency bound for the ATT
Reading fidelity high
Study strength high
not reported
0.2
Adding a covariate that predicts the control-arm outcome can strictly increase the ATT efficiency bound when the treatment propensity is below one half and treatment effects are arm-asymmetric. Organizational Efficiency negative ATT semiparametric efficiency bound
Reading fidelity high
Study strength high
not reported
0.2
The ATT-optimal valid adjustment set cannot generally be determined from the causal graph alone. Organizational Efficiency negative Optimal adjustment-set choice for ATT efficiency
Reading fidelity high
Study strength high
not reported
0.2
The ATE-optimal adjustment set can be strictly suboptimal for the ATT. Organizational Efficiency negative ATT asymptotic variance or efficiency bound
Reading fidelity high
Study strength high
not reported
0.2
Under no effect modification, the ATE-optimal adjustment set is also ATT-optimal among graphically valid adjustment sets. Organizational Efficiency positive ATT efficiency relative to alternative graphically valid adjustment sets
Reading fidelity high
Study strength high
not reported
0.2
For overlap weights, adding a control-arm-only outcome predictor harms efficiency when the propensity score is below 1/3, while adding a treated-arm-only predictor harms efficiency when the propensity score is above 2/3. Organizational Efficiency negative Overlap-weighted average treatment effect efficiency bound
Reading fidelity high
Study strength high
thresholds e < 1/3 and e > 2/3
0.2
Constant weights are the only smooth positive propensity-score weights for which outcome-side supplementation never increases the efficiency bound, regardless of the joint predictive structure of the two treatment arms. Organizational Efficiency null_result Weighted average treatment effect efficiency bound under outcome-side covariate supplementation
Reading fidelity high
Study strength high
not reported
0.2
For overlap-type weights, adjusting for an instrument can change the estimand itself, not merely its variance. Decision Quality mixed Definition of the overlap-weighted treatment effect estimand
Reading fidelity high
Study strength high
not reported
0.2
In a scan of random four-covariate structures, the ATE- and ATT-optimal graphically valid adjustment sets differed in 10 of 49 structures with arm-asymmetric effect modification. Organizational Efficiency mixed Agreement between ATE- and ATT-optimal adjustment-set choices
Reading fidelity high
Study strength medium
n=49
10 of 49 structures
0.12
In the random-structure scan, using the ATE-optimal set incurred exact ATT penalties ranging from 2.9% to 25.9%. Organizational Efficiency negative ATT efficiency penalty from using the ATE-optimal adjustment set
Reading fidelity high
Study strength medium
n=49
2.9% to 25.9% penalty
0.12
In the Dehejia–Wahba LaLonde subsample, the seven leave-one-out specifications that left the estimate essentially unchanged had standard errors at most 1.6% below those of the full model. Organizational Efficiency negative ATT estimator standard error under leave-one-out covariate specifications
Reading fidelity high
Study strength medium
n=16177
standard errors at most 1.6% below the full model's
0.12

Notes