0 cumulative citations
View corpus contextWhen excluded instruments are unavailable, small local differences in the distribution of first-stage shocks can identify endogenous coefficients: by approximating the control function with a sieve and bounding approximation/misspecification error, the paper derives moment inequalities that pin down structural parameters and provides a bootstrap-based inference routine.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
This paper studies identification and inference in a triangular system with an endogenous regressor when exclusion restrictions are unavailable and the dependence between structural disturbances is modeled through an unrestricted control function. In this setting, standard orthogonality conditions do not deliver point identification, as the unknown control function can rationalize a wide range of structural coefficients. We show that identifying information can be extracted from distributional shifts in the first-stage disturbance induced by an auxiliary variable that may directly affect the outcome. Imposing a local restriction on the log density ratio, together with an explicit bound on the sieve approximation error of the control function, we derive moment inequalities that restrict the structural parameter. We develop a practical inference procedure based on test inversion and multiplier bootstrap that accommodates generated regressors, cross-fitted sieve estimation, and locally estimated density-ratio nuisances. The results clarify how identification can be recovered from weak local distributional structure in the absence of classical instruments.
Summary
Main Finding
The paper shows that in a triangular system with an endogenous regressor and an unrestricted nonparametric control function, researchers can recover identifying information about the structural coefficient on the endogenous regressor even when no classical exclusion instrument exists. The identifying content comes from small, local distributional shifts in the first-stage disturbance induced by an auxiliary variable that may directly affect the outcome. By (i) imposing a local approximation of the log-density ratio (a quadratic tilt plus a bounded remainder), (ii) bounding sieve approximation error for the unknown control function, and (iii) applying Cauchy–Schwarz to residualized, locally weighted cross-group moment differences, the authors derive moment inequalities that restrict the structural parameter. As the sieve bias vanishes and under a local relevance condition, these inequalities deliver point identification in the limit. They also provide a practical inference procedure (test inversion with multiplier bootstrap) that accommodates generated regressors, cross-fitted sieve estimation, and locally estimated density-ratio nuisances.
Key Points
- Model and problem
- Triangular system: Y1 = X′β1 + γ1 Y2 + ε1 and Y2 = X′β2 + ε2, with ε1 = h(ε2) + η. The control-function h(·) is unknown and unrestricted; η satisfies E[η|X,ε2]=0.
- With no valid excluded instrument, baseline orthogonality moments E[f(ε2)(Y1 − X′β1 − γ1 Y2 − h(ε2))] = 0 do not point-identify γ1 because h can absorb endogeneity.
- New identifying idea
- Introduce a binary auxiliary variable Z0 that may change the distribution of the first-stage shock ε2 but is not required to be excluded from the outcome equation.
- Work locally around an expansion point e0 and set u = ε2 − e0. Model the log density ratio ℓ(u) = log(f1(u)/f0(u)) on a local window |u| ≤ b as ℓ(u) = a + λ u^2 + r(u) with sup|r(u)| ≤ δb (local quadratic tilting with bounded remainder).
- Intuition: the λu^2 term captures local variance/dispersion shifts across Z0; δb is a sensitivity parameter for local misspecification.
- Sieve approximation and residualization
- Use Frisch–Waugh–Lovell residualization to remove X (so ε2 is observable as the residual).
- Locally approximate h(·) by a sieve projection hj over the same local window and let rj = h − hj be the sieve error; assume E0[ψb(u) rj(u)^2] ≤ δh(j)^2 with δh(j) → 0 as sieve dimension j → ∞.
- Moment inequalities
- Define a locally weighted cross-group difference ∆E[ψb(u) Wj(γ)], where Wj(γ) is the residualized outcome after subtracting γ · ε2 and hj.
- Using the density-ratio identity and Cauchy–Schwarz, obtain the inequality |∆E[ψb(u) Wj(γ)]| ≤ δh(j) √Vb, where Vb = E0[ψb(u) (e^{ℓ(u)} − 1)^2].
- Rearranged, these yield two one-sided moment inequalities that a candidate γ must satisfy. Under a local relevance condition ∆E[ψb(u) ε2] ≠ 0, the identified set for γ has diameter bounded by (2 δh(j) √Vb) / |∆E[ψb(u) ε2]|; it shrinks to {γ1} as δh(j) → 0.
- Inference approach
- Construct sample analogues using generated first-stage residuals (from estimated β2), cross-fitted local sieve regression for hj, and locally estimated density-ratio nuisances.
- Obtain confidence sets by inverting a max-type test statistic with multiplier bootstrap critical values; intersect results across grids of tuning parameters (e0, b, sieve dimension, sensitivity parameters).
- Procedure accounts for generated regressors and nuisance estimation while preserving asymptotic coverage.
Data & Methods
- Econometric setup
- Triangular system with endogenous regressor Y2 and unrestricted control function h linking ε1 and ε2.
- Key assumptions: mean independence E[ε1|X]=E[ε2|X]=0, E[η|X,ε2]=0, finite moments, ΣXX positive definite.
- Distributional-shift modeling
- Binary auxiliary Z0 induces conditional densities f0, f1 for u = ε2 − e0.
- Log-density ratio ℓ(u) approximated locally as ℓ(u) = a + λ u^2 + r(u), with sup_{|u|≤b} |r(u)| ≤ δb.
- Normalization constraint E0[e^{ℓ(u)}] = 1 pins down a up to the δb margin.
- Residualization and local weighting
- Use population/empirical FWL residuals to remove X: ˜Y2 = ε2, ˜Y1 = γ1 ε2 + h(ε2) + η.
- Apply a local bounded weight ψb(u) = K(u/b) 1{|u|≤b} to focus moments on the local window.
- Sieve approximation
- Choose basis {pk} and project h onto Hj (dimension j) using weighted L2(µ0,b) projection; bound the L2 error by δh(j).
- Derivation of moment inequalities
- Combine density-ratio identity ∆E[g(u)] = E0[g(u) (e^{ℓ(u)} − 1)] with the residualized decomposition Wj to obtain |∆E[ψb(u) Wj]| ≤ δh(j) √Vb.
- Estimation & inference
- Replace latent ε2 by generated residuals from first-stage regression.
- Estimate hj using cross-fitted sieve regression; estimate local density ratios (and hence Vb) locally.
- Compute sample moment inequalities, use a max-statistic and multiplier bootstrap to obtain critical values, invert tests to get confidence sets; intersect across tuning-grid choices.
- Key tuning and sensitivity elements
- e0 (centering), b (bandwidth/window), sieve dimension j, and sensitivity parameter δb control the strength and interpretation of the identifying assumptions. δh(j) captures approximation error and is central to the identified-set size.
Implications for AI Economics
- Identification without classical instruments
- Many AI-economics settings (algorithmic personalization, platform interventions, A/B tests with imperfect exclusion, market design changes) lack credible exclusion restrictions. This paper provides a viable alternative: exploit auxiliary variables or context changes that induce local distributional shifts in first-stage residuals—even if they also affect outcomes directly.
- Practical strategy for empirical work on algorithmic/firm behavior
- Researchers can look for sources of local dispersion/variance shifts (for example, policy changes, time-of-day effects, or operational regime switches) that meaningfully change the distribution of unobserved inputs into an endogenous decision (ε2). Even if those sources are not excluded from outcomes, they can deliver identifying content under the paper’s local tilting + sieve-bias control framework.
- Transparent sensitivity and trade-offs
- The framework makes explicit two sources of fragility: (i) local misspecification of the log-density ratio (δb) and (ii) sieve approximation error for the control function (δh(j)). Those are interpretable tuning knobs that researchers can vary and report, enabling structured sensitivity analysis in empirical AI-economics studies.
- Methodological recommendations for applied researchers
- Use local weighting (focus on windows where the density-ratio is well-approximated), cross-fitted sieve methods, and account for generated regressors when estimating residuals—these steps are necessary to obtain valid inference as per the proposed bootstrap-based test inversion.
- Collect or exploit auxiliary variation that plausibly affects dispersion (heteroskedasticity) of first-stage shocks; even if that variation affects outcomes directly, it may still be useful for identification under the paper’s assumptions.
- Limitations and cautions
- Identification relies on a local relevance condition (∆E[ψb(u) ε2] ≠ 0) and on the density-ratio being well-approximated by the quadratic tilt within the chosen window. In practice, insufficient support in the local window, large δb, or slow decay of δh(j) can leave the identified set wide.
- The approach yields point identification only asymptotically as sieve error vanishes; finite-sample performance depends on tuning (b, e0, sieve size) and the strength of the local shift.
- Empirical application requires careful sensitivity reporting (vary δb, b, e0, and sieve dimension) and possibly external validation of the local-tilt plausibility.
Overall, the paper offers a principled, implementable route for recovering causal parameters in settings common in AI economics where exclusion restrictions are doubtful but auxiliary distributional shifts (e.g., changes in volatility or dispersion) can be observed and credibly modeled locally.
Assessment
Claims (7)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| With an unrestricted, potentially nonlinear control function h(ε2) and no valid excluded instrument, the structural coefficient γ1 is generally not point identified because the control function can absorb a wide range of structural parameter values. Other | null_result | Point identification of the structural coefficient on the endogenous regressor |
Reading fidelity
high
Study strength
high
|
not reported
|
| The baseline orthogonality conditions do not provide identifying power by themselves, even when multiplied by a large dictionary of functions of ε2, because the unrestricted control function remains inside the moments. Other | null_result | Identifying power of baseline orthogonality moments |
Reading fidelity
high
Study strength
high
|
not reported
|
| The first-stage coefficient β2 is point identified under the exogeneity and full-rank assumptions, and the OLS estimator of β2 is consistent; consequently, ε2 can be asymptotically proxied by the generated residual Y2 − X′β̂2. Other | positive | Identification and consistency of the first-stage coefficient and residual proxy |
Reading fidelity
high
Study strength
high
|
not reported
|
| A cross-group distributional shift in the first-stage disturbance, together with a bounded local approximation error for the control function, yields moment inequalities that restrict the structural parameters. Other | positive | Restriction of the identified set for the structural parameter |
Reading fidelity
high
Study strength
high
|
|∆E[ψb(u) W̃j]| ≤ δh(j)√Vb
|
| Under the stated local relevance condition, the identified set for γ1 has diameter bounded by 2δh(j)√Vb / |∆E[ψb(u)ε2]|. Other | positive | Diameter of the identified set for γ1 |
Reading fidelity
high
Study strength
high
|
diam(Γj) ≤ 2δh(j)√Vb / |∆E[ψb(u)ε2]|
|
| If the sieve approximation error δh(j) converges to zero as the sieve dimension increases, the identified set for γ1 shrinks to the singleton containing the true coefficient γ1, provided local relevance holds. Other | positive | Point identification of γ1 in the limit |
Reading fidelity
high
Study strength
medium
|
diam(Γj) → 0 as δh(j) → 0
|
| The proposed inference procedure constructs confidence sets by inverting a max-type test with multiplier-bootstrap critical values and intersecting across a finite grid of tuning parameters, while accommodating generated regressors, cross-fitted sieve estimation, and locally estimated density-ratio nuisances. Other | positive | Inference procedure's accommodation of nuisance estimation and generated regressors |
Reading fidelity
high
Study strength
medium
|
not reported
|