0 cumulative citations
View corpus contextConcentrated within-variation can break t-tests in saturated fixed-effect regressions, and no single critical value cures the boundary failure; constructing design-only contrasts (cycle-space sign-flips) gives finite-sample exact p-values under symmetric errors and delivers practical power via a cycle-packing algorithm.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
In saturated fixed-effects regressions, Gaussian inference depends not on total identifying variation but on its concentration, measured by the self-normalized leverage $λ_n$ of the residualized treatment. When finitely many score weights remain persistent, the $t$-statistic converges to a convolution of raw errors and a Gaussian component. At full concentration, its null distribution varies across symmetric error laws with equal variance, so no fixed critical value is uniformly valid. We instead construct nuisance-annihilating contrasts from the design alone. These eliminate the fixed effects identically and yield finite-sample exact sign-flip inference under symmetric, arbitrarily heteroskedastic errors, with no homogeneity assumptions or restrictions on the fixed-effect dimension. In two-way designs, admissible contrasts form the cycle space of the observation multigraph. Their efficiency is summarized by an observable capture ratio $κ$, which equals Pitman efficiency. The resulting design problem involves a capture--granularity trade-off: coarse supports maximize capture but reduce the number of randomization signs. Cycle packing provides sufficiently granular supports. On matched employer--employee data, a structure-exploiting algorithm achieves $κ\approx 0.51$, compared with $0.26$ for naive packing. In the Grunfeld investment regression, realized score concentration is $0.739$, corresponding to $N_{\mathrm{eff}}^{\mathrm{score}}=1.80$, while $32$ valid supports attain $κ=0.627$. The resulting exact $95%$ confidence interval is $[0.150,0.450]$. A worker--firm application demonstrates scalability to large networks.
Summary
Main Finding
When most treatment variation is absorbed by many fixed effects, conventional Gaussian (t-test) inference can fail because identification depends on how residualized treatment variation is concentrated across observations, not its aggregate size. Halkiewicz develops a finite-sample, design-based alternative: construct nuisance-annihilating contrasts (q with q′D = 0) from the design, use sign-flip randomization on disjoint supports to obtain exact inference under symmetric (and blockwise symmetric) errors, and—for two-way (worker–firm, firm–time) designs—shows these contrasts correspond exactly to the cycle space of the observation multigraph. He introduces an observable capture ratio κC (a Pitman-efficiency analogue) that governs power, provides cycle-packing algorithms to produce high-capture granular supports, and supplies diagnostics (λn, Hn, Neff, score-based analogues) and software. Conventional leverage-corrected tests can severely over-reject under concentrated identifying variation; the cycle/sign-flip tests are exact and competitive in power when appropriately packed.
Key Points
- Concentration parameter (design-only): λn = maxi exi^2 / Vn (where ex = MDx, Vn = ex′ex). λn → 0 is the usual negligibility condition for asymptotic normality; when λn is not small, the asymptotic null may be a convolution mixing few large-weighted errors and a Gaussian remainder (Proposition 2.1).
- Impossibility at the boundary (Proposition 2.3): At the fully concentrated boundary (one persistent weight), the limiting null law depends on the unknown symmetric error distribution, so no fixed critical value (tabulated test) is uniformly valid across symmetric equal-variance laws. Valid inference must adapt to the realized error law.
- Nuisance-annihilating contrasts: any q measurable from (x, D) with q′D = 0 extracts a linear combination free of fixed effects: q′(Y − xβ0) = q′ε under H0. If contrasts have disjoint supports and errors are blockwise independent and symmetric across those blocks (Assumption 2), flipping signs on contrasts yields exact finite-sample randomization tests (Theorem 3.3).
- Two-way designs ↔ cycle space: In bipartite observation graphs (edges = observations, vertices = firms/workers or firm/time), the space of annihilating contrasts equals the graph’s cycle space. The within-projection norm Vn = ||ΠZ x||^2 is the squared projection of treatment onto the cycle space (Proposition 4.1).
- Capture ratio κC: For a system of C disjoint supports with local normalized projections vc, capture is κC = (1/Vn) ∑c ||ΠZ_Ac x||^2 ∈ [0,1]. In diffuse regimes κC is the Pitman relative efficiency; κC^−1/2 is the multiplicative standard-error price (Theorem 3.9).
- Capture–granularity trade-off: Merging supports increases capture (can reach κC = 1 by taking one projection per biconnected block) but reduces the randomization group size (one sign per support), possibly wrecking power because the attainable minimal p-value grows (orbit-size floor). Fine, granular supports preserve randomization granularity but may lower capture. The paper formalizes and prices this trade-off.
- Cycle packing algorithms: constructs (digons, edge-disjoint firm-pair four-cycles, recursive contraction) that produce granular, high-capture supports. On public matched employer–employee extract, cycle packing achieves κC ≈ 0.51 (match-level treatment) vs 0.26 for naive greedy packing; for a time-varying covariate it achieves 0.91 vs 0.88.
- Empirical diagnostics & examples:
- Grunfeld investment regression: realized score concentration λn = 0.739 (N_score_eff = 1.80); 32 valid supports capture κC = 0.627; exact 95% CI is [0.150, 0.450].
- Monte Carlo: with concentrated design (λn = 0.29, effective sample 6.0), df-corrected t rejects up to 58.5% (size badly inflated), HC2 up to 33.4%; cycle test empirical size ~5% everywhere; cycle-test power often higher once oracle size correction applied to conventional tests.
- Software and reproducibility: implementations in PanelAdequacy.jl and an R companion (panelcert) are provided.
Data & Methods
- Model and conditioning: Linear model Yi = xiβ + di′γ + εi, conditional on the design (x, D). Focus is on inference for β with γ unrestricted nuisance.
- Key design objects:
- ex = MDx (treatment residualized on fixed effects)
- Vn = ex′ex
- λn = maxi exi^2 / Vn (self-normalized leverage)
- Herfindahl Hn = ∑(exi^2 / Vn)^2, Neff = 1/Hn
- Score-based realized analogues using fitted residuals (bλscore, bNscore_eff) as warnings.
- Limit theory: Proposition 2.1 gives convolutional limits when a finite number of weights persist; Corollary 2.2 shows studentization cannot generally restore normality.
- Exact-test construction:
- Build annihilating contrast system {q1,...,qC} with q′c D = 0 and pairwise disjoint supports Sc (measurable from design).
- Under Assumption 2 (blocks of observations that are independent and centrally symmetric; supports unions of blocks), contrast scores Uc = vc′(Y − xβ0) are invariant to flipping signs on each support.
- Define test-statistic T(U1,...,UC) (default T = |∑c bc Uc| with bc = vc′x) and compute randomization p-value by enumerating or Monte Carlo sampling sign patterns s ∈ {±1}^C (Theorem 3.3). Studentized variants are also exact.
- Two-way graph methods:
- Observations are edges of bipartite multigraph. Annihilating contrasts correspond to ±1 alternations around cycles (including digons for repeated matches).
- Projection of x onto cycle space gives Vn; local projections on supports give the captured portion.
- Cycle packing: algorithmic constructions (digons, firm-pair 4-cycles, recursive contraction) to get many edge-disjoint or nearly edge-disjoint cycles to maximize κC with maintained granularity.
- Monte Carlo setups:
- Simulation designs include concentrated and diffuse regimes, heteroskedastic symmetric errors, and dependence structures consistent with Assumption 2 at various granularities.
- Benchmarks included df-corrected t, HC2, and oracle-adjusted versions.
Implications for AI Economics
- Many AI-economics settings have high-dimensional fixed effects and networked observations (platform-user, user-item, firm-worker, time-varying policy adoption). When controlling for those effects, remaining identifying variation for treatment or instrument often concentrates on few observations; this paper shows standard asymptotic t- or cluster-robust inference can be invalid in that regime.
- Practical diagnostics: report λn, Hn, Neff and the score-based analogues (bλscore, bNscore_eff). If effective sample size is small or λn large, conventional inference is suspect and the exact randomization methods are a principled alternative.
- Exact randomization tests adapt to the realized (unknown) error law and remain valid under arbitrary heteroskedasticity so long as blockwise symmetry/exchangeability assumptions are plausible at the support granularity chosen. This is useful for AI policy evaluation (e.g., A/B tests with fixed effects, algorithmic interventions with network spillovers) where dependence and heteroskedasticity are common.
- Design guidance: thinking in graph terms (cycles) suggests how to design data collection or treatment assignment to increase identifiability—pack cycles or create movement that produces cycle-space variation. The capture–granularity trade-off is directly relevant to clustering choices in practice: coarser clusters may increase captured variation but can destroy randomization granularity and power.
- Algorithmic/statistical complementarity: these exact, design-based tests provide a robustness layer that complements ML-based heterogeneity/exogeneity diagnostics. They can be used to validate inference from models that rely on asymptotic approximations, or as replacements when asymptotics fail due to concentration.
- Accessibility: the paper supplies code (PanelAdequacy.jl, panelcert) and practical algorithms for cycle-packing, so applied AI economists can compute diagnostics and run exact tests on large panels or matched datasets.
- Caveats: exactness requires blockwise symmetry (or exchangeability via antithetic pairing) at the support level—practitioners must judge plausibility given the error-generating process (e.g., residual interactive fixed effects may violate the assumption). The method handles random-effect-like within-col reinterpretations that lie in the fixed-effect column space without cost, but complex cross-block dependence remains outside its scope.
Suggested practical workflow for AI-econ applied work: 1. Compute design diagnostics (λn, Hn, Neff, and score-based analogues). 2. If concentration is substantial (high λn or small Neff), construct annihilating contrasts; in two-way settings, attempt cycle packing targeting high κC while keeping supports compatible with plausible error-blocking. 3. Run the sign-flip randomization test (exact or Monte Carlo) and invert for confidence intervals. 4. Report diagnostics, κC, chosen support granularity, and sensitivity to block partitioning (e.g., observation vs match vs firm-level blocks).
Overall, this paper supplies both a diagnostic lens and a practically implementable, finite-sample valid inference procedure for the concentrated-identifying-variation regime common in modern panel and networked datasets relevant to AI economics.
Assessment
Claims (12)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| The self-normalized leverage of residualized treatment, λn = max_i(ex_i^2/Vn), rather than aggregate identifying variation, governs the reliability of Gaussian inference in saturated fixed-effect regressions. Decision Quality | negative | Reliability of Gaussian-based statistical inference |
Reading fidelity
high
Study strength
high
|
not reported
|
| At the fully concentrated boundary, the limiting null distribution of the studentized score varies across symmetric error distributions with the same variance, so no fixed critical value is uniformly valid over that class. Decision Quality | negative | Uniform asymptotic size control of fixed-critical-value tests |
Reading fidelity
high
Study strength
high
|
not reported
|
| Under blockwise central symmetry and independent blocks, sign-flip randomization using design-based nuisance-annihilating contrasts is finite-sample exact under the null, even with arbitrary heteroskedasticity and without identical-distribution or moment conditions. Decision Quality | positive | Finite-sample null rejection probability of the randomization test |
Reading fidelity
high
Study strength
high
|
Exact conditional size control for every n
|
| In two-way fixed-effect designs, the space of nuisance-annihilating contrasts is exactly the cycle space of the associated bipartite observation multigraph. Decision Quality | positive | Availability and characterization of valid fixed-effect-annihilating contrasts |
Reading fidelity
high
Study strength
high
|
not reported
|
| The observable capture ratio κC equals the Pitman efficiency of the contrast-based test relative to the infeasible oracle Gaussian test in the diffuse benchmark. Decision Quality | positive | Asymptotic testing efficiency relative to an oracle Gaussian test |
Reading fidelity
high
Study strength
high
|
κC^−1/2 asymptotic standard-error price
|
| Capture reaches one when supports are merged to the granularity of biconnected blocks, but coarse supports can reduce power because the randomization group has only one sign per support. Decision Quality | mixed | Capture of residualized treatment variation and randomization-test power |
Reading fidelity
high
Study strength
high
|
κC = 1 at block granularity
|
| On the public Kline et al. (2020) matched employer–employee data, the structure-exploiting cycle-packing algorithm achieves κC = 0.51 for a match-level treatment, compared with 0.26 for greedy packing. Decision Quality | positive | Treatment variation captured by valid contrast supports |
Reading fidelity
high
Study strength
medium
|
κC = 0.51 versus 0.26
|
| In the concentrated Monte Carlo design with λn = 0.29 and effective sample size 6.0, conventional degrees-of-freedom-corrected t-tests reject a true null at rates up to 58.5%, while HC2 t-tests reject at rates up to 33.4% under symmetric heteroskedastic errors. Error Rate | negative | Empirical type-I error rate of conventional t-tests |
Reading fidelity
high
Study strength
medium
|
up to 58.5% for df-corrected t; up to 33.4% for HC2 t
|
| In every simulated configuration, the cycle randomization test has empirical size between 4.9% and 5.0% under the null. Error Rate | positive | Empirical type-I error rate of the cycle randomization test |
Reading fidelity
high
Study strength
medium
|
4.9–5.0% empirical size
|
| After oracle size correction in the two heteroskedastic simulation designs, cycle-test power ranges from 49.5% to 52.0%, compared with 27.5–49.6% for the degrees-of-freedom t-test and 23.5–51.1% for HC2. Decision Quality | positive | Power against alternatives after size correction |
Reading fidelity
high
Study strength
medium
|
cycle test 49.5–52.0%; df-t 27.5–49.6%; HC2 23.5–51.1%
|
| Across 49 prespecified positive regressor pairs from ten public panel distributions, taking logarithms lowers λn in 37 pairs, or 75.5% of the comparisons; the paired mean falls from 0.170 to 0.097. Automation Exposure | positive | Residualized-treatment concentration as measured by λn |
Reading fidelity
high
Study strength
medium
|
n=49
λn lower in 37 of 49 pairs (75.5%); paired mean 0.170 to 0.097
|
| In the canonical Grunfeld investment regression, realized score concentration is 0.739 with an effective score sample size of 1.80, while 32 valid nuisance-annihilating supports capture κC = 0.627; the resulting exact 95% confidence interval is [0.150, 0.450]. Decision Quality | positive | Treatment-effect inference and confidence interval under fixed effects |
Reading fidelity
high
Study strength
medium
|
n=32
score concentration 0.739; Neff^score = 1.80; κC = 0.627; exact 95% interval [0.150, 0.450]
|