The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Concentrated within-variation can break t-tests in saturated fixed-effect regressions, and no single critical value cures the boundary failure; constructing design-only contrasts (cycle-space sign-flips) gives finite-sample exact p-values under symmetric errors and delivers practical power via a cycle-packing algorithm.

Exact Inference in Fixed-Effect Regressions with Concentrated Identifying Variation
Stanisław M. S. Halkiewicz · August 05, 2026
arxiv theoretical medium evidence 7/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

Arxiv

Latest observation:

  1. Stanisław M. S. Halkiewicz unresolved corpus identity

Semantic Scholar

Latest observation:

  1. S. Halkiewicz provider ID
When identifying variation in saturated fixed-effect regressions is highly concentrated, conventional Gaussian t-tests can fail and no fixed critical value is uniformly valid; design-based, nuisance-annihilating sign-flip contrasts (the cycle-space construction for two-way panels) yield finite-sample exact inference under blockwise symmetric errors and provide an observable Pitman-like efficiency κC with practical packing algorithms.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

In saturated fixed-effects regressions, Gaussian inference depends not on total identifying variation but on its concentration, measured by the self-normalized leverage $λ_n$ of the residualized treatment. When finitely many score weights remain persistent, the $t$-statistic converges to a convolution of raw errors and a Gaussian component. At full concentration, its null distribution varies across symmetric error laws with equal variance, so no fixed critical value is uniformly valid. We instead construct nuisance-annihilating contrasts from the design alone. These eliminate the fixed effects identically and yield finite-sample exact sign-flip inference under symmetric, arbitrarily heteroskedastic errors, with no homogeneity assumptions or restrictions on the fixed-effect dimension. In two-way designs, admissible contrasts form the cycle space of the observation multigraph. Their efficiency is summarized by an observable capture ratio $κ$, which equals Pitman efficiency. The resulting design problem involves a capture--granularity trade-off: coarse supports maximize capture but reduce the number of randomization signs. Cycle packing provides sufficiently granular supports. On matched employer--employee data, a structure-exploiting algorithm achieves $κ\approx 0.51$, compared with $0.26$ for naive packing. In the Grunfeld investment regression, realized score concentration is $0.739$, corresponding to $N_{\mathrm{eff}}^{\mathrm{score}}=1.80$, while $32$ valid supports attain $κ=0.627$. The resulting exact $95%$ confidence interval is $[0.150,0.450]$. A worker--firm application demonstrates scalability to large networks.

Summary

Main Finding

When most treatment variation is absorbed by many fixed effects, conventional Gaussian (t-test) inference can fail because identification depends on how residualized treatment variation is concentrated across observations, not its aggregate size. Halkiewicz develops a finite-sample, design-based alternative: construct nuisance-annihilating contrasts (q with q′D = 0) from the design, use sign-flip randomization on disjoint supports to obtain exact inference under symmetric (and blockwise symmetric) errors, and—for two-way (worker–firm, firm–time) designs—shows these contrasts correspond exactly to the cycle space of the observation multigraph. He introduces an observable capture ratio κC (a Pitman-efficiency analogue) that governs power, provides cycle-packing algorithms to produce high-capture granular supports, and supplies diagnostics (λn, Hn, Neff, score-based analogues) and software. Conventional leverage-corrected tests can severely over-reject under concentrated identifying variation; the cycle/sign-flip tests are exact and competitive in power when appropriately packed.

Key Points

  • Concentration parameter (design-only): λn = maxi exi^2 / Vn (where ex = MDx, Vn = ex′ex). λn → 0 is the usual negligibility condition for asymptotic normality; when λn is not small, the asymptotic null may be a convolution mixing few large-weighted errors and a Gaussian remainder (Proposition 2.1).
  • Impossibility at the boundary (Proposition 2.3): At the fully concentrated boundary (one persistent weight), the limiting null law depends on the unknown symmetric error distribution, so no fixed critical value (tabulated test) is uniformly valid across symmetric equal-variance laws. Valid inference must adapt to the realized error law.
  • Nuisance-annihilating contrasts: any q measurable from (x, D) with q′D = 0 extracts a linear combination free of fixed effects: q′(Y − xβ0) = q′ε under H0. If contrasts have disjoint supports and errors are blockwise independent and symmetric across those blocks (Assumption 2), flipping signs on contrasts yields exact finite-sample randomization tests (Theorem 3.3).
  • Two-way designs ↔ cycle space: In bipartite observation graphs (edges = observations, vertices = firms/workers or firm/time), the space of annihilating contrasts equals the graph’s cycle space. The within-projection norm Vn = ||ΠZ x||^2 is the squared projection of treatment onto the cycle space (Proposition 4.1).
  • Capture ratio κC: For a system of C disjoint supports with local normalized projections vc, capture is κC = (1/Vn) ∑c ||ΠZ_Ac x||^2 ∈ [0,1]. In diffuse regimes κC is the Pitman relative efficiency; κC^−1/2 is the multiplicative standard-error price (Theorem 3.9).
  • Capture–granularity trade-off: Merging supports increases capture (can reach κC = 1 by taking one projection per biconnected block) but reduces the randomization group size (one sign per support), possibly wrecking power because the attainable minimal p-value grows (orbit-size floor). Fine, granular supports preserve randomization granularity but may lower capture. The paper formalizes and prices this trade-off.
  • Cycle packing algorithms: constructs (digons, edge-disjoint firm-pair four-cycles, recursive contraction) that produce granular, high-capture supports. On public matched employer–employee extract, cycle packing achieves κC ≈ 0.51 (match-level treatment) vs 0.26 for naive greedy packing; for a time-varying covariate it achieves 0.91 vs 0.88.
  • Empirical diagnostics & examples:
    • Grunfeld investment regression: realized score concentration λn = 0.739 (N_score_eff = 1.80); 32 valid supports capture κC = 0.627; exact 95% CI is [0.150, 0.450].
    • Monte Carlo: with concentrated design (λn = 0.29, effective sample 6.0), df-corrected t rejects up to 58.5% (size badly inflated), HC2 up to 33.4%; cycle test empirical size ~5% everywhere; cycle-test power often higher once oracle size correction applied to conventional tests.
  • Software and reproducibility: implementations in PanelAdequacy.jl and an R companion (panelcert) are provided.

Data & Methods

  • Model and conditioning: Linear model Yi = xiβ + di′γ + εi, conditional on the design (x, D). Focus is on inference for β with γ unrestricted nuisance.
  • Key design objects:
    • ex = MDx (treatment residualized on fixed effects)
    • Vn = ex′ex
    • λn = maxi exi^2 / Vn (self-normalized leverage)
    • Herfindahl Hn = ∑(exi^2 / Vn)^2, Neff = 1/Hn
    • Score-based realized analogues using fitted residuals (bλscore, bNscore_eff) as warnings.
  • Limit theory: Proposition 2.1 gives convolutional limits when a finite number of weights persist; Corollary 2.2 shows studentization cannot generally restore normality.
  • Exact-test construction:
    • Build annihilating contrast system {q1,...,qC} with q′c D = 0 and pairwise disjoint supports Sc (measurable from design).
    • Under Assumption 2 (blocks of observations that are independent and centrally symmetric; supports unions of blocks), contrast scores Uc = vc′(Y − xβ0) are invariant to flipping signs on each support.
    • Define test-statistic T(U1,...,UC) (default T = |∑c bc Uc| with bc = vc′x) and compute randomization p-value by enumerating or Monte Carlo sampling sign patterns s ∈ {±1}^C (Theorem 3.3). Studentized variants are also exact.
  • Two-way graph methods:
    • Observations are edges of bipartite multigraph. Annihilating contrasts correspond to ±1 alternations around cycles (including digons for repeated matches).
    • Projection of x onto cycle space gives Vn; local projections on supports give the captured portion.
    • Cycle packing: algorithmic constructions (digons, firm-pair 4-cycles, recursive contraction) to get many edge-disjoint or nearly edge-disjoint cycles to maximize κC with maintained granularity.
  • Monte Carlo setups:
    • Simulation designs include concentrated and diffuse regimes, heteroskedastic symmetric errors, and dependence structures consistent with Assumption 2 at various granularities.
    • Benchmarks included df-corrected t, HC2, and oracle-adjusted versions.

Implications for AI Economics

  • Many AI-economics settings have high-dimensional fixed effects and networked observations (platform-user, user-item, firm-worker, time-varying policy adoption). When controlling for those effects, remaining identifying variation for treatment or instrument often concentrates on few observations; this paper shows standard asymptotic t- or cluster-robust inference can be invalid in that regime.
  • Practical diagnostics: report λn, Hn, Neff and the score-based analogues (bλscore, bNscore_eff). If effective sample size is small or λn large, conventional inference is suspect and the exact randomization methods are a principled alternative.
  • Exact randomization tests adapt to the realized (unknown) error law and remain valid under arbitrary heteroskedasticity so long as blockwise symmetry/exchangeability assumptions are plausible at the support granularity chosen. This is useful for AI policy evaluation (e.g., A/B tests with fixed effects, algorithmic interventions with network spillovers) where dependence and heteroskedasticity are common.
  • Design guidance: thinking in graph terms (cycles) suggests how to design data collection or treatment assignment to increase identifiability—pack cycles or create movement that produces cycle-space variation. The capture–granularity trade-off is directly relevant to clustering choices in practice: coarser clusters may increase captured variation but can destroy randomization granularity and power.
  • Algorithmic/statistical complementarity: these exact, design-based tests provide a robustness layer that complements ML-based heterogeneity/exogeneity diagnostics. They can be used to validate inference from models that rely on asymptotic approximations, or as replacements when asymptotics fail due to concentration.
  • Accessibility: the paper supplies code (PanelAdequacy.jl, panelcert) and practical algorithms for cycle-packing, so applied AI economists can compute diagnostics and run exact tests on large panels or matched datasets.
  • Caveats: exactness requires blockwise symmetry (or exchangeability via antithetic pairing) at the support level—practitioners must judge plausibility given the error-generating process (e.g., residual interactive fixed effects may violate the assumption). The method handles random-effect-like within-col reinterpretations that lie in the fixed-effect column space without cost, but complex cross-block dependence remains outside its scope.

Suggested practical workflow for AI-econ applied work: 1. Compute design diagnostics (λn, Hn, Neff, and score-based analogues). 2. If concentration is substantial (high λn or small Neff), construct annihilating contrasts; in two-way settings, attempt cycle packing targeting high κC while keeping supports compatible with plausible error-blocking. 3. Run the sign-flip randomization test (exact or Monte Carlo) and invert for confidence intervals. 4. Report diagnostics, κC, chosen support granularity, and sensitivity to block partitioning (e.g., observation vs match vs firm-level blocks).

Overall, this paper supplies both a diagnostic lens and a practically implementable, finite-sample valid inference procedure for the concentrated-identifying-variation regime common in modern panel and networked datasets relevant to AI economics.

Assessment

Paper Typetheoretical Evidence Strengthmedium — The paper presents rigorous theoretical results (limit theorems, impossibility at the concentration boundary, finite-sample exactness), simulation evidence demonstrating size/power performance under a range of designs, and empirical illustrations (Grunfeld, matched employer–employee data). It lacks large-scale empirical causal applications but provides strong theoretical justification and practical demonstrations. Methods Rigorhigh — The paper offers formal propositions and theorems with proofs (convolution limits, impossibility result, exactness of sign-flip tests, cycle-space characterization, Pitman-efficiency analogue κC), provides algorithmic constructions for cycle packing with performance guarantees, and supplies Monte Carlo experiments and reproducible code; assumptions and limitations (notably symmetry and block structure) are stated and analyzed. SampleThe paper is primarily theoretical. Empirical material includes Monte Carlo experiments (two-way designs with concentrated and diffuse identifying variation; a concentrated design with λn = 0.29 and Neff = 6.0 is reported), an end-to-end concentrated Grunfeld specification, and applications to public matched employer–employee data (Kline et al. 2020) showing cycle-packing performance (e.g., κC ≈ 0.51 for match-level treatment). Code and replication materials are provided in public repositories (PanelAdequacy.jl and an R package). Themeslabor_markets productivity IdentificationIdentification of the scalar treatment β relies on within-unit (residualized) variation after absorbing high-dimensional fixed effects: the effective identifying variation is the projection of x onto the complement of the fixed-effect column space (ex = MD x); in two-way worker–firm panels identification requires mobility that maps to nonzero projection onto the graph's cycle space. The paper frames identification in terms of the concentration of the self-normalized leverage weights (λn) and uses design-measurable, nuisance-annihilating contrasts (q with q' D = 0) to isolate score variation that identifies β while eliminating fixed-effect nuisances. GeneralizabilityFinite-sample exactness requires blockwise central symmetry and independence across blocks (Assumption 2); violations (e.g., general serial dependence, interactive fixed effects coupling blocks) invalidate exactness., Method is conditional on design (x, D) and targets inference for β in saturated fixed-effect regressions; it does not solve endogeneity from omitted time-varying confounders or instrument identification., Power depends on the chosen packing/supports and capture–granularity trade-off; coarse supports can give high capture but low randomization granularity and poor power., Cycle-packing is NP-hard in general; practical algorithms approximate but may not attain oracle capture on very large, complex networks., Results are most directly applicable to one- and two-way fixed-effect panels and matched worker–firm designs; extensions to other dependence structures require care.

Claims (12)

ClaimDirectionOutcomeConfidence & EvidenceDetails
The self-normalized leverage of residualized treatment, λn = max_i(ex_i^2/Vn), rather than aggregate identifying variation, governs the reliability of Gaussian inference in saturated fixed-effect regressions. Decision Quality negative Reliability of Gaussian-based statistical inference
Reading fidelity high
Study strength high
not reported
0.2
At the fully concentrated boundary, the limiting null distribution of the studentized score varies across symmetric error distributions with the same variance, so no fixed critical value is uniformly valid over that class. Decision Quality negative Uniform asymptotic size control of fixed-critical-value tests
Reading fidelity high
Study strength high
not reported
0.2
Under blockwise central symmetry and independent blocks, sign-flip randomization using design-based nuisance-annihilating contrasts is finite-sample exact under the null, even with arbitrary heteroskedasticity and without identical-distribution or moment conditions. Decision Quality positive Finite-sample null rejection probability of the randomization test
Reading fidelity high
Study strength high
Exact conditional size control for every n
0.2
In two-way fixed-effect designs, the space of nuisance-annihilating contrasts is exactly the cycle space of the associated bipartite observation multigraph. Decision Quality positive Availability and characterization of valid fixed-effect-annihilating contrasts
Reading fidelity high
Study strength high
not reported
0.2
The observable capture ratio κC equals the Pitman efficiency of the contrast-based test relative to the infeasible oracle Gaussian test in the diffuse benchmark. Decision Quality positive Asymptotic testing efficiency relative to an oracle Gaussian test
Reading fidelity high
Study strength high
κC^−1/2 asymptotic standard-error price
0.2
Capture reaches one when supports are merged to the granularity of biconnected blocks, but coarse supports can reduce power because the randomization group has only one sign per support. Decision Quality mixed Capture of residualized treatment variation and randomization-test power
Reading fidelity high
Study strength high
κC = 1 at block granularity
0.2
On the public Kline et al. (2020) matched employer–employee data, the structure-exploiting cycle-packing algorithm achieves κC = 0.51 for a match-level treatment, compared with 0.26 for greedy packing. Decision Quality positive Treatment variation captured by valid contrast supports
Reading fidelity high
Study strength medium
κC = 0.51 versus 0.26
0.12
In the concentrated Monte Carlo design with λn = 0.29 and effective sample size 6.0, conventional degrees-of-freedom-corrected t-tests reject a true null at rates up to 58.5%, while HC2 t-tests reject at rates up to 33.4% under symmetric heteroskedastic errors. Error Rate negative Empirical type-I error rate of conventional t-tests
Reading fidelity high
Study strength medium
up to 58.5% for df-corrected t; up to 33.4% for HC2 t
0.12
In every simulated configuration, the cycle randomization test has empirical size between 4.9% and 5.0% under the null. Error Rate positive Empirical type-I error rate of the cycle randomization test
Reading fidelity high
Study strength medium
4.9–5.0% empirical size
0.12
After oracle size correction in the two heteroskedastic simulation designs, cycle-test power ranges from 49.5% to 52.0%, compared with 27.5–49.6% for the degrees-of-freedom t-test and 23.5–51.1% for HC2. Decision Quality positive Power against alternatives after size correction
Reading fidelity high
Study strength medium
cycle test 49.5–52.0%; df-t 27.5–49.6%; HC2 23.5–51.1%
0.12
Across 49 prespecified positive regressor pairs from ten public panel distributions, taking logarithms lowers λn in 37 pairs, or 75.5% of the comparisons; the paired mean falls from 0.170 to 0.097. Automation Exposure positive Residualized-treatment concentration as measured by λn
Reading fidelity high
Study strength medium
n=49
λn lower in 37 of 49 pairs (75.5%); paired mean 0.170 to 0.097
0.12
In the canonical Grunfeld investment regression, realized score concentration is 0.739 with an effective score sample size of 1.80, while 32 valid nuisance-annihilating supports capture κC = 0.627; the resulting exact 95% confidence interval is [0.150, 0.450]. Decision Quality positive Treatment-effect inference and confidence interval under fixed effects
Reading fidelity high
Study strength medium
n=32
score concentration 0.739; Neff^score = 1.80; κC = 0.627; exact 95% interval [0.150, 0.450]
0.12

Notes