The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Staggered-DiD estimators can badly over-reject when clusters are few or imbalanced; a cluster-jackknife correction substantially improves standard errors and inference and is available in csdidjack/didjack software.

Improved Inference for CSDID Using the Cluster Jackknife
Sunny R. Karim, Morten Ørregaard Nielsen, James G. MacKinnon, Matthew D. Webb · February 12, 2026
arxiv theoretical n/a evidence 7/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

Arxiv

Latest observation:

  1. Sunny R. Karim unresolved corpus identity
  2. Morten Ørregaard Nielsen unresolved corpus identity
  3. James G. MacKinnon unresolved corpus identity
  4. Matthew D. Webb unresolved corpus identity

Semantic Scholar

Latest observation:

  1. Sunny Karim provider ID
  2. Morten Ørregaard Nielsen provider ID
  3. J. MacKinnon provider ID
  4. Matthew D. Webb provider ID
The paper shows that popular staggered-DiD estimators (CSDID) suffer from severe over-rejection in inference when clusters or treated clusters are few or imbalanced and demonstrates that a cluster-jackknife correction markedly improves inference, providing Stata and R implementations.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

Obtaining reliable inferences with traditional difference-in-differences (DiD) methods can be difficult. Problems can arise when both outcomes and errors are serially correlated, when there are few clusters or few treated clusters, when cluster sizes vary greatly, and in various other cases. In recent years, recognition of the ``staggered adoption'' problem has shifted the focus away from inference towards consistent estimation of treatment effects. One of the most popular new estimators is the CSDID procedure of Callaway and Sant'Anna (2021). We find that the issues of over-rejection with few clusters and/or few treated clusters are at least as severe for CSDID as for traditional DiD methods. We also propose using a cluster jackknife for inference with CSDID, which simulations suggest greatly improves inference. We provide software packages in Stata csdidjack and R didjack to calculate cluster-jackknife standard errors easily.

Summary

Main Finding

CSDID (Callaway & Sant’Anna, 2021), a widely used estimator for staggered DiD designs, suffers the same finite-sample inference problems as conventional TWFE DiD — notably severe over-rejection when clusters are few, when few clusters are treated, when cluster sizes vary, and when outcomes/errors are serially correlated. Applying a cluster jackknife to CSDID substantially improves inference in Monte Carlo evidence. The authors provide easy-to-use implementations (Stata: csdidjack; R: didjack).

Key Points

  • Problem statement
    • Recent work focused on consistent estimation under staggered adoption (e.g., CSDID) but paid less attention to finite-sample inference.
    • Cluster dependence (serial correlation within regions), few clusters, few treated clusters, and heterogeneous cluster sizes can all induce over-rejection of standard tests when using existing CSDID inference methods.
  • What the paper shows
    • Simulation evidence indicates CSDID is at least as vulnerable as TWFE methods to over-rejection under realistic clustering problems.
    • A cluster jackknife applied to the aggregated CSDID ATT greatly reduces size distortion in many designs.
    • The cluster jackknife does not completely solve problems that arise when the number of treated clusters is very small — some settings remain challenging.
  • Practical contribution
    • Implementation: Stata package csdidjack and R package didjack for computing cluster-jackknife standard errors for CSDID estimates.
    • The jackknife is computationally straightforward, conceptually simple (leave-one-cluster-out recomputation), and integrates with the usual CSDID two-step procedure (estimate group × time 2×2 ATT(g,t) blocks, then aggregate).
  • Relation to other methods
    • Builds on literature showing cluster-robust variance estimators and wild cluster bootstraps can fail with few clusters (Bertrand et al. 2004; Cameron et al. 2008; MacKinnon et al. 2023).
    • The cluster jackknife is presented as a complementary/superior finite-sample inference tool for CSDID in many realistic cases; but it is not a panacea for extremely small numbers of treated clusters.

Data & Methods

  • Estimator studied
    • CSDID: constructs 2×2 DiD comparisons by cohort (first-treatment time g) vs. chosen comparison groups (never-treated or not-yet-treated), estimates ATT(g,t) using outcomes in periods g−1 and t, then aggregates via user-specified weights to an overall ATT.
  • Inference approaches compared
    • Standard cluster-robust variance estimators (CRVE) applied to the CSDID aggregated ATT.
    • Wild cluster bootstrap variants in the literature.
    • Cluster jackknife: leave-one-cluster-out recalculation of the aggregated ATT; standard jackknife variance formula across cluster leave-outs used to form SEs and t-tests.
  • Monte Carlo design (overview)
    • Simulations with staggered adoption designs (various numbers of clusters, treated clusters, and time periods).
    • Designs include serially correlated outcomes and errors, unbalanced cluster sizes, and cases with/without never-treated controls.
    • Performance metrics: empirical rejection rates (size), power, and coverage of confidence intervals.
  • Main simulation findings (summary)
    • CRVE and some bootstrap methods display severe over-rejection (empirical sizes well above nominal) in many small-cluster or few-treated-cluster settings.
    • Cluster jackknife markedly reduces over-rejection in most simulated designs, improving size control and interval coverage.
    • Residual failure modes remain when the number of treated clusters is extremely small, where even the jackknife can be unreliable.

Implications for AI Economics

  • Prevalence of staggered adoption in AI policy and adoption work
    • Many AI-economics studies use staggered DiD designs to evaluate policy changes, AI deployments across regions/firms, or phased rollouts of AI-based programs. These designs often cluster at region/firm level and can have few treated clusters.
  • Practical guidance
    • Use modern, heterogeneous-robust estimators (e.g., CSDID) to avoid bias from forbidden comparisons and heterogeneous treatment effects.
    • For inference, augment CSDID with the cluster jackknife (csdidjack/didjack) to improve finite-sample validity — especially important when:
      • Total clusters are modest (e.g., dozens or fewer),
      • Cluster sizes vary a lot,
      • Outcomes/errors are serially correlated.
    • Report details that matter for inference: number of clusters, number of treated clusters, whether never-treated controls exist, and sensitivity checks (jackknife vs. CRVE vs. bootstrap).
  • When to be cautious
    • If the number of treated clusters is very small, be sceptical of conventional p-values and even jackknife-adjusted inference. Consider alternative or complementary approaches: randomization/permutation tests where feasible, transparent reporting of group-time ATT components, or conservative interpretation of results.
  • Research agenda for AI economics
    • Further work is needed on robust inference when treated clusters are few (the most common remaining failure mode).
    • Development and adoption of inference diagnostics and standardized reporting (e.g., showing ATT(g,t) components, cluster counts) would raise credibility in empirical AI-economics work.
    • The availability of csdidjack and didjack makes it straightforward to adopt better inference practices in applied AI-economics papers.

References cited/related (select) - Callaway, B. & Sant’Anna, P. (2021). Difference-in-Differences with multiple time periods. Journal of Econometrics. - Goodman-Bacon, A. (2021). Difference-in-differences with variation in treatment timing. - Bertrand, M., Duflo, E., & Mullainathan, S. (2004). How much should we trust differences-in-differences estimates? - Cameron, A. C., Gelbach, J. B., & Miller, D. L. (2008). Bootstrap-based improvements for inference with clustered errors. - MacKinnon, J. G., Nielsen, M. Ø., & Webb, M. D. (2023). Cluster-jackknife and related methods for clustered inference.

If you want, I can: - Provide the cluster-jackknife variance formula and a short example of how it is computed for the aggregated CSDID ATT; or - Show a short code snippet demonstrating csdidjack (Stata) or didjack (R) usage with typical options.

Assessment

Paper Typetheoretical Evidence Strengthn/a — The paper is methodological: it does not present new causal empirical findings about economic outcomes but evaluates inference procedures via simulations and theoretical arguments, so traditional evidence-strength grading for causal claims is not applicable. Methods Rigorhigh — Builds on state-of-the-art DiD identification (Callaway & Sant'Anna), systematically documents failure modes (few clusters, few treated clusters, serial correlation, cluster-size heterogeneity), evaluates performance across Monte Carlo designs, and proposes a standard, well-motivated resampling correction (cluster jackknife) with software implementations—showing substantial simulation improvements relative to common alternatives. SampleMonte Carlo simulation experiments varying number of clusters, number of treated clusters, cluster-size heterogeneity, outcome and error serial correlation, and treatment timing; no primary real-world empirical dataset reported in the abstract; software packages (Stata csdidjack and R didjack) implement the estimator and jackknife inference. Themesadoption productivity IdentificationUses staggered-adoption difference-in-differences identification (group-time average treatment effects under parallel trends conditional on covariates, no anticipatory effects, SUTVA). Estimation follows Callaway and Sant'Anna (2021) CSDID for consistent ATT estimation under heterogeneous treatment timing; causal inference is achieved by comparing treated and not-yet-treated groups appropriately and aggregating group-time effects. For inference the paper proposes a cluster jackknife (leave-one-cluster-out) procedure to approximate the sampling distribution of the CSDID estimator and correct standard errors in presence of few clusters, cluster-size imbalance, and serial correlation. GeneralizabilitySimulation results depend on the range of data-generating processes considered and may not cover all empirically relevant DGPs., Cluster jackknife assumes independent clusters; performance may deteriorate with cross-cluster spillovers or strong cross-cluster dependence., Very small numbers of clusters (e.g., <5) or extreme imbalance may still limit reliable inference., Does not address violations of parallel trends or pervasive time-varying unobserved confounders; consistent estimation still requires standard DiD identification assumptions., Performance with very long panels, nonstationary outcomes, or complex missing data patterns is not guaranteed without further validation.

Claims (6)

ClaimDirectionOutcomeConfidence & EvidenceDetails
Obtaining reliable inferences with traditional difference-in-differences (DiD) methods can be difficult when outcomes and errors are serially correlated, when there are few clusters or few treated clusters, when cluster sizes vary greatly, and in various other cases. Research Productivity negative reliability of statistical inference (e.g., Type I error control) in DiD designs under serial correlation, few clusters, and heterogeneous cluster sizes
Reading fidelity high
Study strength medium
not reported
0.12
Recognition of the 'staggered adoption' problem has shifted the focus in recent years away from inference towards consistent estimation of treatment effects. Research Productivity mixed research focus/priorities in the econometrics literature (estimation vs inference) concerning staggered adoption designs
Reading fidelity high
Study strength medium
not reported
0.12
One of the most popular new estimators for staggered adoption designs is the CSDID procedure of Callaway and Sant'Anna (2021). Research Productivity positive popularity/adoption of the CSDID estimator in applied/stated methodological work
Reading fidelity high
Study strength low
not reported
0.06
The issues of over-rejection with few clusters and/or few treated clusters are at least as severe for CSDID as for traditional DiD methods. Research Productivity negative over-rejection rate (Type I error rate) of inferential procedures for CSDID versus traditional DiD
Reading fidelity high
Study strength medium
not reported
0.12
Using a cluster jackknife for inference with CSDID greatly improves inference (according to simulations). Research Productivity positive improvement in inference quality (e.g., better Type I error control, coverage) when using cluster-jackknife standard errors with CSDID
Reading fidelity high
Study strength medium
greatly improves inference
0.12
We provide software packages in Stata (csdidjack) and R (didjack) to calculate cluster-jackknife standard errors easily. Research Productivity positive availability of software tools for implementing cluster-jackknife inference with CSDID
Reading fidelity high
Study strength high
not reported
0.2

Notes