0 cumulative citations
View corpus contextStaggered-DiD estimators can badly over-reject when clusters are few or imbalanced; a cluster-jackknife correction substantially improves standard errors and inference and is available in csdidjack/didjack software.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
Obtaining reliable inferences with traditional difference-in-differences (DiD) methods can be difficult. Problems can arise when both outcomes and errors are serially correlated, when there are few clusters or few treated clusters, when cluster sizes vary greatly, and in various other cases. In recent years, recognition of the ``staggered adoption'' problem has shifted the focus away from inference towards consistent estimation of treatment effects. One of the most popular new estimators is the CSDID procedure of Callaway and Sant'Anna (2021). We find that the issues of over-rejection with few clusters and/or few treated clusters are at least as severe for CSDID as for traditional DiD methods. We also propose using a cluster jackknife for inference with CSDID, which simulations suggest greatly improves inference. We provide software packages in Stata csdidjack and R didjack to calculate cluster-jackknife standard errors easily.
Summary
Main Finding
CSDID (Callaway & Sant’Anna, 2021), a widely used estimator for staggered DiD designs, suffers the same finite-sample inference problems as conventional TWFE DiD — notably severe over-rejection when clusters are few, when few clusters are treated, when cluster sizes vary, and when outcomes/errors are serially correlated. Applying a cluster jackknife to CSDID substantially improves inference in Monte Carlo evidence. The authors provide easy-to-use implementations (Stata: csdidjack; R: didjack).
Key Points
- Problem statement
- Recent work focused on consistent estimation under staggered adoption (e.g., CSDID) but paid less attention to finite-sample inference.
- Cluster dependence (serial correlation within regions), few clusters, few treated clusters, and heterogeneous cluster sizes can all induce over-rejection of standard tests when using existing CSDID inference methods.
- What the paper shows
- Simulation evidence indicates CSDID is at least as vulnerable as TWFE methods to over-rejection under realistic clustering problems.
- A cluster jackknife applied to the aggregated CSDID ATT greatly reduces size distortion in many designs.
- The cluster jackknife does not completely solve problems that arise when the number of treated clusters is very small — some settings remain challenging.
- Practical contribution
- Implementation: Stata package csdidjack and R package didjack for computing cluster-jackknife standard errors for CSDID estimates.
- The jackknife is computationally straightforward, conceptually simple (leave-one-cluster-out recomputation), and integrates with the usual CSDID two-step procedure (estimate group × time 2×2 ATT(g,t) blocks, then aggregate).
- Relation to other methods
- Builds on literature showing cluster-robust variance estimators and wild cluster bootstraps can fail with few clusters (Bertrand et al. 2004; Cameron et al. 2008; MacKinnon et al. 2023).
- The cluster jackknife is presented as a complementary/superior finite-sample inference tool for CSDID in many realistic cases; but it is not a panacea for extremely small numbers of treated clusters.
Data & Methods
- Estimator studied
- CSDID: constructs 2×2 DiD comparisons by cohort (first-treatment time g) vs. chosen comparison groups (never-treated or not-yet-treated), estimates ATT(g,t) using outcomes in periods g−1 and t, then aggregates via user-specified weights to an overall ATT.
- Inference approaches compared
- Standard cluster-robust variance estimators (CRVE) applied to the CSDID aggregated ATT.
- Wild cluster bootstrap variants in the literature.
- Cluster jackknife: leave-one-cluster-out recalculation of the aggregated ATT; standard jackknife variance formula across cluster leave-outs used to form SEs and t-tests.
- Monte Carlo design (overview)
- Simulations with staggered adoption designs (various numbers of clusters, treated clusters, and time periods).
- Designs include serially correlated outcomes and errors, unbalanced cluster sizes, and cases with/without never-treated controls.
- Performance metrics: empirical rejection rates (size), power, and coverage of confidence intervals.
- Main simulation findings (summary)
- CRVE and some bootstrap methods display severe over-rejection (empirical sizes well above nominal) in many small-cluster or few-treated-cluster settings.
- Cluster jackknife markedly reduces over-rejection in most simulated designs, improving size control and interval coverage.
- Residual failure modes remain when the number of treated clusters is extremely small, where even the jackknife can be unreliable.
Implications for AI Economics
- Prevalence of staggered adoption in AI policy and adoption work
- Many AI-economics studies use staggered DiD designs to evaluate policy changes, AI deployments across regions/firms, or phased rollouts of AI-based programs. These designs often cluster at region/firm level and can have few treated clusters.
- Practical guidance
- Use modern, heterogeneous-robust estimators (e.g., CSDID) to avoid bias from forbidden comparisons and heterogeneous treatment effects.
- For inference, augment CSDID with the cluster jackknife (csdidjack/didjack) to improve finite-sample validity — especially important when:
- Total clusters are modest (e.g., dozens or fewer),
- Cluster sizes vary a lot,
- Outcomes/errors are serially correlated.
- Report details that matter for inference: number of clusters, number of treated clusters, whether never-treated controls exist, and sensitivity checks (jackknife vs. CRVE vs. bootstrap).
- When to be cautious
- If the number of treated clusters is very small, be sceptical of conventional p-values and even jackknife-adjusted inference. Consider alternative or complementary approaches: randomization/permutation tests where feasible, transparent reporting of group-time ATT components, or conservative interpretation of results.
- Research agenda for AI economics
- Further work is needed on robust inference when treated clusters are few (the most common remaining failure mode).
- Development and adoption of inference diagnostics and standardized reporting (e.g., showing ATT(g,t) components, cluster counts) would raise credibility in empirical AI-economics work.
- The availability of csdidjack and didjack makes it straightforward to adopt better inference practices in applied AI-economics papers.
References cited/related (select) - Callaway, B. & Sant’Anna, P. (2021). Difference-in-Differences with multiple time periods. Journal of Econometrics. - Goodman-Bacon, A. (2021). Difference-in-differences with variation in treatment timing. - Bertrand, M., Duflo, E., & Mullainathan, S. (2004). How much should we trust differences-in-differences estimates? - Cameron, A. C., Gelbach, J. B., & Miller, D. L. (2008). Bootstrap-based improvements for inference with clustered errors. - MacKinnon, J. G., Nielsen, M. Ø., & Webb, M. D. (2023). Cluster-jackknife and related methods for clustered inference.
If you want, I can: - Provide the cluster-jackknife variance formula and a short example of how it is computed for the aggregated CSDID ATT; or - Show a short code snippet demonstrating csdidjack (Stata) or didjack (R) usage with typical options.
Assessment
Claims (6)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| Obtaining reliable inferences with traditional difference-in-differences (DiD) methods can be difficult when outcomes and errors are serially correlated, when there are few clusters or few treated clusters, when cluster sizes vary greatly, and in various other cases. Research Productivity | negative | reliability of statistical inference (e.g., Type I error control) in DiD designs under serial correlation, few clusters, and heterogeneous cluster sizes |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Recognition of the 'staggered adoption' problem has shifted the focus in recent years away from inference towards consistent estimation of treatment effects. Research Productivity | mixed | research focus/priorities in the econometrics literature (estimation vs inference) concerning staggered adoption designs |
Reading fidelity
high
Study strength
medium
|
not reported
|
| One of the most popular new estimators for staggered adoption designs is the CSDID procedure of Callaway and Sant'Anna (2021). Research Productivity | positive | popularity/adoption of the CSDID estimator in applied/stated methodological work |
Reading fidelity
high
Study strength
low
|
not reported
|
| The issues of over-rejection with few clusters and/or few treated clusters are at least as severe for CSDID as for traditional DiD methods. Research Productivity | negative | over-rejection rate (Type I error rate) of inferential procedures for CSDID versus traditional DiD |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Using a cluster jackknife for inference with CSDID greatly improves inference (according to simulations). Research Productivity | positive | improvement in inference quality (e.g., better Type I error control, coverage) when using cluster-jackknife standard errors with CSDID |
Reading fidelity
high
Study strength
medium
|
greatly improves inference
|
| We provide software packages in Stata (csdidjack) and R (didjack) to calculate cluster-jackknife standard errors easily. Research Productivity | positive | availability of software tools for implementing cluster-jackknife inference with CSDID |
Reading fidelity
high
Study strength
high
|
not reported
|