0 cumulative citations
View corpus contextEvent-window studies that align on user-initiated AI prompts can misattribute ongoing task activity to the AI: in same-user cross-surface logs, engineered zero-effect timestamps reproduce the majority of the apparent post-event search and browsing lift, showing that single-surface event windows lack identifying power without additional assumptions or cross-surface controls.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
Studies of digital behavior often align users at moments they choose, such as opening an AI assistant, clicking a recommendation, or visiting a product page, and interpret higher activity afterward as an event effect. We show how this creates an endogenous time zero: the event occurs during an ongoing task episode, so the aligned curve can trace episode continuation rather than a response to the event. In same-user, cross-surface web logs, AI, shopping, news, coding, and reference events are all preceded by broad activity increases that peak before time zero. Our strongest test uses known-null timestamps that cause nothing. Among the 5.8% of AI responses meeting strict pre-event activity and washout criteria, these timestamps show 3.42 times the post-event search activity of a within-user placebo, compared with 4.32 times for real events. The fraction of excess reproduced by the known null falls from 0.56 at detectably active moments to -0.04 at quiet moments, where the design detects none. We formalize this episode-selection bias, prove that a single-surface event window cannot separate it from a genuine effect without additional assumptions, and show in zero-effect simulations why user fixed effects and coarse activity matching can fail: the confound is within-user and time-varying. We provide a diagnostic protocol, public-data benchmarks, and burstcheck, a lightweight audit tool. User-timed events may have real effects, but post-event volume does not identify them by default; studies should compare similar episodes with and without the event.
Summary
Main Finding
When studies align users on moments they choose (e.g., opening an AI assistant, clicking a recommendation) and interpret higher activity afterward as a causal “event effect,” much of the measured lift can instead be continuation of an already‑ongoing task episode. The paper formalizes this “endogenous time zero” / episode‑selection bias, shows it survives common within‑user fixes, demonstrates it on real cross‑surface web logs using a known‑null (pseudo‑event) test, proves a single‑surface event‑window cannot identify the effect without extra assumptions, and supplies diagnostics and a lightweight audit tool (burstcheck).
Key Points
- Problem statement
- Endogenous time zero: the aligned timestamp (time zero) is chosen by the same latent process (task/intent/episode intensity) that drives outcomes, creating a backdoor path from latent episode state to post‑event outcomes.
- Episode‑selection bias: user‑timed focal events are more likely to occur during latent task episodes; the episode’s own continuation can be mistaken for an event effect.
- Why common defenses fail
- The confounder is within‑user and time‑varying, so user fixed effects and within‑user comparisons do not remove it.
- Coarse matching (e.g., on recent activity, hour‑of‑day) or anchoring to page views can even exacerbate the bias by oversampling dense spells (inspection paradox / regression to the mean).
- Formal result
- Proposition (informal): No functional of single‑surface event‑aligned data identifies the event’s causal effect without additional assumptions (e.g., a proxy that responds to episode state but is unaffected by the focal event).
- Empirical demonstration (known‑null pseudo‑event experiment)
- Data: opt‑in cross‑surface panel recording same users’ browsing, search, and conversational‑AI events.
- Strict pre‑event construction: an “active landmark” is a moment with ≥5 non‑search, non‑AI page views in the prior 30 minutes (only pre‑t information used).
- Headline result (per‑user matched set, 60‑minute prior‑AI washout): real AI events show a post‑event search lift of 4.32× (95% CI [3.34, 5.29]) versus a within‑user placebo; known‑null timestamps (guaranteed to cause nothing) matched on pre‑event activity show a lift of 3.42× ([3.01, 3.84]). The known‑null reproduces ~0.73 of the real excess.
- Anchor‑class gradient: reproduced fraction falls with pre‑event quietness — landmark‑active anchors: 0.56 (±), mildly active: 0.43, quiet: −0.04. (31% of responses are landmark‑active.)
- Many standard adjustments (hour‑of‑day matching, richer matching on pre‑activity) do not eliminate the false association.
- Simulations and sensitivity
- Zero‑effect simulations explain why unit fixed effects and coarse activity matching fail: the confound is within‑user and time‑varying.
- Conditioning on cross‑surface proxies reduces confounding but does not automatically identify the effect unless they satisfy additional conditions (e.g., respond to episode state but are unaffected by the treatment).
- Tools and protocols
- Provides a diagnostic protocol, public plasmode benchmarks (real data with injected effects), and burstcheck, a light audit tool whose screening statistic is validated on an external event set.
Data & Methods
- Data: same‑user, cross‑surface web logs (browsing, search, conversational‑AI events) from an opt‑in panel; analyses use per‑user capping in some constructions and an equal‑user matched set in the headline.
- Primary known‑null construction (strictly pre‑t):
- Define active landmarks using only pre‑t activity (≥5 non‑search/non‑AI page views in prior 30 minutes).
- Real moments: AI responses that are landmarks. Pseudo moments: uniform‑random times that pass the same pre‑t filter and have no AI response in prior hour.
- Match each user’s real responses (after 60‑minute prior‑AI washout) to their own pseudo moments; measure 30‑minute follow‑up search rate against a within‑user placebo baseline.
- Robustness and sensitivity:
- Multiple draws of pseudo pools, bootstrap over matched sets, sensitivity grid varying washout, landmark threshold, pre‑window, matching granularity, follow‑up window.
- Auxiliary completed‑episode construction (sensitivity) where pseudo‑events inserted into AI‑free episodes reproduce ~73% of real in‑episode excess.
- Theoretical work:
- Causal graph and formalization of endogenous time zero (latent episode state B_it → event A_it and outcome Y_{i,t+h}); proof that single‑surface event‑window data cannot identify treatment effect without additional assumptions.
- Simulations:
- Zero‑effect simulations demonstrate failure modes of common adjustments and quantify limits of proxy conditioning.
Implications for AI Economics
- Event‑window studies that anchor on user‑timed events (e.g., opening an assistant, submitting a prompt) can substantially overstate causal effects on downstream activity (searches, pageviews, purchases) unless they diagnose and adjust for episode selection.
- Within‑user designs and simple pre‑event matching are not sufficient safeguards: the confound is intra‑person and time‑varying.
- Practical recommendations for researchers and policymakers studying AI impacts:
- Treat time zero as endogenous by default. Report diagnostics that test for pre‑event episode structure (e.g., cross‑surface pre‑activity peaks).
- Use engineered negative controls (known‑null/pseudo‑event timestamps) to measure the floor of episode‑selection bias in your data and design.
- Where possible, compare similar task episodes with and without the focal event (episode‑level contrasts), rather than comparing time‑aligned windows around user‑timed events.
- Leverage cross‑surface proxies for episode state and test the key assumption: proxies must respond to episode state but not be affected by the focal event (e.g., “leave‑one‑surface‑out” indices). If that assumption is plausible and validated, panel covariate‑instrument estimators can help.
- Prefer randomized or quasi‑experimental designs (e.g., randomized access, denial, or timing of the assistant) when the policy question requires causal attribution of a single event’s incremental effect.
- Use the paper’s audit protocol and burstcheck as a routine screening step before interpreting event‑anchored volume lifts.
- Scope and limits:
- The critique applies most directly to count/volume lifts (search/pageview rates). Other outcomes (compositional shifts, first‑time events, routing changes) may be less or differently affected but are not automatically safe.
- The paper does not claim all event‑window estimates are invalid; rather, the burden of proof is on studies to show they have separated episode shape from treatment response.
Summary takeaway: in high‑frequency behavioral logs, user‑timed events often occur inside ongoing task episodes. Without diagnostics or stronger identification assumptions, post‑event increases in activity can largely reflect episode continuation rather than causal effects of the event — researchers should use known‑null tests, cross‑surface proxies, episode comparisons, or randomized designs to credibly estimate causal impacts of AI interactions.
Assessment
Claims (9)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| Around conversational-AI responses, browsing and search activity peak before the event rather than jumping at the event; the pre-window activity is approximately 3.2× for browsing and 3.3× for search relative to a within-user placebo. Other | positive | Page views and search activity around AI-response timestamps |
Reading fidelity
high
Study strength
medium
|
3.2× and 3.3× over placebo in the pre-window
|
| Known-null timestamps matched to active AI-response moments produce a large post-event search association: 3.42× relative to a within-user placebo, compared with 4.32× for real AI events. Other | positive | Search activity during the 30-minute post-event follow-up |
Reading fidelity
high
Study strength
high
|
3.42× [3.01, 3.84] for pseudo-events versus 4.32× [3.34, 5.29] for real events
|
| The known-null timestamps reproduce approximately 73% of the excess association observed for real AI events. Other | positive | Excess post-event search activity relative to a within-user placebo |
Reading fidelity
high
Study strength
high
|
0.73 [0.54, 0.92] of the real excess
|
| The fraction of the real-event excess reproduced by known-null timestamps is highest at detectably active moments and declines to approximately zero at quiet moments: 0.56 for landmark-active anchors, 0.43 for sub-threshold anchors, and −0.04 for quiet anchors. Other | mixed | Fraction of real post-event search excess reproduced by known-null timestamps |
Reading fidelity
high
Study strength
high
|
0.56 [0.436, 0.724]; 0.43 [0.316, 0.603]; −0.04 [−0.083, −0.004]
|
| A naive visit-versus-ordinary-time event-window design generates a 5.57× search association even when the event is a null timestamp; adding pre-event state matching does not reduce it and produces associations of 6.11–6.44×. Other | positive | Search activity lift associated with event-anchored moments |
Reading fidelity
high
Study strength
high
|
5.57× [4.81, 6.42]; 6.11–6.44× after pre-event state matching
|
| The event-time confound is within-user and time-varying, so user fixed effects and within-person comparisons do not remove it. Other | positive | Post-event behavioral activity attributed to the focal event |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Across AI, shopping, news, coding, and reference events, focal events occur inside broad multi-surface activity episodes that rise before the event. Other | positive | Pre-event activity across multiple digital surfaces |
Reading fidelity
high
Study strength
medium
|
not reported
|
| The paper's completed-episode sensitivity analysis found that pseudo-events inserted into AI-free episodes reproduced 73% of the excess associated with real in-episode events, with an upper-bound estimate of 108% under whole-episode intensity matching. Other | positive | Excess activity associated with in-episode AI events |
Reading fidelity
high
Study strength
medium
|
73%; upper bound 108%
|
| In the authors' formal setup, a single-surface event-aligned data functional cannot identify the focal event's causal effect without additional assumptions. Other | null_result | Identification of the causal effect of the focal event on the post-event outcome window |
Reading fidelity
high
Study strength
medium
|
not reported
|