The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Event-window studies that align on user-initiated AI prompts can misattribute ongoing task activity to the AI: in same-user cross-surface logs, engineered zero-effect timestamps reproduce the majority of the apparent post-event search and browsing lift, showing that single-surface event windows lack identifying power without additional assumptions or cross-surface controls.

Event-Time Confounding Under Bursty Human Dynamics
Michael Iannelli, Alan Ai · August 21, 2026
arxiv other high evidence 8/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

Arxiv

Latest observation:

  1. Michael Iannelli unresolved corpus identity
  2. Alan Ai unresolved corpus identity

Semantic Scholar

Latest observation:

  1. Michael Iannelli provider ID
  2. Alan Ai provider ID
Anchoring analyses on user-timed AI events often mistakes continuation of an ongoing task episode for an event effect: known-null pseudo-timestamps that should cause nothing reproduce most of the observed post-event activity lift, particularly at 'active' anchors.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

Studies of digital behavior often align users at moments they choose, such as opening an AI assistant, clicking a recommendation, or visiting a product page, and interpret higher activity afterward as an event effect. We show how this creates an endogenous time zero: the event occurs during an ongoing task episode, so the aligned curve can trace episode continuation rather than a response to the event. In same-user, cross-surface web logs, AI, shopping, news, coding, and reference events are all preceded by broad activity increases that peak before time zero. Our strongest test uses known-null timestamps that cause nothing. Among the 5.8% of AI responses meeting strict pre-event activity and washout criteria, these timestamps show 3.42 times the post-event search activity of a within-user placebo, compared with 4.32 times for real events. The fraction of excess reproduced by the known null falls from 0.56 at detectably active moments to -0.04 at quiet moments, where the design detects none. We formalize this episode-selection bias, prove that a single-surface event window cannot separate it from a genuine effect without additional assumptions, and show in zero-effect simulations why user fixed effects and coarse activity matching can fail: the confound is within-user and time-varying. We provide a diagnostic protocol, public-data benchmarks, and burstcheck, a lightweight audit tool. User-timed events may have real effects, but post-event volume does not identify them by default; studies should compare similar episodes with and without the event.

Summary

Main Finding

When studies align users on moments they choose (e.g., opening an AI assistant, clicking a recommendation) and interpret higher activity afterward as a causal “event effect,” much of the measured lift can instead be continuation of an already‑ongoing task episode. The paper formalizes this “endogenous time zero” / episode‑selection bias, shows it survives common within‑user fixes, demonstrates it on real cross‑surface web logs using a known‑null (pseudo‑event) test, proves a single‑surface event‑window cannot identify the effect without extra assumptions, and supplies diagnostics and a lightweight audit tool (burstcheck).

Key Points

  • Problem statement
    • Endogenous time zero: the aligned timestamp (time zero) is chosen by the same latent process (task/intent/episode intensity) that drives outcomes, creating a backdoor path from latent episode state to post‑event outcomes.
    • Episode‑selection bias: user‑timed focal events are more likely to occur during latent task episodes; the episode’s own continuation can be mistaken for an event effect.
  • Why common defenses fail
    • The confounder is within‑user and time‑varying, so user fixed effects and within‑user comparisons do not remove it.
    • Coarse matching (e.g., on recent activity, hour‑of‑day) or anchoring to page views can even exacerbate the bias by oversampling dense spells (inspection paradox / regression to the mean).
  • Formal result
    • Proposition (informal): No functional of single‑surface event‑aligned data identifies the event’s causal effect without additional assumptions (e.g., a proxy that responds to episode state but is unaffected by the focal event).
  • Empirical demonstration (known‑null pseudo‑event experiment)
    • Data: opt‑in cross‑surface panel recording same users’ browsing, search, and conversational‑AI events.
    • Strict pre‑event construction: an “active landmark” is a moment with ≥5 non‑search, non‑AI page views in the prior 30 minutes (only pre‑t information used).
    • Headline result (per‑user matched set, 60‑minute prior‑AI washout): real AI events show a post‑event search lift of 4.32× (95% CI [3.34, 5.29]) versus a within‑user placebo; known‑null timestamps (guaranteed to cause nothing) matched on pre‑event activity show a lift of 3.42× ([3.01, 3.84]). The known‑null reproduces ~0.73 of the real excess.
    • Anchor‑class gradient: reproduced fraction falls with pre‑event quietness — landmark‑active anchors: 0.56 (±), mildly active: 0.43, quiet: −0.04. (31% of responses are landmark‑active.)
    • Many standard adjustments (hour‑of‑day matching, richer matching on pre‑activity) do not eliminate the false association.
  • Simulations and sensitivity
    • Zero‑effect simulations explain why unit fixed effects and coarse activity matching fail: the confound is within‑user and time‑varying.
    • Conditioning on cross‑surface proxies reduces confounding but does not automatically identify the effect unless they satisfy additional conditions (e.g., respond to episode state but are unaffected by the treatment).
  • Tools and protocols
    • Provides a diagnostic protocol, public plasmode benchmarks (real data with injected effects), and burstcheck, a light audit tool whose screening statistic is validated on an external event set.

Data & Methods

  • Data: same‑user, cross‑surface web logs (browsing, search, conversational‑AI events) from an opt‑in panel; analyses use per‑user capping in some constructions and an equal‑user matched set in the headline.
  • Primary known‑null construction (strictly pre‑t):
    • Define active landmarks using only pre‑t activity (≥5 non‑search/non‑AI page views in prior 30 minutes).
    • Real moments: AI responses that are landmarks. Pseudo moments: uniform‑random times that pass the same pre‑t filter and have no AI response in prior hour.
    • Match each user’s real responses (after 60‑minute prior‑AI washout) to their own pseudo moments; measure 30‑minute follow‑up search rate against a within‑user placebo baseline.
  • Robustness and sensitivity:
    • Multiple draws of pseudo pools, bootstrap over matched sets, sensitivity grid varying washout, landmark threshold, pre‑window, matching granularity, follow‑up window.
    • Auxiliary completed‑episode construction (sensitivity) where pseudo‑events inserted into AI‑free episodes reproduce ~73% of real in‑episode excess.
  • Theoretical work:
    • Causal graph and formalization of endogenous time zero (latent episode state B_it → event A_it and outcome Y_{i,t+h}); proof that single‑surface event‑window data cannot identify treatment effect without additional assumptions.
  • Simulations:
    • Zero‑effect simulations demonstrate failure modes of common adjustments and quantify limits of proxy conditioning.

Implications for AI Economics

  • Event‑window studies that anchor on user‑timed events (e.g., opening an assistant, submitting a prompt) can substantially overstate causal effects on downstream activity (searches, pageviews, purchases) unless they diagnose and adjust for episode selection.
  • Within‑user designs and simple pre‑event matching are not sufficient safeguards: the confound is intra‑person and time‑varying.
  • Practical recommendations for researchers and policymakers studying AI impacts:
    • Treat time zero as endogenous by default. Report diagnostics that test for pre‑event episode structure (e.g., cross‑surface pre‑activity peaks).
    • Use engineered negative controls (known‑null/pseudo‑event timestamps) to measure the floor of episode‑selection bias in your data and design.
    • Where possible, compare similar task episodes with and without the focal event (episode‑level contrasts), rather than comparing time‑aligned windows around user‑timed events.
    • Leverage cross‑surface proxies for episode state and test the key assumption: proxies must respond to episode state but not be affected by the focal event (e.g., “leave‑one‑surface‑out” indices). If that assumption is plausible and validated, panel covariate‑instrument estimators can help.
    • Prefer randomized or quasi‑experimental designs (e.g., randomized access, denial, or timing of the assistant) when the policy question requires causal attribution of a single event’s incremental effect.
    • Use the paper’s audit protocol and burstcheck as a routine screening step before interpreting event‑anchored volume lifts.
  • Scope and limits:
    • The critique applies most directly to count/volume lifts (search/pageview rates). Other outcomes (compositional shifts, first‑time events, routing changes) may be less or differently affected but are not automatically safe.
    • The paper does not claim all event‑window estimates are invalid; rather, the burden of proof is on studies to show they have separated episode shape from treatment response.

Summary takeaway: in high‑frequency behavioral logs, user‑timed events often occur inside ongoing task episodes. Without diagnostics or stronger identification assumptions, post‑event increases in activity can largely reflect episode continuation rather than causal effects of the event — researchers should use known‑null tests, cross‑surface proxies, episode comparisons, or randomized designs to credibly estimate causal impacts of AI interactions.

Assessment

Paper Typeother Evidence Strengthhigh — The authors demonstrate the bias on real same-user cross-surface panel data using a constructed known-null (guaranteed-zero) timestamp experiment that reproduces a large fraction of the apparent AI-event lift; they run extensive pre-specified sensitivity checks, bootstrap inference, class-stratified analyses, simulations, and present formal propositions showing non-identifiability absent extra assumptions. Methods Rigorhigh — Carefully pre-specified strict pre-event constructions, washouts, within-user matching, multiple matching/clock controls, bootstrapped uncertainty that rebuilds matched sets each replicate, influence diagnostics, complementary simulations, and formal theoretical results together form a rigorous methods package; limitations arise from reliance on an opt-in panel and limited public disclosure of raw counts. SampleAn opt-in cross-surface research panel recording the same users' web browsing, search, and conversational-AI events over a measurement window; analyses use per-user caps (except for headline matched set where both arms are equally sampled), landmark-active filters (e.g., >=5 non-search/non-AI page views in prior 30 minutes), pseudo moments drawn from same users with no AI in the prior hour, and follow-ups typically of 30–60 minutes; exact panel size and raw counts are not disclosed in the provided excerpt. Themeshuman_ai_collab productivity IdentificationThe paper does not claim to identify a causal effect; rather it uses a known-null/negative-control pseudo-event design and within-user matched comparisons (strictly pre-event matching, washouts, per-user baselines, and reweighted pooled controls) to test whether episode-selection (endogenous time-zero) can generate apparent event effects. It supplements empirical diagnostics with formal proofs, zero-effect simulations, and sensitivity grids; it argues that single-surface event-window contrasts are not identified without extra assumptions (e.g., a valid cross-surface proxy or exclusion). GeneralizabilityOpt-in, proprietary panel — may not represent the general population or all user types., Observations are limited to the surfaces the panel records (web browsing, search, conversational-AI) and may miss in-app or off-platform activity, biasing the observable proxy for latent episode state., Findings focus on count-based volume outcomes over minute-to-hour windows and may not generalize to compositional outcomes, long-run effects, or first-observed events., Results are most directly applicable to user-timed events that fall inside bursty task episodes; scheduled or instrumented events may behave differently., Some diagnostics rely on matching on observables; unmeasured dimensions of intent/task type may still differ between real and pseudo moments.

Claims (9)

ClaimDirectionOutcomeConfidence & EvidenceDetails
Around conversational-AI responses, browsing and search activity peak before the event rather than jumping at the event; the pre-window activity is approximately 3.2× for browsing and 3.3× for search relative to a within-user placebo. Other positive Page views and search activity around AI-response timestamps
Reading fidelity high
Study strength medium
3.2× and 3.3× over placebo in the pre-window
0.12
Known-null timestamps matched to active AI-response moments produce a large post-event search association: 3.42× relative to a within-user placebo, compared with 4.32× for real AI events. Other positive Search activity during the 30-minute post-event follow-up
Reading fidelity high
Study strength high
3.42× [3.01, 3.84] for pseudo-events versus 4.32× [3.34, 5.29] for real events
0.2
The known-null timestamps reproduce approximately 73% of the excess association observed for real AI events. Other positive Excess post-event search activity relative to a within-user placebo
Reading fidelity high
Study strength high
0.73 [0.54, 0.92] of the real excess
0.2
The fraction of the real-event excess reproduced by known-null timestamps is highest at detectably active moments and declines to approximately zero at quiet moments: 0.56 for landmark-active anchors, 0.43 for sub-threshold anchors, and −0.04 for quiet anchors. Other mixed Fraction of real post-event search excess reproduced by known-null timestamps
Reading fidelity high
Study strength high
0.56 [0.436, 0.724]; 0.43 [0.316, 0.603]; −0.04 [−0.083, −0.004]
0.2
A naive visit-versus-ordinary-time event-window design generates a 5.57× search association even when the event is a null timestamp; adding pre-event state matching does not reduce it and produces associations of 6.11–6.44×. Other positive Search activity lift associated with event-anchored moments
Reading fidelity high
Study strength high
5.57× [4.81, 6.42]; 6.11–6.44× after pre-event state matching
0.2
The event-time confound is within-user and time-varying, so user fixed effects and within-person comparisons do not remove it. Other positive Post-event behavioral activity attributed to the focal event
Reading fidelity high
Study strength medium
not reported
0.12
Across AI, shopping, news, coding, and reference events, focal events occur inside broad multi-surface activity episodes that rise before the event. Other positive Pre-event activity across multiple digital surfaces
Reading fidelity high
Study strength medium
not reported
0.12
The paper's completed-episode sensitivity analysis found that pseudo-events inserted into AI-free episodes reproduced 73% of the excess associated with real in-episode events, with an upper-bound estimate of 108% under whole-episode intensity matching. Other positive Excess activity associated with in-episode AI events
Reading fidelity high
Study strength medium
73%; upper bound 108%
0.12
In the authors' formal setup, a single-surface event-aligned data functional cannot identify the focal event's causal effect without additional assumptions. Other null_result Identification of the causal effect of the focal event on the post-event outcome window
Reading fidelity high
Study strength medium
not reported
0.12

Notes