The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

X's AI replies displaced community fact-checking: enabling Grok to generate reply-style fact checks reduced both requests for Community Notes and user contributions, disproportionately pushing away the platform's most active volunteer verifiers.

From Notes to Bots: How Generative AI Impacts Human-Led Fact-Checking
Yingxin Zhou, Jingbo Hou · January 01, 2026 · Proceedings of the ... Annual Hawaii International Conference on System Sciences/Proceedings of the Annual Hawaii International Conference on System Sciences
openalex quasi_experimental medium evidence 7/10 relevance Full text usable extracted full text DOI Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

OpenAlex

Latest observation:

  1. Yingxin Zhou provider ID
  2. Jingbo Hou provider ID

Semantic Scholar

Latest observation:

  1. Yingxin Zhou provider ID
  2. Jingbo Hou provider ID
When Grok's reply function became available, user participation in X's Community Notes fell on both the demand and supply sides, with especially large disengagement among highly active contributors crucial to the system's sustainability.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

Human-led approaches, such as crowdsourced fact-checking, have been central to combating misinformation, but generative AI (GenAI) has recently emerged as an alternative, offering faster fact-check-like outputs. We examine how GenAI affects human-led fact-checking in the context of Grok (a GenAI chatbot) and Community Notes (X's crowdsourced fact-checking system). Leveraging the rollout of Grok's reply function, which enabled users to summon AI-generated fact-check-like replies, we find that its availability reduced user participation in Community Notes on both the demand and supply sides. Additionally, this disengagement effect is more pronounced for highly active contributors who are critical to the sustainability of Community Notes. This study enhances the understanding of the interdependence between fact-checking approaches.

Summary

Main Finding

The public rollout of Grok’s inline reply function (March 7, 2025) substantially reduced user participation in X’s crowdsourced fact-checking system (Community Notes) on both demand and supply sides. Daily note requests, new contributor registrations, note ratings, and notes written all fell sharply after the intervention. The substitution effect is stronger for highly active contributors, threatening the sustainability of crowdsourced fact-checking.

Key Points

  • Context: Grok (a GenAI chatbot on X) began producing public, inline "fact-check-like" replies when mentioned (@grok). Users appropriated Grok for verifying tweets (e.g., “Is this tweet true?”).
  • Hypothesis: GenAI fact-checking may substitute for human-led, crowdsourced fact-checking because of GenAI’s speed and scalability, despite higher hallucination risk.
  • Empirical finding: Interrupted Time Series (ITS) analysis shows statistically significant, large declines after Grok’s reply rollout:
    • Daily NoteRequests: −5,619 (p<0.01)
    • NewContributors (daily): −833 (p<0.01)
    • NotesRating (daily): −87,980 (p<0.01)
    • NotesWritten (daily): −694 (p<0.01)
  • Heterogeneity: High-activity contributors reduced participation more than low-activity contributors. This is important because highly active users are critical for maintaining crowdsourced systems.
  • Robustness: Model-free checks (forecasting via PatchTST) and placebo tests support that declines are associated with the Grok rollout rather than spurious trends.
  • Trade-offs emphasized: GenAI brings speed and convenience but risks reducing human verification labor and may propagate hallucinations or misinformation if used as a substitute.

Data & Methods

  • Data:
    • Source: X’s Community Notes program activity logs.
    • Aggregation: Daily time-series over a 6-month window (3 months pre- and 3 months post-intervention), 180 observations.
    • Outcomes:
      • NoteRequests (daily user requests for notes)
      • NewContributors (daily new contributor registrations)
      • NotesRating (daily note helpfulness ratings)
      • NotesWritten (daily notes authored)
  • Identification strategy:
    • Interrupted Time Series (ITS) with robust standard errors to estimate immediate level change and change in trend after March 7, 2025.
    • Model specification includes Post indicator, Time trend, and Post × (Time − T) interaction to capture slope changes.
    • Placebo test using a pseudo-intervention date in the pre-treatment window.
    • Model-free visual checks against PatchTST counterfactual forecasts.
  • Heterogeneity analysis:
    • Contributor-week level panel.
    • Users split into high- and low-activeness by median pre-intervention activity; subgroup ITS analysis to estimate differential impacts.
  • Limitations noted by authors:
    • The study documents participation changes but does not directly measure the accuracy or downstream misinformation consequences of substituting human fact-checks with Grok replies.
    • Grok was not explicitly a fact-checking tool; user appropriation may differ across platforms or AI systems.

Implications for AI Economics

  • Substitution in voluntary public-good provision: Generative AI can crowd out unpaid human contributions to public-good tasks (fact-checking), not just paid labor — implying negative externalities on community-based quality production.
  • Asymmetric supply elasticity: The larger drop among highly active contributors suggests AI can disproportionately affect the most productive contributors; models of platform supply should allow for heterogeneous displacement risk rather than uniform reductions.
  • Quality vs. speed trade-offs and welfare effects: Faster, cheaper AI responses may reduce aggregate human verification capacity, potentially lowering the overall factual accuracy of available signals (if AI outputs are noisier). Welfare analyses must account for externalities from degraded public verification.
  • Policy and platform design levers:
    • Incentives/subsidies: Platforms might need to incentivize human verification (monetary rewards, reputation boosts, visibility guarantees) to counteract AI substitution.
    • Hybrid designs: Pair AI-generated replies with human validation workflows (AI-first triage + human confirmation for high-risk content) to capture scale while preserving quality.
    • Visibility controls & provenance: Label AI replies clearly; limit their primacy for content flagged as high-stakes; surface human Community Notes prominently to preserve complementary roles.
    • Regulation & guardrails: Accuracy guarantees, auditing, and redress mechanisms for AI fact-check outputs could mitigate negative externalities.
  • Research directions for AI economists:
    • Quantify net welfare consequences (speed gains vs. quality losses) from AI substitution in fact-checking and other civic/public-good domains.
    • Structural models of contributor behavior with heterogeneous endowments and intrinsic motivations, to simulate policy interventions (e.g., subsidies, reputation inflation).
    • Cross-platform and long-run studies: persistence of substitution effects, impacts on misinformation prevalence, and generalizability across different GenAI systems and user bases.
    • Market design experiments testing hybrid verification mechanisms, incentive schemes, and provenance signals to measure ability to sustain human participation and quality.

Summary takeaway: The paper provides evidence that readily accessible GenAI fact-checking substitutes for human crowdsourced fact-checking by reducing both requests and contributions—especially from the most active users—posing challenges for the provision and sustainability of high-quality public verification. For AI economics, this highlights a novel channel of AI-induced displacement in voluntary, public-good production and calls for economic analysis and platform policy responses to manage trade-offs between scale and accuracy.

Assessment

Paper Typequasi_experimental Evidence Strengthmedium — The study leverages a plausibly exogenous product rollout to estimate causal effects and likely includes robustness checks and heterogeneous analyses, but it is observational (not randomized) and vulnerable to concurrent platform changes, selection into using Grok, spillovers, and other unobserved confounders that could bias estimates. Methods Rigormedium — Apparent use of standard quasi-experimental tools (event study / DiD, fixed effects, heterogeneity analysis) and focus on both demand and supply increases credibility, but without random assignment or clear instruments, the design cannot fully rule out alternative explanations; the rigor depends on how thoroughly authors test for parallel trends, address selection, and rule out contemporaneous shocks. SampleUser- and post-level activity data from X (formerly Twitter), covering Community Notes contributions and requests as well as Grok reply availability/use, observed over a window spanning the rollout; includes both occasional and highly active Community Notes contributors and the content/timestamps of notes and replies. Themeshuman_ai_collab adoption IdentificationExploits the staggered rollout/availability of Grok's reply function as an exogenous shock to users' ability to summon AI-generated fact-check-like replies and compares Community Notes participation before and after rollout (and/or between exposed and unexposed users), using longitudinal user- and time-fixed effects and event-study/difference-in-differences style comparisons to attribute changes to Grok availability. GeneralizabilitySingle-platform (X) study; platform dynamics and community norms may not match other social networks., Findings pertain to a specific GenAI (Grok) and time period; different models/UXs could produce different effects., Focuses on crowdsourced fact-checking volunteers (non-paid contributors), so results may not generalize to paid moderation or professional fact-checkers., Short- to medium-run rollout effects may differ from long-run equilibrium behavior once users adapt.

Claims (4)

ClaimDirectionOutcomeConfidence & EvidenceDetails
The availability of Grok's reply function reduced user participation in Community Notes on both the demand and supply sides. Adoption Rate negative user participation in Community Notes (demand-side engagement and supply-side contributions)
Reading fidelity high
Study strength medium
not reported
0.48
The disengagement effect following Grok's reply rollout is more pronounced for highly active contributors who are critical to the sustainability of Community Notes. Adoption Rate negative activity/engagement of highly active Community Notes contributors
Reading fidelity high
Study strength medium
not reported
0.48
Generative AI (Grok) emerged as an alternative to human-led fact-checking by offering faster fact-check-like outputs. Task Completion Time positive speed of producing fact-check-like responses (time-to-reply)
Reading fidelity high
Study strength medium
not reported
0.48
Human-led approaches, such as crowdsourced fact-checking, have been central to combating misinformation. Decision Quality positive effectiveness/role of human-led fact-checking in combating misinformation
Reading fidelity high
Study strength low
not reported
0.24

Notes