The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Prompting users to spot assumptions in AI plans curbs overreliance without extra cognitive burden; hypothetical 'what-if' prompts feel more useful to users but are less effective at reducing blind trust.

An Experimental Comparison of Cognitive Forcing Functions for Execution Plans in AI-Assisted Writing: Effects On Trust, Overreliance, and Perceived Critical Thinking
Ahana Ghosh, Advait Sarkar, Siân Lindley, Christian Poelitz · January 25, 2026
arxiv rct medium evidence 8/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

Arxiv

Latest observation:

  1. Ahana Ghosh unresolved corpus identity
  2. Advait Sarkar unresolved corpus identity
  3. Siân Lindley unresolved corpus identity
  4. Christian Poelitz unresolved corpus identity

Semantic Scholar

Latest observation:

  1. Ahana Ghosh provider ID
  2. Advait Sarkar provider ID
  3. Siân Lindley provider ID
  4. Christian Poelitz provider ID
In a randomized experiment, requiring users to identify assumptions in AI-generated plans (Assumption CFF) reduced overreliance without raising cognitive load, while WhatIf prompts were perceived as most helpful by participants.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

Generative AI (GenAI) tools improve productivity in knowledge workflows such as writing, but also risk overreliance and reduced critical thinking. Cognitive forcing functions (CFFs) mitigate these risks by requiring active engagement with AI output. As GenAI workflows grow more complex, systems increasingly present execution plans for user review. However, these plans are themselves AI-generated and prone to overreliance, and the effectiveness of applying CFFs to AI plans remains underexplored. We conduct a controlled experiment in which participants completed AI-assisted writing tasks while reviewing AI-generated plans under four CFF conditions: Assumption (argument analysis), WhatIf (hypothesis testing), Both, and a no-CFF control. A follow-up think-aloud and interview study qualitatively compared these conditions. Results show that the Assumption CFF most effectively reduced overreliance without increasing cognitive load, while participants perceived the WhatIf CFF as most helpful. These findings highlight the value of plan-focused CFFs for supporting critical reflection in GenAI-assisted knowledge work.

Summary

Main Finding

Assumption-focused cognitive forcing functions (CFFs) applied to AI-generated execution plans for writing tasks most effectively reduced users’ overreliance on AI outputs without increasing cognitive load. Targeting argument-analysis (identifying the AI’s implicit assumptions) outperformed targeting hypothesis-testing (WhatIf) or combining both. Users nevertheless perceived the WhatIf prompt as more helpful, revealing a divergence between subjective preference and objective effectiveness. Participant traits (cognitive disposition, GenAI familiarity) influenced whether users revised their judgments, but CFF design primarily drove overreliance behavior.

Key Points

  • Research question: Which plan-focused CFFs help people critically evaluate AI plans and prevent overreliance in AI-assisted writing?
  • Interventions tested (between-subjects):
    • Assumptions: prompts to identify implicit assumptions in the AI plan (targets argument analysis).
    • WhatIf: prompts to reason about the impact of changing a key plan step (targets hypothesis testing).
    • Both: Assumptions then WhatIf.
    • None: control (no CFF).
  • Studies:
    • Controlled experiment: n = 214 participants completing AI-assisted writing tasks with plan review.
    • Qualitative think-aloud + interviews: n = 12 participants comparing conditions.
  • Primary outcomes:
    • Revision of initial readiness assessment of AI output (did users change their readiness ratings after CFF engagement).
    • Behavioral overreliance (tendency to accept AI output without critical revision).
    • Cognitive load and perceived helpfulness/trust of CFFs.
  • Main behavioral result: Assumptions reduced overreliance most effectively while not increasing measured cognitive load. Combining CFFs (Both) was not additive and could be counterproductive.
  • Perception gap: users tended to rate WhatIf as more helpful, despite poorer objective performance relative to Assumptions.
  • Individual heterogeneity: cognitive disposition and GenAI familiarity shaped revision behavior, suggesting CFF effects interact with user traits.

Data & Methods

  • Experimental design:
    • Realistic AI-assisted writing scenarios where the system generated an execution plan and a draft.
    • Participants reviewed the AI plan, completed the assigned CFF (or none), then assessed readiness of AI output and could request changes.
    • Plan-centered CFFs were implemented as short interactive prompts/quizzes tied to the AI plan (e.g., select assumptions, identify the most critical step and its impact).
  • Sample:
    • Large-scale controlled experiment: n = 214 (between-subjects across four conditions).
    • Qualitative follow-up: n = 12 think-aloud interviews covering all conditions for comparative subjective feedback.
  • Measures:
    • Behavioral measures of overreliance (acceptance vs revision of draft/plan).
    • Changes in readiness judgments (did CFF prompt reassessment).
    • Cognitive load (measured comparatively across conditions).
    • Subjective ratings of helpfulness and trust.
    • Recorded participant traits (cognitive disposition, GenAI familiarity) to assess moderation effects.
  • Analysis:
    • Between-group comparisons to identify which CFFs changed behavior and cognitive load.
    • Qualitative coding of think-aloud data to understand perceived usefulness and friction.

Implications for AI Economics

  • Productivity vs quality trade-off at the interface level:
    • Plan-focused CFFs can reduce costly errors from overreliance without imposing measurable extra cognitive load (Assumptions CFF), implying a way to preserve gains in productivity while improving output quality.
    • However, perceived helpfulness (WhatIf) does not equal effectiveness; product teams that optimize for user satisfaction alone may underdeliver on error reduction and quality control.
  • Labor quality, deskilling, and human capital:
    • Effective CFFs that preserve critical evaluation (argument analysis) can slow or prevent deskilling that accompanies passive use of generative tools, supporting longer-term maintenance of worker skills and human capital.
    • Firms that deploy such CFFs may avoid downstream costs tied to degraded judgment (rework, reputational risk, regulatory fines).
  • Adoption dynamics and product design incentives:
    • Designers face a trade-off: some CFFs reduce overreliance but may be perceived as less useful. Adoption and continued use depend on both objective benefits and subjective user experience. A/B testing and education may be needed to align perception with effectiveness.
    • The ineffectiveness (or harm) of combining multiple CFFs cautions against stacking friction indiscriminately; incremental, well-targeted interventions are preferable.
  • Heterogeneity and targeted deployment:
    • Effects vary by worker traits (cognitive disposition, GenAI familiarity). Economic evaluations of GenAI deployment should account for heterogeneous returns—CFFs may be more valuable for less experienced users or in contexts where critical thinking has higher payoff.
    • Customizable or adaptive CFFs (e.g., enabled based on user profile, task criticality) could maximize net benefits.
  • Organizational policy and regulation:
    • CFFs are a light-touch, design-based safeguard that firms could adopt to improve trust calibration and reduce automation bias. Regulators or industry standards for high-risk domains might recommend plan-focused CFFs (especially argument-analysis prompts) as part of human-in-the-loop UI requirements.
  • Measurement and valuation in AI economics:
    • Economic impact assessments of GenAI should include behavioral-intervention effects (e.g., CFFs) because interface design materially alters realized productivity and error rates.
    • Cost–benefit analyses should quantify both time costs (perceived friction) and error-avoidance gains; this paper suggests some CFFs can yield net quality gains without measurable time/cognitive penalties.
  • Product-market signals and mismatch risk:
    • The divergence between perceived helpfulness and objective effectiveness can distort market signals—user ratings may favor interfaces that feel helpful but deliver worse calibration. Firms and policymakers should rely on outcome-based metrics (error rates, revision frequency) when evaluating tool safety and effectiveness.

Limitations & open questions relevant to economics - Domain and task scope: experiment focused on writing tasks; generalization to other knowledge work (e.g., coding, medical synthesis, legal drafting) requires validation. - Short-term vs long-term effects: experiment measured immediate behavior; long-term impacts on skill retention, training needs, and labor productivity are not yet known. - Field deployment costs and behavioral adaptation: real-world adoption dynamics, compliance, and possible gaming or superficial responses to prompts need study. - Optimal targeting and personalization: how to cost-effectively adapt CFFs to user types and task criticality remains an open design/economic optimization problem.

Practical takeaways for economists and practitioners - When evaluating or deploying GenAI in organizations, measure both subjective satisfaction and objective calibration/quality; prefer outcome-based metrics. - Consider integrating light, plan-focused CFFs that target argument analysis to reduce overreliance without detectable extra cognitive load. - Account for heterogeneity in workforce skill and tool familiarity; allocate training and UI personalization resources where they yield the largest return in quality. - Use controlled A/B tests that track error rates and revision behavior (not just engagement) to decide which CFF designs to scale.

Assessment

Paper Typerct Evidence Strengthmedium — Random assignment provides credible internal validity for causal claims about the effects of CFFs on overreliance and cognitive load, and the qualitative follow-up strengthens interpretation; however, external validity is limited by laboratory/short-term tasks, likely modest sample sizes, potential demand effects, single task domain (writing), and possible reliance on subjective measures. Methods Rigormedium — The study uses a rigorous experimental design with multiple treatment arms and complementary qualitative methods, but likely suffers from typical lab-experiment limitations (unclear sample size and representativeness, ecological validity of the AI tool and tasks, and possible measurement reliance on self-reports rather than long-run performance metrics). SampleHuman participants completed AI-assisted writing tasks in a controlled experiment (four between-subject CFF conditions); a subset participated in think-aloud protocols and interviews—exact N, recruitment source, and demographics are not specified in the abstract. Themeshuman_ai_collab productivity IdentificationRandomized controlled experiment: participants were randomly assigned to one of four cognitive forcing function (CFF) conditions (Assumption, WhatIf, Both, control); causal effects are estimated by comparing behavioral and self-report outcomes across these randomized groups, supplemented by qualitative think-aloud and interview data. GeneralizabilityLab/short-term experimental tasks may not reflect real-world, long-term knowledge work, Single task domain (writing) limits transfer to other knowledge workflows, Likely single GenAI system and specific plan presentation format used — results may not hold for other tools/interfaces, Unclear participant representativeness (e.g., students or crowdworkers) limits population generalizability, Cultural and language context not specified; may not generalize across languages or work cultures

Claims (10)

ClaimDirectionOutcomeConfidence & EvidenceDetails
Generative AI (GenAI) tools improve productivity in knowledge workflows such as writing. Developer Productivity positive productivity in knowledge workflows (e.g., writing)
Reading fidelity high
Study strength medium
not reported
0.6
GenAI also risks overreliance and reduced critical thinking. Decision Quality negative overreliance on AI outputs / reduced critical thinking
Reading fidelity high
Study strength medium
not reported
0.6
Cognitive forcing functions (CFFs) mitigate these risks by requiring active engagement with AI output. Decision Quality positive reduction in overreliance / increased critical engagement
Reading fidelity high
Study strength medium
not reported
0.6
As GenAI workflows grow more complex, systems increasingly present execution plans for user review; these plans are themselves AI-generated and prone to overreliance. Task Allocation negative presentation and user reliance on AI-generated execution plans
Reading fidelity high
Study strength medium
not reported
0.6
The effectiveness of applying CFFs to AI-generated plans remains underexplored. Other null_result empirical evidence on CFFs applied to AI plans
Reading fidelity high
Study strength speculative
not reported
0.1
We conducted a controlled experiment in which participants completed AI-assisted writing tasks while reviewing AI-generated plans under four CFF conditions: Assumption (argument analysis), WhatIf (hypothesis testing), Both, and a no-CFF control. Other null_result experimental manipulation / CFF condition exposure
Reading fidelity high
Study strength high
not reported
1.0
A follow-up think-aloud and interview study qualitatively compared these conditions. Other null_result qualitative comparison of participant experiences across CFF conditions
Reading fidelity high
Study strength high
not reported
1.0
Results show that the Assumption CFF most effectively reduced overreliance without increasing cognitive load. Decision Quality positive reduction in overreliance; no increase in cognitive load
Reading fidelity high
Study strength medium
not reported
0.6
Participants perceived the WhatIf CFF as most helpful. Worker Satisfaction positive participant perceived helpfulness of CFF conditions (WhatIf rated most helpful)
Reading fidelity high
Study strength medium
not reported
0.6
These findings highlight the value of plan-focused CFFs for supporting critical reflection in GenAI-assisted knowledge work. Decision Quality positive support for critical reflection when reviewing AI-generated plans
Reading fidelity high
Study strength medium
not reported
0.6

Notes