The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Cheap, on-demand AI boosts help-seeking but can weaken short-term skill acquisition: participants with low-cost access asked for more AI help and performed worse on unassisted post-tests, while learning gains tracked with preserved independent problem-solving rather than request frequency.

How AI Assistance Affects Human Skill Development: A Study of Learning with Logic Puzzles
Shang Wu, Catarina G Belem, Shuyuan Fu, Mark Steyvers, Padhraic Smyth · August 24, 2026
arxiv rct medium evidence 7/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

Arxiv

Latest observation:

  1. Shang Wu unresolved corpus identity
  2. Catarina G Belem unresolved corpus identity
  3. Shuyuan Fu unresolved corpus identity
  4. Mark Steyvers unresolved corpus identity
  5. Padhraic Smyth unresolved corpus identity
In a randomized logic-puzzle experiment, making on-demand AI cheap led to more requests, and participants who used AI showed weaker unaided post-test performance, while greater independent problem-solving during AI access predicted larger gains in latent ability.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

While AI assistance can improve human task performance in the short term, it may also undermine the development of skills in the longer term. We examine this tension in a controlled logic-puzzle experiment involving on-demand AI assistance, where participants complete tasks before, during, and after AI is available. By experimentally varying AI request costs, we find that lower-cost assistance induces more frequent AI use. We also find that participants who request AI assistance during the AI-access phase perform worse at the task after assistance is removed, and their subsequent unassisted performance is overestimated when predicted from earlier AI-assisted performance. We use a Bayesian latent ability model to separate initial ability, post-AI ability, and participant-specific skill change, while estimating how independent reasoning during the AI-access phase relates to skill development. The results show that greater independent problem-solving effort is associated with larger gains in latent ability, consistent with the interpretation that skill development is weaker when AI assistance substitutes for independent reasoning.

Summary

Main Finding

AI assistance that is easy/cheap to access increases short-term reliance and can weaken short-term skill development when it substitutes for independent problem solving. Crucially, it is not mere frequency of AI use but the displacement of independent effort (measured as time spent solving before requesting AI) that predicts weaker gains in latent ability. Also, AI-assisted performance tends to overestimate subsequent unaided performance.

Key Points

  • Experimental manipulation: varying per-request AI costs (no-AI, low-cost 0.1 point, high-cost 0.18 point) led to different help-seeking rates; lower cost → more requests (6.67 vs 3.33 on average).
  • Short-run effects: AI requests improved immediate task success when available (AI was simulated as 100% correct).
  • Post-AI (Phase 3) performance: participants who used AI in Phase 2 performed worse, on average, in the unassisted post-test than participants who did not use AI; Phase-2 AI-assisted performance overestimated later unaided performance for AI users.
  • Mechanism: a “solo share” metric (fraction of time spent working independently before requesting help) was positively associated with gains in latent ability. After controlling for solo share and initial ability, the raw count of AI requests was not predictive of weaker skill gains.
  • Reward-rate improvements (accuracy/time) were large among non-AI users (about 90% increase from Phase 1 to Phase 3), while improvement was smaller in the low-cost AI condition.
  • Secondary observations: low-cost AI participants had longer response times in Phase 3; task feedback and revealed solutions were provided after each problem (to allow learning).

Data & Methods

  • Task: constrained logic puzzles (6 items, 5 constraints) solved under time limits. Participants could make up to two attempts per problem; feedback given (number correct) after first attempt; correct solution revealed after each problem.
  • Design: three-phase within-subjects assessment
    • Phase 1: pre-AI baseline (no AI).
    • Phase 2: 20-minute treatment phase (no-AI control or AI available on demand with cost).
    • Phase 3: post-AI assessment (no AI).
  • Conditions: No-AI; Low-cost AI (0.1-point deduction per AI reveal); High-cost AI (0.18-point deduction). AI assistance revealed the correct location of one randomly selected object per request; AI simulated as perfectly correct.
  • Sample: recruited 150 US adults via Prolific; final N = 124 after exclusions (42 No-AI, 43 Low-cost, 39 High-cost). Sample skewed toward college-educated participants (66 BA, 58 MA+), mean age ~40.
  • Outcome measures:
    • Accuracy (0–6 correct objects per problem),
    • Response time (seconds),
    • Reward rate = average accuracy / average response time (correct objects per minute),
    • AI usage = number of requests in Phase 2,
    • Solo share = fraction of problem time spent working independently before requesting AI (averaged over Phase 2 problems).
  • Analysis:
    • Descriptive statistics and hypothesis tests (ANOVA, t-tests) to evaluate usage and performance differences across conditions.
    • Regression showing Phase-2 performance tends to overpredict Phase-3 performance for AI users.
    • Bayesian latent ability model: treats Phase 1 and Phase 3 performance (accuracy, response time) as noisy signals of latent ability, models participant-specific skill change as a function of initial ability and behavioral measures (notably solo share). This model separates initial ability, post-AI ability, and skill-change components to identify relationships between independent effort and learning gains.

Implications for AI Economics

  • Human capital accumulation vs. short-term productivity trade-off:
    • Cheap/easy access to highly accurate AI increases immediate productivity but can reduce acquisition of durable skills when it substitutes for independent reasoning. Economists and organizations must weigh short-term gains against potential long-term depreciation of human capital.
  • Pricing and subsidy design for AI tools matters:
    • Small per-use costs (or free access) change usage behavior; policy levers (pricing, quotas, or earned-access schemes) can be used to moderate reliance and encourage independent practice where skill-building is desired.
  • Measurement and incentive distortions:
    • Performance metrics based on AI-assisted outputs can overstate true human ability, producing biased signals for hiring, promotion, certification, or payments tied to task performance. Incentive systems should distinguish assisted vs. unaided performance or adjust evaluations accordingly.
  • Complementarity vs. substitution design:
    • The economic value of AI depends on whether it complements human reasoning (scaffolding that preserves learning) or substitutes it. Product and platform design (e.g., forcing periods of independent work, requiring explanation, partial hints instead of full answers, delayed reveals) can nudge toward complementarity and preserve human skill accumulation.
  • Policy and training programs:
    • In education and workforce training, regulators and program designers should consider guidelines ensuring AI tools augment learning (e.g., assist-as-feedback, graded hints) rather than acting as shortcuts that reduce learning outcomes.
  • Empirical evaluation frameworks:
    • The paper’s Bayesian latent-ability approach illustrates how to infer hidden skill dynamics from observed assisted/unassisted performance and behavior; such methods can be used to evaluate interventions (pricing, interface changes, pedagogical scaffolds) in labor and education economics.

Caveats and directions for further work - External validity: controlled logic-puzzle task and a lab-like short-term study may not generalize to complex, real-world tasks or long-run skill dynamics. - Simulated perfect AI: using 100%-accurate AI isolates cost effects but does not capture trade-offs when AI is imperfect. - Short horizon: the study measures immediate post-treatment performance; longer-run follow-ups are needed to assess persistent effects on human capital. - Sample limitations: online Prolific sample, educated adults; different populations (students, professionals) may respond differently.

Overall, the paper highlights that AI availability and its pricing shape reliance behavior and that the critical mechanism undermining short-term learning is displacement of independent effort rather than raw frequency of requests. For economic policy, product design, and training programs, interventions that preserve or incentivize independent problem solving while leveraging AI for targeted scaffolding will better balance short-term productivity with durable skill formation.

Assessment

Paper Typerct Evidence Strengthmedium — Randomized assignment to AI-cost and inclusion of a no-AI control give strong internal validity for causal claims about how cost affects help-seeking; evidence that AI exposure reduces subsequent unaided performance is consistent but relies partly on non-randomized AI use and modestly powered comparisons (many effects reported at p<0.1), so causal claims about AI use -> learning are weaker. Methods Rigormedium — Well-designed pre-post RCT with attention/comprehension checks, incentive alignment, and principled latent-trait Bayesian modeling; strengths include randomized cost, control group, and clear behavioral measures (solo share, requests, reward rate). Limitations: AI-use during Phase 2 is endogenous (not randomized), some reported effects are marginal, the AI was simulated as perfectly accurate (limits ecological validity), and the sample is an online convenience cohort. Sample124 U.S.-based adult participants recruited on Prolific (final N=124 after exclusions; 42 No-AI, 43 Low-cost AI, 39 High-cost AI), mean age 40 (SD=12), educationally skewed (all at least undergraduate degree; 58 with master's+), paid base $9 plus performance bonuses; task: timed logic-puzzle problems across three phases with on-demand simulated AI assistance in Phase 2. Themesskills_training human_ai_collab IdentificationBetween-subjects randomized assignment to three Phase-2 conditions (No-AI, Low-cost AI, High-cost AI) with pre- (Phase 1) and post- (Phase 3) unassisted assessments; simulated AI with fixed (100%) accuracy to isolate effect of request cost on help-seeking; Bayesian latent-ability model uses Phase-1 and Phase-3 performance as signals of latent skill and conditions/behavior in Phase-2 to predict skill change. Note: AI request decisions during Phase 2 were not themselves randomized, so inference about causal effect of individual AI requests on learning is associational. GeneralizabilityTask-specific: controlled logic-puzzle paradigm under time pressure — may not generalize to complex, open-ended, or domain-specific tasks (e.g., programming, writing, professional work)., Simulated AI was perfectly accurate (100%), unlike real-world assistants that vary in quality, so findings may not translate when AI is noisy or provides partial/incorrect help., Short-term study: measured immediate/short-term learning within a single session; long-term retention and transfer not assessed., Convenience online sample (Prolific) of educated U.S. adults limits population generalizability (age, education, culture, professional background)., Monetary incentives and the artificial UI/time constraints may alter effort and strategy compared with real-world settings.

Claims (10)

ClaimDirectionOutcomeConfidence & EvidenceDetails
Among participants who made no AI requests during Phase 2, mean reward rate increased from 2.03 to 3.86 correct objects per minute between Phase 1 and Phase 3, a 90.2% increase. Skill Acquisition positive Change in unassisted logic-puzzle reward rate across study phases
Reading fidelity high
Study strength high
n=75
90.2% increase in correct objects per minute
1.0
Lower-cost AI assistance led participants to make more AI requests during Phase 2 than higher-cost AI assistance. Adoption Rate positive Number of AI assistance requests during Phase 2
Reading fidelity high
Study strength low
n=82
6.67 vs. 3.33 AI requests
0.3
The increase in reward rate from Phase 1 to Phase 3 was smaller for the Low-cost AI condition than for the No-AI condition. Skill Acquisition negative Change in reward rate from pre-AI to post-AI assessment
Reading fidelity high
Study strength low
n=85
1.32 versus 1.95 correct objects per minute
0.3
Participants who requested AI assistance during Phase 2 had lower average reward rates in the unassisted Phase 3 than participants who did not request AI. Skill Acquisition negative Post-assistance unassisted reward rate
Reading fidelity high
Study strength medium
n=124
3.42 vs. 3.86 correct items per minute
0.6
Predictions based on Phase 2 AI-assisted performance overestimated AI users' subsequent unassisted Phase 3 performance by 0.22 reward-rate units on average. Skill Acquisition negative Prediction error for subsequent unassisted reward rate
Reading fidelity high
Study strength medium
0.22 reward-rate units overestimated
0.6
Predictions based on Phase 2 unassisted performance underestimated non-users' subsequent Phase 3 performance by 0.15 reward-rate units on average. Skill Acquisition positive Prediction error for subsequent unassisted reward rate
Reading fidelity high
Study strength medium
0.15 reward-rate units underestimated
0.6
Greater independent problem-solving effort during the AI-access phase was associated with larger gains in latent ability. Skill Acquisition positive Gain in latent problem-solving ability from before to after AI access
Reading fidelity high
Study strength medium
n=124
0.6
After accounting for initial ability and independent effort, the frequency of AI assistance requests was not associated with gains in skill development. Skill Acquisition null_result Latent skill-development gains
Reading fidelity high
Study strength medium
n=124
0.6
Phase 3 accuracy did not differ across the No-AI, Low-cost AI, and High-cost AI conditions. Skill Acquisition null_result Accuracy in the post-AI unassisted assessment
Reading fidelity high
Study strength medium
n=124
0.6
Phase 3 response times were longer in the Low-cost AI condition than in either the High-cost AI or No-AI condition. Task Completion Time negative Response time during the post-AI unassisted assessment
Reading fidelity high
Study strength medium
n=124
111.81s vs. 86.78s and 92.21s
0.6

Notes