0 cumulative citations
View corpus contextSimple behavioral fixes — asking employees to generate brief explanations and reminding them with a personal cue — make human oversight of LLM output measurably more effective: self-explanations raised error detection by about ten percentage points, and daily personalized cues slowed the erosion of vigilance across a working week.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
Large language models (LLMs) are increasingly embedded in organizational work, yet their errors often pass human review. Prior research locates such failures in users' capability to review LLM output or their engagement in doing so. We develop an alternative, retrieval-based account of human oversight and posit that error detection is more effective when oversight-relevant information is accessible to users at the moment of review. Across two randomized lab-in-the-field experiments with 640 customer-facing employees, we show that self-generated explanations improve error detection and strengthen recall of verification-relevant reasoning, while cues that reactivate such reasoning help sustain detection under repeated LLM use. Theoretically, we identify information retrievability as a distinct precondition for effective oversight and specify generative encoding and cue-supported reactivation as mechanisms that build and sustain it. Practically, lightweight onboarding self-explanations and daily retrieval cues can make human oversight more resilient as LLM use becomes routine.
Summary
Main Finding
Retrievability — whether oversight-relevant information encoded earlier is accessible at the moment of review — is a distinct precondition for effective human oversight of LLM output. Two randomized lab-in-the-field experiments (N = 640 customer-facing employees) show that (1) generative encoding via self-explanation improves later error detection (≈10 percentage points) and strengthens recall of verification-relevant reasoning, and (2) brief personalized retrieval cues presented during repeated LLM use slow the decline in error detection by reactivating that prior reasoning. Together these results identify generative encoding and cue-supported reactivation as low-cost, effective mechanisms to build and sustain human oversight.
Key Points
- Problem addressed: Users often miss LLM errors even when they know the system can err and are motivated to review. The paper argues this can occur because the oversight-relevant information is not retrieved at the moment of review.
- Concept introduced: Retrievability — distinct from capability (having the information) and engagement (willingness to check) — denotes the potential for encoded oversight information to be accessed when needed.
- Theoretical basis: Encoding Specificity principle — retrieval depends on match between cues at encoding and cues at retrieval. Two actionable pathways:
- Generative encoding: actively generating an explanation (self-explanation) produces more elaborated, organized verification-relevant reasoning that is more retrievable later.
- Cue-supported reactivation: brief cues associated with the prior encoding episode can reinstate that reasoning without restating the full explanation.
- Empirical evidence:
- Study 1 (N = 400): Participants who generated their own explanations before receiving a standardized explanation showed more elaborated recall and ~10 percentage point higher subsequent error detection than participants who only read the explanation.
- Study 2 (N = 240, 8-day protocol): All participants generated an explanation at baseline; over seven subsequent days participants received either a personalized cue tied to their earlier reasoning or a personalized neutral phrase. Error detection declined with repeated use in both groups but declined significantly more slowly for the cue condition. Additional measures (cue-evocation and cue–episode association) support that the cue reactivated earlier reasoning rather than merely functioning as a salience boost.
- Practical design insight: Lightweight onboarding that prompts self-explanation plus a small, user-linked retrieval cue in daily messages can make oversight more resilient as LLM use becomes routine.
Data & Methods
- Setting: Lab-in-the-field experiments with customer-facing employees at a large enterprise-software firm in China.
- Studies:
- Study 1: Between-subjects randomized experiment, N = 400. Manipulation: self-generated explanation before receiving a standardized explanation vs. receiving the standardized explanation only. Outcomes: error detection rate on subsequent LLM outputs; measures of recall/elaboration of verification-relevant reasoning.
- Study 2: Randomized field experiment, N = 240. All participants generated an explanation at baseline. Over the next seven workdays participants received a daily personalized message: either a self-chosen cue linked to their earlier reasoning (treatment) or an equally personalized but verification-neutral phrase (control). Outcomes: daily error detection rates over time; measures of cue-evocation and cue–episode association to test mechanism.
- Key measures: binary identification of errors in LLM outputs (error detection), elaboration/recall of verification-relevant reasoning, trajectory of detection rates across repeated use, and association/evocation metrics for cues.
- Analysis: Randomization-based comparisons of detection rates and mixed models of detection decline over time; mediation/auxiliary analyses using cue measures to support reactivation mechanism.
- Limitations in methods reported or implied: samples from one firm/industry and country; focus on error detection (not other oversight outcomes like calibration, revision quality); follow-up limited to seven days in Study 2.
Implications for AI Economics
- Cost-effective oversight design: Generative encoding (short self-explanation tasks) and simple personalized retrieval cues are low-cost interventions compared with extensive retraining, continuous system warnings, or intrusive cognitive forcing functions. Organizations can improve oversight ROI by reallocating modest onboarding time and in-app nudges rather than heavier interventions.
- Scalability of human-in-the-loop models: The retrievability lens helps explain why oversight quality degrades as LLM use routinizes. Implementing lightweight cue-based supports can slow that degradation, enabling firms to scale LLM deployment while keeping human safeguards effective and less resource-intensive.
- Risk management and compliance: Slower decay of error detection reduces the probability of erroneous outputs reaching customers or regulators. This decreases expected liability and compliance costs; the approach may be particularly valuable in regulated or high-stakes domains where continuous expert oversight is costly.
- Labor productivity and task allocation: By improving the efficiency and resilience of reviewers’ verification, retrievability interventions could reduce time spent on exhaustive checks while preserving catch rates for substantive errors — affecting staffing models, throughput, and the value proposition of augmenting workers with LLMs.
- Trade-offs and complementarities: Retrievability interventions complement, not replace, capability- and engagement-based measures. For example, investing in AI literacy (capability) and interface design that encourages scrutiny (engagement) remains important; retrievability increases the likelihood that those investments are actually invoked when needed.
- Empirical agenda for economic evaluation: Future work should quantify the monetary benefits (error-costs averted, time saved), compare cost-effectiveness across interventions, and model long-term impacts on labor demand and task reallocation across industries and regulatory regimes.
Notes and avenues for further research: generalize findings across domains and error types; measure long-term durability of self-explanation effects; study interactions between retrievability supports and system-side transparency (e.g., provenance/evidence exposure); and perform cost–benefit analyses to inform deployment decisions in firms and regulated sectors.
Reference: Fu, X.; Ramasubbu, N.; Galletta, D. F. (Preprint; forthcoming in Journal of Management Information Systems, Sep. 2026) — "Knowing Is Not Enough: Information Retrievability as a Precondition to Effective LLM Oversight."
Assessment
Claims (6)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| Self-generated explanations of LLM errors improve subsequent error detection compared with passively receiving the same standardized explanation. Error Rate | positive | The number or proportion of errors detected in subsequent LLM output |
Reading fidelity
high
Study strength
high
|
n=400
approximately a 10-percentage-point improvement
|
| Self-generated explanations produce more elaborated recall of verification-relevant reasoning than passive exposure to an explanation. Skill Acquisition | positive | Recall and elaboration of verification-relevant reasoning |
Reading fidelity
high
Study strength
medium
|
n=400
|
| A personalized daily retrieval cue tied to a participant's earlier oversight reasoning slows the decline in error detection during repeated LLM use. Error Rate | positive | Trajectory of error detection across repeated LLM use over seven workdays |
Reading fidelity
high
Study strength
high
|
n=240
Error detection declined significantly more slowly
|
| The beneficial effect of personalized retrieval cues is consistent with reactivation of participants' earlier verification reasoning, rather than merely with receiving a personalized message. Skill Acquisition | positive | Cue evocation and association with the prior verification episode, alongside the repeated-use error-detection trajectory |
Reading fidelity
high
Study strength
medium
|
n=240
|
| The paper identifies information retrievability as a distinct user-side precondition for effective oversight of LLM output, in addition to capability and engagement. Ai Safety And Ethics | positive | Accessibility of oversight-relevant information during review and its contribution to error detection |
Reading fidelity
high
Study strength
medium
|
n=640
|
| Oversight can fail even when users possess the relevant capability and motivation because oversight-relevant information may be retained but not retrieved during review. Ai Safety And Ethics | negative | Successful access to verification-relevant information during review |
Reading fidelity
high
Study strength
speculative
|
not reported
|