The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Clear explanations, not a human face, drive acceptance of AI performance reviews: high-quality explanations sharply raise perceived fairness and trust, while the evaluator being an AI versus a human matters little when explanations are strong.

Beyond Algorithm Aversion: How Explanation Quality Enables Fairness and Trust in AI-Based Performance Evaluations
McGraw, Darren · January 01, 2026 · Digital Commons - George Fox University (George Fox University)
openalex rct medium evidence 7/10 relevance Summary only summary available; pdf_status=paywall Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

OpenAlex

Latest observation:

  1. McGraw, Darren provider ID
In a randomized vignette experiment of 275 U.S. workers, high-quality explanations greatly increased perceived fairness and trust in AI-based performance evaluations, while whether the evaluator was human or AI had little independent effect when explanations were high-quality.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

Performance evaluation systems increasingly incorporate artificial intelligence, yet employee acceptance remains limited. This experimental study examined how explanation quality and evaluator type influenced fairness perceptions and trust in AI-based performance evaluations among 275 U.S. working adults. Using a 2 × 2 vignette design, results demonstrated that explanation quality produced large effects on perceived fairness (d = 1.79) and trust, while evaluator type—human versus AI—had negligible independent effects (d = −0.08) when explanations were high quality. Fairness mediated the explanation-trust relationship (indirect effect = 1.95, 95% CI [1.65, 2.26]), with inconsistent mediation indicated by a negative direct effect (b = −0.81, p < .001) when fairness was controlled. Findings suggest organizations should prioritize explanation quality over evaluator source when implementing AI evaluation systems.

Summary

Main Finding

High-quality explanations drive large increases in perceived fairness and trust for performance-evaluation systems; whether the evaluator is an AI or a human matters little once explanation quality is high. Fairness mediates the effect of explanation quality on trust, but an inconsistent mediation (negative direct effect) suggests there are additional, countervailing dynamics when fairness is held constant.

Key Points

  • Sample and design: 275 U.S. working adults, between-subjects 2 × 2 vignette experiment (evaluator type: human vs AI; explanation quality: high vs low).
  • Primary outcomes: perceived fairness and trust in the evaluator/system.
  • Effect sizes:
    • Explanation quality → perceived fairness: very large (d = 1.79).
    • Explanation quality → trust: large effect (reported as strong; mediation results below quantify relationship).
    • Evaluator type (human vs AI) → trust or fairness: negligible independent effect when explanations are high (d = −0.08).
  • Mediation:
    • Fairness mediates the explanation → trust effect (indirect effect = 1.95, 95% CI [1.65, 2.26]).
    • Inconsistent mediation observed: a negative direct effect of explanation quality on trust when fairness is controlled (b = −0.81, p < .001), indicating other processes that reduce trust conditional on fairness.
  • Practical summary: improving explanation quality substantially increases perceived fairness and trust; swapping a human for an AI evaluator has little impact if explanations are good.

Data & Methods

  • Participants: 275 working adults in the United States (online vignette respondents; demographic representativeness not specified).
  • Experimental design: 2 (evaluator: human vs AI) × 2 (explanation quality: high vs low) vignette study; participants assigned to one of four conditions and reported perceptions of fairness and trust.
  • Measures: self-report scales for perceived fairness and trust (standard psychometric outcome measures typical for vignette research).
  • Analysis: between-group effect size estimates (Cohen’s d), mediation analysis estimating indirect effect of explanation quality on trust via fairness with confidence intervals; report of direct effect controlling for mediator.
  • Limitations of method: vignette and self-report design limits ecological validity and behavioral inference; sample size moderate; possible selection or demographic biases; short-term perceptions rather than longitudinal outcomes.

Implications for AI Economics

  • Adoption and investment priorities: Firms should prioritize investments in high-quality explainability features (transparent, comprehensible justifications of evaluations) over focusing on whether evaluations are labeled “AI” or “human.” The large effect of explanation quality implies high returns to spending on explanation design, UX, and explanation auditing.
  • Labor-market friction and turnover risk: Perceived fairness strongly influences trust; poor explanations may raise perceived unfairness and increase turnover or reduce cooperation. Explainability investments can reduce these frictions and associated costs (recruitment, productivity loss).
  • Compensation and bargaining: Explanations that increase perceived fairness can stabilize acceptance of algorithmic evaluation systems in pay-for-performance schemes and automated promotion or bonus decisions, lowering bargaining frictions and litigation risk.
  • Regulatory and compliance incentives: High-quality explanations align with regulatory transparency demands (e.g., informational requirements under data-protection regimes), reducing legal and compliance risk and potential costs.
  • Signaling and adoption externalities: Because evaluator source has little independent effect when explanations are strong, firms can deploy AI evaluators without a major reputational penalty—provided they offer high-quality explanations. This shapes the signaling calculus: discloseable explanation quality becomes the key signal to employees and stakeholders.
  • Design trade-offs and unexpected effects: The inconsistent mediation (negative direct effect when fairness is controlled) cautions that very detailed explanations may reveal imperfections or set expectations that undermine trust in ways not mediated by fairness. Designers should test explanation content and salience to avoid backfiring (e.g., exposing uncertainty, error rates, or mechanistic detail that reduces confidence).
  • Research and policy priorities: Economic evaluations (cost–benefit analyses) of AI adoption should include the value of explanation quality (reduced turnover, higher productivity, lower litigation probability) and consider heterogeneity across occupations and worker demographics. Policy should encourage standards for explanation quality rather than focusing solely on whether humans or machines make decisions.

Suggestions for next steps (research & practice) - Field experiments evaluating behavioral outcomes (turnover, grievance filings, performance) under real AI evaluation systems with varying explanation quality. - Costing studies to estimate ROI of explainability investments versus other interventions (training, appeals processes). - Decomposing what aspects of explanation quality (comprehensibility, completeness, actionable guidance) drive effects and testing for heterogeneous treatment effects across job types and worker characteristics.

Assessment

Paper Typerct Evidence Strengthmedium — Internal validity is strong because of random assignment and clear experimental manipulations with large reported effect sizes; however, external validity is limited by the hypothetical vignette design, reliance on self-reported perceptions rather than observed behavior or outcomes, a moderate sample size (N=275), and limited information about sample representativeness. Methods Rigormedium — The study uses a well-specified factorial randomized design and reports effect sizes and mediation analysis, indicating solid analytic practice; but rigor is reduced by reliance on a single vignette study (no field validation), potential measurement and manipulation-check issues not reported here, and lack of information about pre-registration, robustness checks, or longer-term behavioral measures. Sample275 U.S. working adults who completed an online between-subjects vignette experiment (2×2: explanation quality high vs low × evaluator type human vs AI); outcomes are self-reported perceived fairness and trust in the evaluation system. Themesadoption org_design human_ai_collab IdentificationRandom assignment to a 2×2 between-subjects vignette experiment manipulating explanation quality (high vs low) and evaluator type (human vs AI); causal effects are inferred from randomization and mediation analysis tests perceived fairness as a mediator of explanation → trust. GeneralizabilityVignette-based, hypothetical scenarios may not reflect real workplace behavior or long-term acceptance, Sample limited to U.S. working adults (likely convenience/online sample) so not nationally representative, Findings pertain specifically to performance-evaluation contexts and may not generalize to other AI applications, Relies on self-reported perceptions rather than objective adoption, performance, or labor-market outcomes, Single study with moderate N; replication across sectors and cultures needed

Claims (8)

ClaimDirectionOutcomeConfidence & EvidenceDetails
Performance evaluation systems increasingly incorporate artificial intelligence, yet employee acceptance remains limited. Worker Satisfaction negative employee acceptance
Reading fidelity medium
Study strength low
not reported
0.18
This experimental study examined how explanation quality and evaluator type influenced fairness perceptions and trust in AI-based performance evaluations among 275 U.S. working adults. Other null_result fairness perceptions and trust
Reading fidelity high
Study strength high
n=275
1.0
Explanation quality produced a large positive effect on perceived fairness (d = 1.79). Worker Satisfaction positive perceived fairness of performance evaluations
Reading fidelity high
Study strength medium
n=275
d = 1.79
0.6
Explanation quality produced large positive effects on trust in AI-based performance evaluations. Worker Satisfaction positive trust in AI-based performance evaluations
Reading fidelity high
Study strength medium
n=275
0.6
Evaluator type (human versus AI) had negligible independent effects (d = −0.08) when explanations were high quality. Worker Satisfaction null_result perceived fairness and trust (negligible independent effect of evaluator source)
Reading fidelity high
Study strength medium
n=275
d = −0.08
0.6
Fairness mediated the explanation–trust relationship (indirect effect = 1.95, 95% CI [1.65, 2.26]). Worker Satisfaction positive trust in AI-based performance evaluations (mediated by perceived fairness)
Reading fidelity high
Study strength medium
n=275
indirect effect = 1.95, 95% CI [1.65, 2.26]
0.6
There was inconsistent mediation: the direct effect of explanation quality on trust, controlling for fairness, was negative (b = −0.81, p < .001). Worker Satisfaction mixed direct effect of explanation quality on trust (controlling for fairness)
Reading fidelity high
Study strength medium
n=275
b = −0.81, p < .001
0.6
Findings suggest organizations should prioritize explanation quality over evaluator source when implementing AI evaluation systems. Adoption Rate positive organizational implementation priority (explanation quality vs evaluator source)
Reading fidelity high
Study strength medium
n=275
0.6

Notes