The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

AI tools subtly shape who succeeds: inaccurate captions and mere suspicion of AI reduce perceived quality and hiring chances, while mismatched writing suggestions erode workers’ sense of agency and belonging.

BEYOND FAIRNESS METRICS: HUMAN EXPERIENCES AND DIFFERENTIAL OUTCOMES IN AI-MEDIATED WORK SETTINGS
Kadoma, Kowe · August 26, 2026 · eCommons (Cornell University)
openalex rct medium evidence 8/10 relevance Full text usable extracted full text DOI Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

OpenAlex

Latest observation:

  1. Kadoma, Kowe provider ID
Randomized experiments show that AI-mediated artifacts—error-prone ASR captions, the suspicion that text was AI-generated, and stylistic mismatches in writing assistants—reduce perceived competence, lower stated hiring likelihood, and undermine feelings of inclusion, control, and ownership.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

140 pages

Summary

Main Finding

AI tools embedded in everyday work—even when not making explicit decisions—shape perceptions, evaluations, and experiences in ways that produce differential outcomes. Technical errors (e.g., ASR subtitle mistakes), the suspicion that work was AI-generated, and stylistic mismatches between AI suggestions and users’ voice each reduce perceived competence, hiring likelihood, inclusion, control, and ownership. These harms often fall outside standard, model-centric fairness metrics and require sociotechnical, human-centered measurement and design.

Key Points

  • Broad thesis: Fairness evaluations that focus only on model errors or group-level performance gaps miss important harms that arise from how AI mediates communication and social judgment in workplaces.
  • Study 1 (ASR subtitles): Inaccurate automatic captions rapidly become social signals—viewers judge speakers and content more negatively when subtitles contain errors. Perceived subtitle quality mediates evaluations; errors can amplify pre-existing penalties tied to accent or speaker characteristics.
  • Study 2 (AI suspicion): The mere suspicion that a writing sample was produced with AI reduces quality assessments and hiring likelihood. AI suspicion levels differ by the presented demographic cues (gender, race, nationality) and interact with those cues to produce disparate evaluative outcomes.
  • Study 3 (Writing assistant style): When AI-generated suggestions stylistically mismatch the user (e.g., overly self-assured versus hesitant tones), users report lower feelings of inclusion, lower control over the writing process, and reduced ownership of outputs. A hesitant-style assistant preserved more perceived agency than a self-assured style.
  • Conceptual contribution: Proposes moving beyond model-centric metrics to measure constructs like perceived inclusion, control, ownership, and AI suspicion as first-class outcomes when assessing workplace AI.
  • Normative implication: Building equitable AI for work requires sociotechnical interventions—improving system robustness, giving users control over style, clearer disclosure and policy, and auditing for social effects—not only optimizing model error rates.

Data & Methods

General approach - Domain: AI-mediated workplace tasks (presentations with captions, freelance hiring evaluation, writing assistance). - Methods: Mixed experimental designs (between- and within-subjects) using controlled manipulations and online participant samples to measure downstream social judgments and subjective experience.

Study 1 — Social consequences of ASR errors - Stimuli: Short speaker videos with varied speaker accents and synchronized/un­synchronized audio/video. Subtitles presented in accurate vs. error-prone versions (derived from different ASR transcriptions; Word Error Rate comparisons reported). - Measures: Viewer ratings of speaker competence, trustworthiness, and perceived quality of content; perceived subtitle quality and speaker attribution. - Key manipulation: subtitle accuracy × speaker accent. - Result summary: Error-prone subtitles reduced both speaker and content evaluations; perceived subtitle quality mediated judgments.

Study 2 — AI suspicion, evaluation, and opportunity - Design: Three preregistered experiments manipulating the presented demographics of a freelancer profile (gender, race, nationality) and the writing style of a sample press release (control vs. AI-inducing style). - Stimuli: Profile pages with photo, name, location and a short press release (edited to contain AI-inducing phrasing in the treatment). - Measures: Participant judgments of whether the text was AI-generated (AI suspicion), perceived quality, and hiring likelihood. - Recruitment: Large online participant pools (demographics reported across experiments). - Result summary: AI suspicion predicted significant reductions in quality ratings and hiring likelihood; suspicion varied by demographic cues and produced interaction effects (i.e., unequal impact across demographic groups).

Study 3 — Inclusion and agency in AI-assisted writing - Platform: A browser-based writing task where participants composed messages with inline AI suggestions. GPT-4 prompted to produce suggestions in two distinct styles: “self-assured” vs. “hesitant.” - Measures: Perceived inclusion (assistant made for me / sounds like me), control (over process and final message), ownership (message is mine / sounds like me), AI reliance, and objective editing behavior (acceptance of suggestions). - Result summary: Participants assisted by hesitant-style suggestions reported higher inclusion, control, and ownership than those assisted by self-assured-style suggestions. Inclusion and perceived reliance on AI interacted to predict perceived agency.

Analytical techniques - Standard inferential statistics for experimental contrasts, mediation analyses where appropriate, and OLS regressions to model perceived agency with interaction terms (style × inclusion, demographics × suspicion).

Implications for AI Economics

  • Labor market allocation effects beyond automated decision systems: AI tools that mediate communication (captions, writing assistants) can create allocative harms by altering who is hired, promoted, or assessed favorably—even when AI is not the decision-maker. Economists should model these indirect channels as sources of occupational sorting and wage impacts.
  • Externalities and signaling: AI errors and the social meanings attached to AI use are negative externalities affecting signaling in markets. For example, subtitle errors reduce perceived productivity/competence; AI suspicion acts as a stigma on submitted work. These effects change equilibrium hiring/firing and search behavior and can disproportionately harm already marginalized groups.
  • Productivity vs. distributive trade-offs: While AI can raise average productivity, heterogeneity in social reception (e.g., based on accent-robustness or stylistic fit) can increase inequality. Policies or firm-level decisions that mandate or ban AI use can produce perverse distributional outcomes if suspicion is unevenly applied.
  • Measurement and evaluation: Economic assessments of AI adoption should expand beyond model accuracy and incorporate human-centered outcomes (perceived quality, hiring likelihood, inclusion, agency). Welfare analyses ought to account for both direct productivity gains and indirect social costs.
  • Policy and firm interventions with economic levers:
    • Invest in robustness where social harms concentrate (e.g., improve ASR for diverse accents) to reduce quality-of-service and downstream labor market harms.
    • Design disclosure and usage policies that minimize discrimination from suspicion (transparent, consistent rules rather than ad hoc bans that rely on unreliable detection).
    • Provide user controls and stylistic customization in tools to preserve worker agency and ownership—this may improve retention/productivity and reduce turnover costs.
    • Audit sociotechnical impacts: cost–benefit analyses and impact assessments should quantify effects on hiring probabilities and perceived worker value across groups.
  • Modeling opportunities for researchers: Incorporate behavioral responses to AI-mediated signals (e.g., suspicion penalties, stylistic mismatch costs) into search-and-matching, employer signaling, and human-capital models to predict distributional impacts of AI uptake.
  • Broader market design considerations: Platforms and firms that interface between workers and evaluators (e.g., freelancing marketplaces, applicant tracking systems) should be evaluated for how their deployment of mediating AI changes market frictions, matching quality, and inequality.

Overall recommendation for economists and policymakers: Treat workplace AI as sociotechnical infrastructure with measurable social externalities. Empirical and theoretical models of AI’s economic effects must include interpersonal interpretation, stigma from AI suspicion, and subjective experiences of agency and inclusion to accurately predict distributional outcomes and to design effective interventions.

Assessment

Paper Typerct Evidence Strengthmedium — Internal validity is strong because of randomized assignment, multiple replications across related experiments, manipulation checks, and careful measurement of perceptions and stated hiring likelihood; however external validity is limited by artificial tasks, low-stakes online evaluation settings (vs real hiring or workplace outcomes), likely convenience panel samples, and measured outcomes are perceptions/self-reports rather than observed long-run economic outcomes. Methods Rigorhigh — Multiple pre-registered randomized experiments with careful stimulus construction (controlled video/audio, ASR manipulations, GPT-4–generated suggestions), manipulation checks, exploration of moderators (accent, gender, race, nationality, AI suspicion), and appropriate analyses; noted limitations include ecological validity, potential demand characteristics, and reliance on stated hiring likelihood rather than behavioral field outcomes. SampleMultiple online experimental samples of adult participants recruited via online panels (details and exact Ns not included in the excerpt). Participants completed one of three tasks: watched short speaker videos with manipulated ASR subtitles and rated speaker/content; evaluated freelancer profiles and press releases with manipulations designed to induce AI suspicion and varying presented demographics (photo/name/location) and reported quality/hiring likelihood; or performed a writing task with in-line AI suggestions (GPT-4) in different stylistic modes and reported inclusion, control, and ownership. Stimuli included actor videos with different accents, engineered ASR transcripts (accurate vs error-prone), GPT-4 generated suggestions in different styles, and experimentally varied profile photos/names. Themeshuman_ai_collab labor_markets inequality org_design productivity adoption IdentificationRandomized controlled experiments varying (a) ASR subtitle accuracy and presence, (b) whether text is AI-inducing (to manipulate observers' AI suspicion) crossed with presented demographic cues (photo, name, location), and (c) writing-assistant suggestion style (hesitant vs self-assured) with preregistration, manipulation checks, and between-subjects assignment to isolate causal effects of the treatments on perception, hiring likelihood, and subjective experience. GeneralizabilityOnline convenience samples may not represent hiring managers or real workplace decision-makers, Low-stakes experimental settings (surveys, single-session tasks) differ from high-stakes, longitudinal hiring and job-performance contexts, Stimuli (short videos, synthetic subtitles, single press-release samples, GPT-4 suggestions) may not capture the complexity and diversity of actual workplace artifacts, Cultural and language scope may be limited (likely Anglophone/US-centric), restricting transfer to other countries/languages, Measured outcomes are perceptions and stated intentions (hiring likelihood), not actual hires, wages, or productivity metrics

Claims (8)

ClaimDirectionOutcomeConfidence & EvidenceDetails
AI systems can produce differential outcomes even when they do not explicitly make decisions about people. Ai Safety And Ethics negative Differential social, evaluative, hiring, and experiential outcomes in AI-mediated work
Reading fidelity high
Study strength medium
not reported
0.6
Inaccurate ASR-generated subtitles can lead audiences to evaluate the speaker and the speaker's content more negatively, turning technical transcription errors into social judgments. Decision Quality negative Audience evaluations of the speaker and the speaker's content
Reading fidelity high
Study strength medium
not reported
0.6
Accurate subtitles are better received than error-prone subtitles. Output Quality positive Reception of subtitle quality
Reading fidelity high
Study strength medium
not reported
0.6
Mere suspicion that a person used an AI tool can negatively affect evaluations of their work, regardless of whether the tool was actually used. Output Quality negative Perceived quality of a freelancer's work
Reading fidelity high
Study strength medium
not reported
0.6
Suspicion that a freelancer used AI negatively affects the likelihood that evaluators will hire the freelancer. Hiring negative Evaluator-reported hiring likelihood
Reading fidelity high
Study strength medium
not reported
0.6
A mismatch between an AI writing assistant's suggestions and the user's writing voice can undermine the user's feelings of inclusion, control, and ownership. Worker Satisfaction negative Perceived inclusion, control over the writing process and message, and ownership of the final text
Reading fidelity high
Study strength medium
not reported
0.6
Participants assisted by a hesitant-writing-style model reported greater control over the final message and the writing process than participants assisted by a self-assured-writing-style model. Worker Satisfaction positive Perceived control over the final message and writing process
Reading fidelity high
Study strength medium
not reported
0.6
Nearly 40% of American workers were already using AI in their workflows, according to a Gallup poll cited by the dissertation. Adoption Rate positive Workplace AI use
Reading fidelity high
Study strength low
nearly 40%
0.3

Notes