The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

A job-task grounded taxonomy finds most workplace AI risks arise from human–agent interaction — incorrect agent actions and gradual skill erosion dominate; augmentation risks worker capability and wellbeing, while automation shifts harms toward operational and financial failures.

Unaccountable Delegation, Fading Skills: Mapping the Risks of Workplace AI Agents
Gabriele La Malfa, Lakmal Meegahapola, Edyta Bogucka, Jie M. Zhang, Michael Luck, Elizabeth Black, Daniele Quercia · August 09, 2026
arxiv descriptive medium evidence 7/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

Arxiv

Latest observation:

  1. Gabriele La Malfa unresolved corpus identity
  2. Lakmal Meegahapola unresolved corpus identity
  3. Edyta Bogucka unresolved corpus identity
  4. Jie M. Zhang unresolved corpus identity
  5. Michael Luck unresolved corpus identity
  6. Elizabeth Black unresolved corpus identity
  7. Daniele Quercia unresolved corpus identity

Semantic Scholar

Latest observation:

  1. G. Malfa provider ID
  2. L. Meegahapola provider ID
  3. E. Bogucka provider ID
  4. Jie M. Zhang provider ID
  5. Michael Luck provider ID
  6. Elizabeth Black provider ID
  7. Daniele Quercia provider ID
Using O*NET job tasks, LLM-generated prospective scenarios, existing incident records, and worker validation, the paper produces a 15-category workplace AI agent risk taxonomy and finds that erroneous agent actions and human capability erosion dominate, with augmentation tending to erode skills/oversight and automation concentrating organizational risks.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

To anticipate socio-technical risks from AI agents, organizations need taxonomies to classify them. However, existing AI risk taxonomies focus on broad risks and do not capture job-specific risks introduced by agents. To address this gap, we make three main contributions. First, we developed a multi-layer framework from a literature review of AI agents. The framework models three core components and their interactions: agents, goals, and environment. Second, we embedded this framework in a structured prompt and applied it to descriptions of 2,078 job tasks from the O*NET database, producing 8,356 risk scenarios labeled by severity and deployment mode (automation or augmentation). We validated these scenarios with 45 workers across 10 job roles and an independent LLM judge, confirming their plausibility and alignment with job tasks. Finally, we extended an existing taxonomy to create a 15-category taxonomy of workplace AI agent risks that covers all our risk scenarios. Our analysis highlights four findings. First, augmentation is not inherently safe because overreliance on agents can gradually erode workers' skills and oversight. Second, Erroneous Agent Actions accounts for the largest share of risk scenarios and has the highest concentration of severe risks. Many arise at the human-agent boundary. Third, automation is associated mainly with organizational risks, while augmentation is associated mainly with risks to workers. Fourth, workers found our taxonomy easier to use for a risk classification task than two other taxonomies and preferred it in 64% of non-tied comparisons with a recent generative AI risk taxonomy. These findings show that workplace AI agent risks do not arise from agents alone; they also depend on how people work with agents and how agents are deployed. Safer workplaces require not only safer agents but also carefully designed human-AI agent collaboration.

Summary

Main Finding

Workplace risks from AI agents are predominantly socio-technical: they arise not only from agent capabilities but from how agents, goals, environments, and workers interact. Augmentation is not inherently safer than automation — it concentrates risks on worker skill erosion and oversight failures, while automation shifts risks toward organizational and financial failures. The authors produce a job-task–grounded, forward-looking 15-category taxonomy of workplace AI agent risks and show that erroneous agent actions and human capability erosion together account for over half of anticipated risk scenarios.

Key Points

  • Multi-layer framing: Risks are organized across Technical Capability (agents, goals, environment), Human Interaction (worker–agent and workplace dynamics), and Systemic Impact (organizational, economic, societal).
  • Large-scale scenario generation: 8,356 prospective, job-task–specific risk scenarios were generated from 2,078 computer-based O*NET job tasks (209 roles, 20 industries) using a structured prompt and an LLM.
  • Taxonomy: Extended a generative-AI taxonomy into a 15-category workplace AI agent risk taxonomy (44 subcategories). Top categories:
    • Erroneous Agent Actions — 30.6% of scenarios (highest concentration of severe risks)
    • Human Capability Erosion — 21.3% (skill loss, oversight erosion)
    • Together these two = 51.9% of scenarios; predominantly appear under augmentation
  • Deployment mode contrasts:
    • Augmentation → worker-centered risks (skill atrophy, psychological/social harms, degraded oversight)
    • Automation → organizational risks (operational failures, financial losses)
  • Human–agent boundary: Many severe risks arise at interfaces where workers rely on, interpret, or override agent outputs.
  • Usability: Workers found this taxonomy easier to use than two alternatives and preferred it over a recent generative-AI taxonomy in 64% of non-tied comparisons.
  • Practical takeaway: Safer workplaces require both safer agents and careful human–AI collaboration design (procedures, oversight, incentives).

Data & Methods

  • Task corpus: 2,078 computer-based job tasks from O*NET (filtered set thought likely to be impacted by AI agents).
  • Scenario generation: Used gpt-4o-mini with a structured prompt embedding the multi-layer AAIS framework to produce 8,356 risk scenarios. Each scenario labeled by:
    • Severity: minimal, limited, high, critical
    • Deployment mode: augmentation (agent assists worker) or automation (agent performs task without worker)
  • Validation:
    • Human validation: 45 workers across 10 job roles evaluated 450 scenarios from their roles.
      • Plausibility: mean 4.03 / 5
      • Task alignment: mean 3.94 / 5
    • LLM-as-judge: Independent LLM assessed the full corpus.
      • Plausibility: mean 4.81 / 5
      • Task alignment: mean 4.60 / 5
  • Grounding & extension: Augmented generated scenarios with 222 public workplace AI incidents and 11 live multi-agent red-team failure cases (combined corpus ≈ 8,589 scenarios). Mapped scenarios to existing generative AI taxonomy and created new categories where needed (resulting in 15 categories, 44 subcategories).
  • Taxonomy evaluation: Assessed structural integrity, coverage against 10 established frameworks, and practical usability (comparative study with 26 workers).
  • External resources: Supplementary materials and full datasets available at authors’ project page.

Implications for AI Economics

  • Labor supply and human capital
    • Skill depreciation: Prolonged augmentation can cause gradual erosion of task-specific skills and oversight capabilities. Economic models of labor supply and human capital accumulation need to incorporate skill depreciation from reliance on AI (not just displacement).
    • Career dynamics: Declines in skill breadth or depth may reduce workers’ future employability and upward mobility, shifting lifetime earnings trajectories.
  • Wages and task valuation
    • Task-level wage effects: Traditional measures that map tasks to AI exposure should differentiate augmentation (risk of skill erosion, potential long-term wage suppression) from automation (direct displacement). Wage impacts may be nonlinear and dynamic.
    • Compensating differentials: Workers may demand higher pay or contractual protections for roles exposed to augmentation-induced erosion or high oversight burdens.
  • Productivity, firm decisions, and TFP measurement
    • Short-run vs long-run productivity: Firms may gain immediate efficiency from agents but incur long-run human-capital costs (hidden externalities) that depress productivity or raise future training costs. Measured TFP gains from adoption may overstate durable gains if human capital loss is not accounted for.
    • Adoption choice under uncertainty: Firms face trade-offs — automation concentrates organizational risk (operational, financial), while augmentation shifts latent risk to workforce capability. Economic models of technology adoption should internalize these different risk profiles and the costs of mitigation (training, audits, redundancy).
  • Incentives, contracting, and liability
    • Contracts and incentives should account for overreliance risks (monitoring, performance metrics that encourage gaming, or de-skilling). Principal–agent problems reappear in new form (who monitors agent outputs; who bears risk of erroneous actions).
    • Insurance and liability markets: New risk categories (erroneous agent actions at human–agent boundaries; systemic operational failures) open demand for tailored liability insurance and create moral-hazard considerations.
  • Measurement and empirical research
    • Need for refined tools: Empirical studies should use task-level, deployment-mode aware measures (augmentation vs automation) and attempt to capture human-capability erosion and oversight costs, not just displacement.
    • Data collection: Longitudinal data on workers’ task performance, re-skilling expenditures, error rates, and oversight hours is critical to quantify economic impacts and to distinguish transitory from persistent effects.
  • Policy and redistribution
    • Training & upskilling subsidies: Public investment may be needed to counteract skill erosion from augmentation (e.g., mandated retraining, certification, rotational task policies).
    • Regulation of deployment: Policies could require human-in-the-loop safeguards, monitoring, and transparency for agentic systems in high-stakes tasks to mitigate externalities that firms may underweight.
    • Labor bargaining: Collective bargaining may shift to include protections against de-skilling (maintenance of manual practice, limits on full automation for certain tasks).
  • Research directions for AI economists
    • Integrate dynamic human capital depreciation into models of technology adoption and wage setting.
    • Estimate firm-level costs of oversight, audits, and mitigation required for safe augmentation — include these in cost-benefit analyses of agent deployment.
    • Study long-term macro effects of widespread augmentation-induced skill erosion on occupational mobility, wage inequality, and structural unemployment.

Short recommendation for economists: when evaluating AI in the workplace, distinguish augmentation from automation in empirical designs; model skill depreciation and oversight costs explicitly; collect longitudinal, task-level data to quantify hidden human-capital externalities and to inform optimal firm and policy responses.

References and materials: see paper and project page (https://social-dynamics.net/ai-risks/workplace-agents) for the taxonomy, scenario corpus, and supplementary methods.

Assessment

Paper Typedescriptive Evidence Strengthmedium — The paper combines multiple evidence sources (2,078 O*NET job tasks, 8,356 LLM-generated risk scenarios, 222 documented incidents, 11 red-team failure cases) and triangulates with human validation (45 workers) and an independent LLM judge; however, the core corpus is prospectively generated by an LLM and validated with a small, non-representative human sample and an LLM judge, limiting empirical strength for claims about prevalence or impact. Methods Rigormedium — Clear, reproducible pipeline (framework definition, O*NET grounding, structured prompting, multi-source corpus assembly, and validation). Use of a single generation model (gpt-4o-mini) and small human validation sample, potential prompt/model biases, and subjective severity labels reduce rigor relative to fully empirical or experimental work. Sample2,078 computer-based job tasks (from 209 job roles across 20 industries) drawn from a filtered O*NET subset; generated 8,356 prospective risk scenarios using gpt-4o-mini via a structured prompt; validation comprised 45 workers across 10 job roles evaluating 450 scenarios and an independent LLM judge assessing the full corpus; supplemented with 222 public workplace AI incident records and 11 multi-agent red-team failure cases to produce a combined corpus (~8,589 scenarios). Themeshuman_ai_collab org_design GeneralizabilityO*NET-based sample is US-centric and restricted to computer-based tasks, limiting representativeness across countries and non-computer jobs, Filter selection (Shao et al.) may bias which tasks were considered AI-relevant, Major reliance on LLM-generated scenarios means results reflect prompt design and model behavior at time of study, Human validation sample is small (45 workers) and covers only 10 roles, limiting external validity across occupations and industries, Incident databases used to supplement scenarios are incomplete and biased toward reported/publicized events, Rapid evolution of AI agents means taxonomy and scenario prevalence may change over short timescales, Cultural, regulatory, and sectoral differences (e.g., healthcare vs. manufacturing) may limit direct applicability of some categories

Claims (11)

ClaimDirectionOutcomeConfidence & EvidenceDetails
The paper introduces a multi-layer framework for workplace AI-agent risks organized into Technical Capability, Human Interaction, and Systemic Impact. Governance And Regulation positive Coverage and organization of workplace AI-agent risk categories
Reading fidelity high
Study strength medium
n=3
0.18
Applying the framework to 2,078 computer-based O*NET job tasks generated 8,356 prospective workplace AI-agent risk scenarios. Automation Exposure positive Number of job-specific AI-agent risk scenarios generated
Reading fidelity high
Study strength low
n=2078
8,356 risk scenarios
0.09
Workers and an independent LLM judge rated the generated risk scenarios as highly plausible. Ai Safety And Ethics positive Perceived plausibility of generated workplace AI-agent risk scenarios
Reading fidelity high
Study strength medium
n=45
workers: 4.03/5; LLM: 4.81/5
0.18
Workers and an independent LLM judge found that the generated risk scenarios aligned strongly with the corresponding job tasks. Task Allocation positive Alignment of generated risk scenarios with source job tasks
Reading fidelity high
Study strength medium
n=45
workers: 3.94/5; LLM: 4.60/5
0.18
Erroneous Agent Actions was the largest risk-taxonomy category, accounting for 30.6% of all generated risk scenarios. Error Rate negative Share of risk scenarios involving erroneous agent actions
Reading fidelity high
Study strength low
n=8356
30.6%
0.09
Human Capability Erosion accounted for 21.3% of all risk scenarios, making it the second-largest taxonomy category. Skill Obsolescence negative Share of risk scenarios involving erosion of worker capabilities
Reading fidelity high
Study strength low
n=8356
21.3%
0.09
Erroneous Agent Actions and Human Capability Erosion together represented 51.9% of all risk scenarios and occurred mostly under augmentation. Skill Obsolescence negative Combined prevalence of agent-error and worker-capability-erosion risks by deployment mode
Reading fidelity high
Study strength low
n=8356
51.9%
0.09
Augmentation is not inherently safe because overreliance on AI agents can gradually erode workers’ skills and oversight capabilities. Skill Obsolescence negative Worker skill retention and ability to oversee AI-agent decisions
Reading fidelity high
Study strength speculative
n=8356
0.03
Automation was associated mainly with organizational risks such as operational failures and financial losses, whereas augmentation was associated mainly with risks to workers, including capability erosion and psychological and social risks. Organizational Efficiency mixed Distribution of workplace AI-agent risks across organizational and worker outcomes by deployment mode
Reading fidelity high
Study strength low
n=8356
0.09
The workplace AI-agent risk taxonomy contains 15 categories and 44 sub-categories. Governance And Regulation positive Number of categories and sub-categories in the workplace AI-agent risk taxonomy
Reading fidelity high
Study strength medium
n=8589
15 categories and 44 sub-categories
0.18
Workers preferred the paper’s workplace AI-agent risk taxonomy in 64% of non-tied comparisons with a recent generative-AI risk taxonomy. Training Effectiveness positive Worker preference for the workplace AI-agent risk taxonomy in comparative classification tasks
Reading fidelity high
Study strength medium
n=26
64% of non-tied comparisons
0.18

Notes