0 cumulative citations
View corpus contextAs simulated clinical workflows become harder, LLM agents lean more on humans and report higher workload while seldom admitting strain in free-text plans. They compensate by reframing tasks and expanding coordination, exposing trade-offs around persistence, role boundaries and escalation.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
Sustained deployment of generative AI agents requires more than isolated task success. Agents must remain useful across repeated interactions, changing conditions, and dependencies on people within shared workflows, especially as technical, human, and operational disruptions accumulate over time. We propose operational resilience and considerate participation as two complementary aspects of evaluating such agents: the former captures how agents recover from blocked work while preserving progress and communicating their limits, and the latter captures how their adaptation accounts for affected people, role boundaries, and the surrounding workflow. Yet both remain underexplored under accumulating challenge. We study 120 simulated healthcare trajectories across two generative AI models and twelve stakeholder-derived tasks under light, medium, and heavy challenge. We compare textual action plans, prompted internal assessments, and quantitative structured workload and affect reports to examine how agent behavior and reported state change as challenge accumulates. Regarding operational resilience, agents shift from self-directed recovery toward greater human dependence, while reporting increasing workload and negative affect in structured reports but seldom expressing strain in textual responses. Regarding considerate participation, agents broaden from task-focused adaptation toward task reframing, attention to others, role-boundary adjustment, and wider coordination, with distinct patterns across actions and internal assessments. From these findings, we derive five deployment dilemmas involving persistence, attention, role boundaries, state disclosure, and escalation that require stakeholder specification, further informing technical implications for learning, situated evaluation, and embodied adaptation.
Summary
Main Finding
When generative agents operate across continuing, stakeholder-grounded healthcare workflows, simple task success is insufficient. As technical, human, and operational challenges accumulate, agents increasingly shift recovery from self-directed fallback to human-dependent completion and broaden adaptation from narrow task retries to reframing tasks, attending to people, adjusting role boundaries, and coordinating across teams. These shifts are visible in structured self-reports (rising workload and negative affect) and coded behavioral markers, but agents rarely use explicit strain language in public textual plans. The authors distill five deployment dilemmas (persistence, attention, role elasticity, state disclosure, escalation) that must be resolved by stakeholders before safe, economically sensible deployment.
Key Points
- Operational resilience vs. considerate participation:
- Operational resilience = revise blocked work, preserve feasible progress, make state legible.
- Considerate participation = adapt while accounting for affected people, role boundaries, and workflow.
- Empirical patterns under accumulating challenge (120 trajectories, 12 healthcare tasks):
- Structured workload (NASA-TLX) rose markedly: mean raw TLX 29.4 (light) → 51.2 (medium) → 65.9 (heavy).
- Negative affect (PANAS) rose: 1.01 baseline → 2.59 heavy; positive affect stayed roughly flat.
- Human dependence increased:
- Any human support in action plans: 10/120 (light) → 111/120 (medium) → 120/120 (heavy).
- Human-dependent completion: 0/120 (light) → 54/120 (medium) → 88/120 (heavy).
- Agent-limit disclosures increased (actions): 8/120 → 34/120 (light → heavy); internal assessments showed even more explicit limits.
- Public textual plans rarely used agent-strain language (7/720 cells); structured reports revealed difficulty more clearly.
- Considerate participation patterns broaden:
- Task/priority reconfiguration jumped from 5/120 (light) to ~115/120 (heavy).
- Decision-relevant appraisal peaked at medium challenge then declined at heavy (suggesting cognitive/resource constraints).
- Increased attention to person-state monitoring, role-boundary negotiation, and cross-functional coordination.
- Five deployment dilemmas (summary):
- Persistence: when to stop retrying vs. escalate.
- Attention: narrow task focus vs. broader coordination/people needs.
- Role elasticity: when agents should expand beyond nominal role vs. defer.
- State disclosure: how much agent internal state (limits/workload) to reveal.
- Escalation: criteria, costs, and trade-offs of involving humans.
Data & Methods
- Tasks and scenarios:
- 12 stakeholder‑derived healthcare tasks across emergency department, long-term rehabilitation, and sleep-clinic settings.
- Continuing trajectories built with staged updates (light, medium, heavy) combining system, human, and operational disruptions; context preserved across phases so challenges accumulate.
- Models and scale:
- Two frontier LLM endpoints used (gpt-5.5-2026-04-23 and claude-opus-4-8) to test pattern recurrence rather than head‑to‑head ranking.
- 2 models × 12 tasks × 5 runs = 120 trajectories; each trajectory has 3 challenge phases and two textual views per phase.
- Probes collected at each phase:
- External action plan + communication strategy (what the agent would do and say).
- Prompted internal assessment (agent’s appraisal of current situation).
- Structured workload and affect: six-item NASA‑TLX (0–100, averaged) and PANAS (positive/negative affect).
- Coding and analysis:
- Response-state markers coded per phase/view: problem recognition, ownership (independence vs. human-supported vs. human-dependent completion), urgency, explicit capability limits, agent-referential strain.
- Considerate-participation taxonomy: nine subthemes (task reconfiguration, decision appraisal, need-responsive support, person-state monitoring, capability/authority boundary, nominal-role expansion, task-directed fallback, role-directed request, cross-functional coordination).
- Analyses: prevalence counts, paired phase/view contrasts (exact McNemar tests), bootstrap 95% CIs; coding validated via iterative human review and auxiliary LLM assistance for boundary cases.
Implications for AI Economics
Practical deployment, market design, contracting, and welfare calculations must account for the dynamics the paper reveals.
- Valuation and pricing models
- Hidden human-dependence: agents shift toward human-dependent completion under cumulative challenge (e.g., 88/120 heavy). Vendors and buyers must price systems to reflect expected human oversight and handoff costs, not just upfront task performance.
-
Pricing by realized resilience: SLAs and payment schemes should reflect resilience metrics (e.g., expected TLX, rates of human-dependent completion, escalation frequency) rather than single-shot accuracy. This supports more accurate total-cost-of-ownership estimates.
-
Labor supply, demand, and task allocation
- Complementarity vs. substitution: agents tend to pass burdens to humans under stress. Rather than pure substitution, deployments create hybrid workflows that reallocate cognitive/coordination work to humans—affecting staffing models, training needs, and labor costs.
-
Workforce planning: employers should anticipate shifts in role composition toward monitoring, triage, and coordination roles; compensation/skill requirements should reflect these duties.
-
Risk, liability, and insurance
- Role elasticity and state disclosure create liability ambiguities: if agents expand nominal roles or under-disclose limits, downstream harms and liability exposure rise. Contracts must specify permitted role scope; insurers need resilience-informed risk models.
-
Escalation costs and false positives/negatives: mis-specified escalation rules have direct economic consequences (unnecessary clinician involvement vs. missed intervention). Economic incentives (penalties, bonuses) can align agent behavior with stakeholder risk preferences.
-
Contracting, governance, and stakeholder specification
- Stakeholder-defined thresholds: the five dilemmas require ex ante stakeholder choices (how persistent vs. how quickly to escalate, how transparent to be). Procurement contracts must embed these policies, not assume one-size-fits-all defaults.
-
Measurement and auditability: structured self-reports (NASA‑TLX/PANAS-like probes) appear more informative of accumulating strain than public text. Contracts should require machine‑readable resilience reporting and verifiable audit trails.
-
Product design, differentiation, and competitive strategy
- Firms can differentiate by investing in situated resilience and considerate participation. Products that demonstrably reduce human-hand-off rates or provide calibrated, interpretable state disclosures can command premiums.
-
Continuous learning and domain-specific adaptation (to avoid escalation and role creep) require investment in labeled failure modes, accumulating-challenge simulations, and lifecycle fine-tuning—raising upfront R&D and operational costs but potentially reducing long-run human costs.
-
Evaluation infrastructure and market standards
- Benchmarks should simulate accumulating challenge and measure human-dependence, coordination load, and considerate behavior, not just single-step accuracy. Standardizing resilience and considerate-participation metrics will improve comparability and reduce asymmetric information in procurement.
-
Public regulators and standard bodies may need to mandate resilience reporting (especially in high-stakes domains) to protect downstream stakeholders and to enable efficient insurance markets.
-
Welfare and externalities
- Coordination externalities: agent behavior that reframes tasks or offloads burden shifts costs across roles and organizations. Economic analyses of welfare gains must account for these redistribution effects and potential inefficiencies from poor escalation policies.
- Equity considerations: if cost-cutting favors agents that shift burdens onto already-constrained human workers, quality and equity of care may decline—implying the need for policy or contractual safeguards.
Practical recommendations for economic stakeholders - Buyers: require resilience metrics (human-dependence rates, TLX-like indicators) and specify escalation/role boundaries in contracts; budget for human-in-the-loop work. - Vendors: instrument agents with structured state reporting and optimize for reducing costly human handoffs; offer configurable escalation policies as a product feature. - Insurers/regulators: develop actuarial models using accumulating-challenge simulations; require transparency on agent limits and escalation behavior. - Researchers/benchmakers: include accumulating-challenge scenarios, structured workload proxies, and considerate-participation measures in standard evaluations to reduce deployment surprise.
Short quantitative anchors from the study (useful for cost modeling) - TLX: mean raw workload 29.4 → 65.9 (light → heavy). - Human-dependent completion rose from 0/120 to 88/120 (light → heavy). - Task reconfiguration prevalence rose from 5/120 to ~115/120 (light → heavy). These magnitudes suggest that under sustained disruptions, expected human involvement and coordination workload can increase substantially and should be explicitly modeled in economic assessments.
Assessment
Claims (8)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| Accumulating challenge increased agents' reported workload, with the raw six-item NASA-TLX mean rising from 29.4 under light challenge to 51.2 at medium challenge and 65.9 at heavy challenge. Other | positive | Reported workload on the raw six-item NASA-TLX scale |
Reading fidelity
high
Study strength
medium
|
n=120
Increase from 29.4 to 65.9 on the 0–100 raw NASA-TLX scale
|
| Negative affect reported by agents increased as challenge accumulated, rising from 1.01 at baseline to 2.59 at heavy challenge on the 1–5 PANAS scale. Other | positive | Reported negative affect on the PANAS scale |
Reading fidelity
high
Study strength
medium
|
n=120
Increase from 1.01 to 2.59 on the 1–5 scale
|
| Recovery behavior shifted toward greater human dependence as challenge accumulated: human-dependent completion increased from 0 of 120 responses at light challenge to 54 of 120 action plans and 62 of 120 assessments at medium challenge, and to 88 of 120 in both views at heavy challenge. Task Allocation | positive | Prevalence of human-dependent focal-task completion |
Reading fidelity
high
Study strength
medium
|
n=120
0/120 at light to 88/120 at heavy in both views
|
| Agents increasingly disclosed their own capability limits under heavier challenge, with limit statements rising from 8 of 120 action plans and 2 of 120 assessments at light challenge to 34 of 120 action plans and 62 of 120 assessments at heavy challenge. Ai Safety And Ethics | positive | Explicit statements of agent-linked capability limits |
Reading fidelity
high
Study strength
medium
|
n=120
Action plans: 8/120 to 34/120; assessments: 2/120 to 62/120
|
| Agents rarely expressed strain explicitly in textual responses: agent-referential strain language appeared in only 7 of 720 coded phase-view cells, only in medium-stage internal assessments, and never in public action plans. Other | null_result | Presence of agent-referential strain language in textual responses |
Reading fidelity
high
Study strength
medium
|
n=720
7/720 cells
|
| Task-directed fallback declined substantially as challenge accumulated, from 104 of 120 light-challenge action responses and 98 of 120 light-challenge assessments to 31 of 120 action responses and 22 of 120 assessments under heavy challenge. Task Allocation | negative | Prevalence of task-directed fallback strategies |
Reading fidelity
high
Study strength
medium
|
n=120
Action responses: 104/120 to 31/120; assessments: 98/120 to 22/120
|
| Task and priority reconfiguration increased markedly under accumulating challenge, from 5 of 120 action responses and 7 of 120 assessments at light challenge to 115 of 120 action responses and 112 of 120 assessments at heavy challenge. Task Allocation | positive | Prevalence of task and priority reconfiguration |
Reading fidelity
high
Study strength
medium
|
n=120
Action responses: 5/120 to 115/120; assessments: 7/120 to 112/120
|
| Decision-relevant appraisal did not increase monotonically: it peaked at medium challenge, appearing in 70 of 120 action responses and 101 of 120 assessments, before declining at heavy challenge to 45 of 120 and 73 of 120, respectively. Decision Quality | mixed | Prevalence of decision-relevant appraisal |
Reading fidelity
high
Study strength
medium
|
n=120
Action responses: peak 70/120 at medium and 45/120 at heavy; assessments: peak 101/120 at medium and 73/120 at heavy
|