The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Microsoft 365 Copilot is seen as reliable and easy to use, with administrative employees reporting highest immediate gains and scientists growing more positive over time; benefits concentrate on clearly structured, text-based knowledge work and suggest learning and routinization effects that call for role-specific training and governance.

Generative AI in Knowledge Work: Perception, Usefulness, and Acceptance of Microsoft 365 Copilot
Carsten F. Schmidt, Sophie Petzolt, Wolfgang Beinhauer, Ingo Weber, Stefan Langer · February 20, 2026
arxiv correlational low evidence 7/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

Arxiv

Latest observation:

  1. Carsten F. Schmidt unresolved corpus identity
  2. Sophie Petzolt unresolved corpus identity
  3. Wolfgang Beinhauer unresolved corpus identity
  4. Ingo Weber unresolved corpus identity
  5. Stefan Langer unresolved corpus identity

Semantic Scholar

Latest observation:

  1. Carsten Schmidt provider ID
  2. Sophie Petzolt provider ID
  3. W. Beinhauer provider ID
  4. Ingo Weber provider ID
  5. S. Langer provider ID
In a single research organization, Microsoft 365 Copilot was perceived as user-friendly and reliable, with administrative staff reporting immediate usefulness while scientific staff showed increasing positive assessments over time, especially for productivity and workload reduction in structured text tasks.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

The study analyzes the introduction of Microsoft 365 Copilot in a non-university research organization using a repeated cross-sectional employee survey. We assess usefulness, ease of use, output quality and reliability, and usefulness for typical knowledge-work activities. Administrative staff report higher usefulness and reliability, whereas scientific staff develop more positive assessments over time, especially regarding productivity and workload reduction. Copilot is widely viewed as user-friendly and technically reliable, with greatest added value for clearly structured, text-based tasks. The findings highlight learning and routinization effects when embedding generative AI into work processes and stress the need for context-sensitive implementation, role-specific training and governance to foster sustainable acceptance of generative AI in knowledge-intensive organizations.

Summary

Main Finding

In a pilot rollout of Microsoft 365 Copilot at a large non-university research organization, employees generally judged the tool as user-friendly, technically reliable, and useful—with the greatest practical gains for clearly structured, text-based tasks. Administrative staff reported higher usefulness and perceived output quality early on, while scientific staff showed significant positive shifts over time (notably in perceived usefulness, ease of use, productivity and workload reduction), consistent with learning and routinization effects. The study highlights the importance of role-specific implementation, training, and governance for sustainable acceptance of generative AI in knowledge-intensive organizations.

Key Points

  • Sample and design

    • Pilot in a large non-university research organization (quota-based selection of license holders).
    • Two repeated cross-sectional survey waves: T01 (Nov–Dec 2024), N = 106 (66 science, 40 admin); T02 (Mar–Apr 2025), N = 90 (51 science, 39 admin).
    • Participants were pseudonymized; overlap across waves was small, so analyses are repeated cross-sections rather than longitudinal individual trajectories.
  • Measures

    • Acceptance constructs drawn from TAM: perceived usefulness (PU, 5 items, α = .97/.95), perceived ease of use (2 items), output quality, reliability, voluntariness; all items on a −3 to +3 Likert scale.
    • Nine task-specific usefulness items (analyzed descriptively).
  • Key quantitative findings (selected)

    • Scientific staff PU increased from 0.42 → 1.09 (d = 0.42, p = .022). Perceived ease of use for scientific staff increased from 0.74 → 1.31 (d = 0.45, p = .016).
    • Administrative staff had higher initial PU (0.94 at T01) and higher output-quality ratings at T01 (0.88 vs 0.26 for science; group d = 0.48, p = .018). Differences largely narrowed by T02.
    • Perceived technical reliability was positive and stable in both groups (admin ≈ 1.13 → 1.15; science ≈ 0.96 → 0.94).
    • Respondents reported low anticipated harm from incorrect outputs (means ≈ −2.3), and high voluntariness of use (means ≈ 2.6–2.8).
    • Task-level pattern: greatest added value for clearly structured, text-based tasks (e.g., drafting, summarization); less perceived benefit for tasks requiring deep domain expertise, experimental design, or compliance/legal judgement.
  • Contextual dynamics

    • Eight dated updates to Copilot occurred during the study period (content and technical improvements), which likely affected perceptions over time.
    • The study is exploratory and descriptive; multiple comparisons for task items were not inferentially tested.

Data & Methods

  • Recruitment and sample

    • 550 employees were provisioned with Copilot licenses as part of a technical pilot and invited to participate; final analysis samples: 106 (T01) and 90 (T02).
    • Quota nomination by organizational units aimed to reflect institute affiliation, age, gender, role, and prior AI experience; not a random or representative sample.
  • Survey instrument

    • Online standardized questionnaire (LimeSurvey), items mainly adapted from TAM (Davis, 1989) and related work.
    • Constructs: perceived usefulness (5 items), perceived ease of use (2 items), output quality (1), reliability (1), harm from incorrect output (1), voluntariness (1), plus nine task-specific usefulness items.
    • Responses coded −3 (“strongly disagree”) to +3 (“strongly agree”).
  • Analysis

    • Descriptive statistics by group and wave; group/time comparisons via Welch t-tests (unequal variances) and Cohen’s d effect sizes.
    • Task-specific usefulness analyzed descriptively to avoid multiple-testing inflation.
    • Missing-data rule: exclude cases with >20% missing; listwise exclusion otherwise. Implausible response patterns removed.
    • Limitations explicitly acknowledged: small and non-random sample, self-reported measures, limited power for population inference, repeated cross-sectional design prevents individual-level change claims, potential selection bias toward AI-interested users.

Implications for AI Economics

  • Heterogeneous returns by role and task

    • Immediate productivity and quality gains appear concentrated in administrative, routine, structured text tasks. For organizations, this implies a faster, clearer ROI from deploying generative-AI tools in back-office and standardized-document workflows than across all knowledge-work functions.
    • Scientific/creative knowledge work shows growing perceived benefits over time—suggesting that complementary investments (learning, prompts/usage practices, integration into workflows) matter: adoption generates increasing returns as users routinize AI-assisted processes.
  • Human capital and training investments

    • The rise in perceived usefulness and ease of use among scientists points to the importance of role-specific training and knowledge-sharing. Economic models of AI adoption should factor in dynamic complementarities: short-run heterogeneity but potential medium-run convergence as firms/train staff.
  • Organizational implementation and governance costs

    • Sensitive legal/compliance contexts (data protection, IP, patents) in research organizations mean governance, auditing, and role-based access controls are economically relevant costs. Effective implementation will require protocols, monitoring, and possibly human-in-the-loop verification—reducing pure automation gains but preserving value through safer deployment.
  • Labor-market effects and task reallocation

    • The pattern is more consistent with task reallocation and augmentation than immediate displacement: routine, structured tasks are augmented (efficiency gains), while domain-expert tasks remain reliant on human oversight. This supports nuanced projections of labor demand changes: decreased time on routinized tasks, increased emphasis on supervision, prompt-engineering, quality assurance, and higher-order tasks.
  • Pricing, licensing, and ROI considerations

    • Given concentrated benefits, organizations should prioritize licencing and deployment where marginal productivity gains are highest (administration, standardized reporting). Cost–benefit analyses must include training, governance, and change management, not just license fees.
  • Policy and research implications

    • Policy to accompany AI diffusion should emphasize upskilling, support for role-differentiated adoption, and funding for governance/verification infrastructure in mission-critical institutions.
    • Research agenda: complement acceptance surveys with objective productivity measures, longitudinal within-subject designs, and task-level decomposition to quantify real output gains, error costs, and net labor-hour effects.

Caveats to keep in mind - Results are exploratory and based on self-reports from a non-random, small pilot sample during a period of rapid product updates; findings should be treated as indicative hypotheses rather than definitive causal estimates. Future work should combine experimental or observational productivity data, representative sampling, and longer follow-up to quantify economic impacts precisely.

Assessment

Paper Typecorrelational Evidence Strengthlow — Findings are based on repeated cross-sectional self-reported survey data from a single organization without random assignment or a clear counterfactual, so causal claims about Copilot's effects on productivity or workload are weak. Methods Rigormedium — The study uses systematic repeated cross-sectional measurement and disaggregates results by role (administrative vs scientific), allowing assessment of temporal trends and heterogeneity; however, it relies on subjective measures, lacks objective productivity metrics, potential response and selection biases are not addressed, and it does not track the same individuals longitudinally. SampleEmployees of a single non-university research organization surveyed in multiple waves around the introduction of Microsoft 365 Copilot; respondents include administrative staff and scientific staff (role-based comparisons reported); exact sample sizes, response rates, and survey timing not specified in the summary. Themeshuman_ai_collab productivity org_design GeneralizabilitySingle-organization sample limits external validity to other institutions or sectors, Findings are specific to Microsoft 365 Copilot and may not generalize to other generative-AI systems, Focus on knowledge-intensive, research-oriented work may not apply to routine or field-based jobs, Results rely on self-reported perceptions rather than objective productivity or output measures, Repeated cross-sectional design prevents tracking individual-level learning trajectories, Organizational culture, digital maturity, and regional context may affect transferability

Claims (6)

ClaimDirectionOutcomeConfidence & EvidenceDetails
Administrative staff report higher usefulness and reliability of Microsoft 365 Copilot than scientific staff. Worker Satisfaction positive perceived usefulness and technical reliability
Reading fidelity high
Study strength medium
not reported
0.3
Scientific staff develop more positive assessments of Copilot over time, particularly regarding productivity gains and reductions in workload. Organizational Efficiency positive perceived productivity and workload reduction
Reading fidelity high
Study strength medium
not reported
0.3
Copilot is widely viewed by employees as user-friendly (easy to use) and technically reliable. Worker Satisfaction positive ease of use and perceived technical reliability
Reading fidelity high
Study strength medium
not reported
0.3
Copilot's greatest added value is for clearly structured, text-based tasks. Task Allocation positive perceived added value by task type (structured text-based tasks)
Reading fidelity high
Study strength medium
not reported
0.3
There are learning and routinization effects when embedding generative AI (Copilot) into work processes. Skill Acquisition positive changes in attitudes and reported use over time (learning/routinization)
Reading fidelity high
Study strength medium
not reported
0.3
Context-sensitive implementation, role-specific training, and governance are needed to foster sustainable acceptance of generative AI in knowledge-intensive organizations. Governance And Regulation positive sustainable acceptance of generative AI (organizational adoption and acceptance)
Reading fidelity high
Study strength speculative
not reported
0.05

Notes