The Commonplace
Home Three-study pilot Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Large language models can speed multilingual coordination in global projects but often mirror dominant cultural norms and lack robust evaluation of culturally appropriate communication, risking pragmatic misalignment in cross‑cultural teams.

Literature review on large language models (LLMs) for cross-cultural project management
Yutong Shen, Dilek Cetindamar, Yi Zhang · September 05, 2026 · AI & Society
openalex review_meta n/a evidence 7/10 relevance Full text usable extracted full text DOI Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

OpenAlex

Latest observation:

  1. Yutong Shen provider ID
  2. Dilek Cetindamar provider ID
  3. Yi Zhang provider ID
A scoping review of 69 studies finds LLMs can improve certain efficiency and coordination outcomes in cross-cultural project teams but tend to reproduce dominant (often Western) cultural norms and are evaluated mostly with semantic/task metrics that poorly capture cultural–pragmatic appropriateness.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

Abstract As large language models (LLMs) become increasingly embedded in global project environments, they are reshaping how people execute tasks, communicate, and collaborate. Although multilingual advances and established translation and interpreting practices can reduce linguistic barriers, semantic accuracy alone does not ensure culturally or pragmatically appropriate interaction in cross-cultural project teams. Misalignment may still arise from implicit norms, values, and context-sensitive expectations that current LLMs do not consistently represent. Using a scoping review methodology, this study synthesises 69 interdisciplinary studies and examines what is currently known about the capacity of culturally adaptive LLMs to support communication processes and collaboration outcomes in cross-cultural project management contexts. The findings are organised into four thematic domains: human–AI collaboration, cross-cultural project communication, LLM cultural and multilingual optimisation strategies, and challenges in evaluating cultural–pragmatic performance. The review indicates that LLMs can contribute to selected efficiency and coordination outcomes, but may also reproduce dominant cultural norms. Moreover, the reviewed technical studies predominantly use semantic or task-oriented metrics, which provide limited evidence about pragmatic alignment and relational effects. The article identifies research gaps in AI-mediated communication, cultural evaluation, and governance-oriented assessment and situates culturally adaptive LLMs within an established socio-technical debate about responsible and inclusive global collaboration.

Summary

Main Finding

LLMs can improve selected efficiency and coordination outcomes in cross-cultural project teams (e.g., drafting, translation, routine coordination), but they also risk reproducing dominant (often Western-centred) cultural norms and producing pragmatically inappropriate outputs when deployed without human contextualisation. Existing evidence is concentrated on semantic/task-oriented metrics, leaving major gaps in understanding LLMs’ effects on cultural‑pragmatic alignment, relational trust, and governance in global projects.

Key Points

  • Four thematic domains synthesised:
    • A: Human–AI collaboration in project management — LLMs as assistants for drafting, summarisation, translation, and routine coordination tasks.
    • B: Cross-cultural communication & cultural intelligence (CQ) — project outcomes depend on culturally sensitive timing, tone, role boundaries and trust; CQ is a useful conceptual lens for what “culturally adaptive” should mean.
    • C: LLM cultural/multilingual optimisation strategies — approaches include multilingual pretraining, fine‑tuning, prompting strategies, localized data augmentation, and human‑in‑the‑loop adaptation; many models still biased toward English/Western corpora.
    • D: Gaps in measuring cultural–pragmatic performance — most technical evaluations use semantic similarity, BLEU/ROUGE-type metrics, or task accuracy; few measure pragmatic appropriateness, relational effects, or culturally contingent misalignment.
  • Risk profile:
    • Semantic adequacy ≠ pragmatic appropriateness: outputs can be fluent yet culturally inappropriate (e.g., tone misfit, misinterpreted politeness or status cues).
    • Reproduction of dominant norms: training data imbalance leads to systematic cultural bias in outputs.
    • Institutional risks when teams rely on LLM outputs without professional mediation or accountability.
  • Methodological and scope notes:
    • Evidence is interdisciplinary but dispersed; technical papers, PM literature, translation/interpreting studies and CQ research have developed largely in parallel.
    • The review highlights the need to integrate translation/interpreting expertise and CQ considerations into LLM deployment in projects.

Data & Methods

  • Review type: Scoping review following Arksey & O’Malley framework, PRISMA-ScR and JBI recommendations; thematic analysis with four a priori domains.
  • Corpus: 69 interdisciplinary studies synthesised.
  • Sources searched: Web of Science, Scopus, ABI/INFORM, EBSCO Business Source, plus targeted screening of ACL/NeurIPS proceedings and relevant arXiv preprints.
  • Timeframe of search: October 2025 – January 2026; no publication-year restriction.
  • Inclusion criteria: English-language publications with sufficient methodological detail; focus on cross-cultural project communication, LLMs, CQ and evaluation.
  • Coding/synthesis: Deductive thematic scaffold (A–D) applied; Zotero for reference management and Excel to document screening and coding.
  • Limitations acknowledged by authors:
    • English‑only search may bias thematic distribution and underrepresent locally published/non‑English evidence.
    • Predefined domains narrowed scope and possibly excluded adjacent, relevant work.
    • The heterogeneous nature of studies precluded meta‑analytic effect estimates.

Implications for AI Economics

  • Productivity vs. relational costs: Economic evaluations of LLM adoption in international projects must include not only task-time savings (drafting, translation) but also potential negative externalities — increased misunderstandings, reduced trust, coordination breakdowns — that can impose transaction and rework costs.
  • Labour-market impacts:
    • Demand shift rather than simple displacement: greater reliance on LLMs for first‑pass work may reduce certain routine translation/drafting tasks but increase demand for culturally specialised reviewers, localisers, and human-in-the-loop QA.
    • Valuation of cultural expertise: translation/interpreting professionals and cross-cultural project managers perform economic value that LLMs do not fully substitute; firms should price and budget for that expertise when deploying LLMs.
  • Investment priorities:
    • Data diversity & localisation: investing in culturally representative training/fine‑tuning data and regionally targeted evaluation yields higher expected returns by reducing cultural-misalignment risks.
    • Human-in-the-loop governance: allocating resources to review, escalation paths, and role definitions (who accepts LLM suggestions) can materially change realized benefits.
  • Measurement & metrics for cost–benefit analysis:
    • Current LLM metrics (BLEU, ROUGE, semantic similarity) are inadequate for economic modelling of cross-cultural deployments. New metrics should quantify pragmatic alignment, impact on trust, frequency/severity of miscommunication, and downstream rework costs.
    • Firms and researchers should run RCTs or quasi‑experimental evaluations comparing project outcomes (time-to-complete, budget variance, stakeholder satisfaction, rework incidence) with and without LLM mediation.
  • Governance & market regulation:
    • Standards and auditing for cultural/pragmatic performance (including documentation of cultural coverage, provenance of localized data) will shape market trust and adoption rates — potential regulatory or procurement requirements could affect costs of adopting LLMs in global projects.
  • Research directions relevant to AI economics:
    • Develop empirical estimates of productivity gains net of cultural-friction costs across sectors and languages.
    • Model complementary labor reallocation (how much human cultural-review capacity is needed per unit of LLM-assisted output).
    • Design pricing models and procurement contracts that internalise cultural‑misalignment risk (e.g., warranties, liability allocations, mandated human sign‑off for high‑stakes communications).
    • Create pragmatic evaluation metrics that can be used to predict economic impacts and incorporate them into ROI frameworks.

Suggested short actionable steps for economic stakeholders: - When piloting LLMs in cross-border projects, measure both efficiency (time/cost savings) and relational outcomes (stakeholder satisfaction, rework). - Budget explicitly for human cultural review and include CQ expertise in project teams. - Prioritise investments in localized training data and evaluation protocols before full deployment. - Collaborate with translators/interpreters to define acceptable pragmatic thresholds; incorporate those thresholds in procurement and auditing.

If you want, I can draft a short economic model outline (inputs, parameters, and equations) to estimate net benefits of deploying LLMs in a cross-cultural project setting that includes cultural‑misalignment risk.

Assessment

Paper Typereview_meta Evidence Strengthn/a — This is a scoping literature review synthesising existing studies rather than providing new causal identification; it summarizes heterogeneous empirical and technical work but does not itself generate causal evidence. Methods Rigormedium — The review follows established scoping-review frameworks (Arksey & O'Malley, PRISMA-ScR, JBI), uses multiple databases, supplementary conference/preprint screening, documented search/screening, and deductive thematic coding; however it is limited by English-only selection, predefined thematic domains that may bias inclusion/coding, and excluding works lacking methodological detail which may omit relevant practice-oriented evidence. SampleScoping review of 69 interdisciplinary studies identified via searches (Oct 2025–Jan 2026) across Web of Science, Scopus, ABI/INFORM, EBSCO Business Source, plus targeted ACL/NeurIPS and arXiv screening; English-language publications only; studies retained required sufficient methodological detail; coded into four deductive domains (human–AI collaboration; cross-cultural project communication; LLM cultural/multilingual optimisation; cultural–pragmatic measurement). Themeshuman_ai_collab adoption org_design productivity GeneralizabilityEnglish-language only — may underrepresent non-English/local scholarship and culturally specific practices, Focused on project-management contexts — limited applicability to other organisational settings (e.g., customer service, education), Deductive thematic framework may have constrained discovery of emergent themes outside predefined domains, Exclusion of studies lacking methodological detail (e.g., some preprints, practitioner reports) may bias toward academically reported methods and metrics, Search window up to Jan 2026 — rapidly evolving LLM work after that date not captured

Claims (8)

ClaimDirectionOutcomeConfidence & EvidenceDetails
The scoping review finds that LLMs can contribute to selected efficiency and coordination outcomes in cross-cultural project management. Organizational Efficiency positive Efficiency and coordination in cross-cultural project work
Reading fidelity high
Study strength low
n=69
0.12
The reviewed literature indicates that LLMs may reproduce dominant cultural norms, particularly because many models are trained disproportionately on English-language and Western-centred corpora. Ai Safety And Ethics negative Representation of cultural norms in LLM outputs
Reading fidelity high
Study strength low
n=69
0.12
Semantic accuracy or multilingual capability alone does not ensure culturally or pragmatically appropriate interaction in cross-cultural project teams. Output Quality negative Cultural–pragmatic appropriateness of communication
Reading fidelity high
Study strength medium
n=69
0.24
The technical studies reviewed predominantly evaluate LLMs using semantic or task-oriented metrics, providing limited evidence about pragmatic alignment and relational effects. Ai Safety And Ethics negative Coverage and validity of cultural–pragmatic evaluation
Reading fidelity high
Study strength medium
n=69
0.24
LLM-mediated communication can generate semantically plausible but pragmatically inappropriate advice or communication when outputs are used without appropriate human review or contextualisation. Error Rate negative Pragmatic appropriateness and reliability of project communication
Reading fidelity high
Study strength medium
n=69
0.24
High performance on semantic or benchmark-based tasks does not necessarily indicate sensitivity to context, pragmatics, or cultural variation. Decision Quality mixed Relationship between benchmark performance and culturally appropriate communication
Reading fidelity high
Study strength medium
n=69
0.24
The review identifies substantial research gaps in AI-mediated communication, cultural evaluation, and governance-oriented assessment of culturally adaptive LLMs. Governance And Regulation negative Availability and adequacy of evidence for cultural evaluation and governance
Reading fidelity high
Study strength low
n=69
0.12
The review was based on 69 interdisciplinary studies identified through a scoping-review process covering Web of Science, Scopus, ABI/INFORM, and EBSCO Business Source, supplemented by selected ACL, NeurIPS, and arXiv materials. Other null_result Scope and composition of the evidence base
Reading fidelity high
Study strength medium
n=69
0.24

Notes