The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

A new collaboration benchmark shows humans and AI replicate some human–human coordination patterns but diverge sharply on establishing common ground and repairing misunderstandings. The lab study validates the task’s theoretical grounding but flags limited generality: human-AI teams break down in predictable ways that could constrain real-world productivity gains.

A Benchmark to Assess Common Ground in Human-AI Collaboration
Christian Poelitz, Finale Doshi-Velez, Siân Lindley · February 24, 2026
arxiv quasi_experimental medium evidence 7/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

Arxiv

Latest observation:

  1. Christian Poelitz unresolved corpus identity
  2. Finale Doshi-Velez unresolved corpus identity
  3. Siân Lindley unresolved corpus identity

Semantic Scholar

Latest observation:

  1. Christian Poelitz provider ID
  2. F. Doshi-Velez provider ID
  3. Siân Lindley provider ID
The paper introduces a theory-grounded collaborative-puzzle benchmark and a confirmatory user study showing it reproduces human–human coordination patterns while also revealing systematic divergences in how humans establish and repair common ground with an AI partner.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

AI is becoming increasingly integrated into everyday life, both in professional work environments and in leisure and entertainment contexts. This integration requires AI to move beyond acting as an assistant for informational or transactional tasks toward a genuine collaborative partner. Effective collaboration, whether between humans or between humans and AI, depends on establishing and maintaining common ground: shared beliefs, assumptions, goals, and situational awareness that enable coordinated action and efficient repair of misunderstandings. While common ground is a central concept in human collaboration, it has received limited attention in studies of human-AI collaboration. In this paper, we introduce a new benchmark grounded in theories and empirical studies of human-human collaboration. The benchmark is based on a collaborative puzzle task that requires iterative interaction, joint action, referential coordination, and repair under varying conditions of situation awareness. We validate the benchmark through a confirmatory user study in which human participants collaborate with an AI to solve the task. The results show that the benchmark reproduces established theoretical and empirical findings from human-human collaboration, while also revealing clear divergences in human-AI interaction.

Summary

Main Finding

The paper introduces and validates a task-based benchmark for measuring whether and how human–AI pairs develop common ground during collaboration. In a confirmatory online study with one fixed model (GPT-4.1), the benchmark reproduces several known patterns from human–human grounding research (e.g., development of shared referential language and efficiency gains) while also revealing systematic divergences in human–AI interaction — notably AI tendencies toward single-turn or overly long replies, lower rates of clarification-seeking, asymmetric grounding effort (humans bearing most repair work), and superficial grounding cues that do not reliably indicate jointly usable understanding.

Key Points

  • Motivation: As AI shifts from one-shot assistants to partners in long-horizon, joint tasks, common ground (shared beliefs, assumptions, goals, situational awareness) becomes essential for coordinated action and efficient repair of misunderstandings.
  • Theoretical foundation: Builds on decades of work in grounding and collaborative communication (Clark & Brennan; Krauss; Clark & Wilkes‑Gibbs), and on human–AI literature highlighting failures in shared mental models and situational understanding.
  • Benchmark task: A collaborative puzzle requiring iterative dialogue, joint actions, referential coordination, and repair under controlled variations in situation awareness (e.g., shared vs. partial visibility of task state).
  • Measured outcomes: Situation awareness, learning effects across repeated interactions, communication efficiency (e.g., message length and turn counts), grounding behaviors (presentation, acceptance, clarification, repair), and task performance.
  • Validation study: An online confirmatory experiment where human participants solved the benchmark task with GPT-4.1 as the partner. Results show (a) emergence of some human–human-like grounding patterns, and (b) clear human–AI mismatches in grounding dynamics and burden of repair.
  • Diagnostic value: The benchmark surfaces interaction patterns known to hinder grounding (single-turn responses, overly long answers, asymmetric grounding effort, shallow surface cues) and thus can evaluate whether an AI system supports genuine shared understanding rather than merely producing language artifacts that look like grounding.

Data & Methods

  • Design rationale: The benchmark is grounded in empirical findings and theory from human–human collaboration. It intentionally requires joint problem-space construction and coordination so that successful performance depends on genuinely shared task representations, not just isolated responses.
  • Task structure: Iterative, multi-turn conversational problem solving with referential elements (objects, states) that must be jointly tracked and updated; includes opportunities for clarification and repair; situation-awareness manipulations (e.g., different visibility or reviewability of task state) to vary grounding costs.
  • Outcome measures (reported/described): task completion and accuracy, time/turns to completion, textual/linguistic metrics (message length, lexical persistence), counts/types of grounding acts (clarification requests, acknowledgments, repairs), measures of situation awareness and learning across trials.
  • Validation sample and model: Confirmatory online user study pairing human participants with GPT-4.1. (The paper reports replication of known empirical grounding patterns and documents qualitative and quantitative divergences; exact sample size and numeric results are reported in the full text.)
  • Contrast with prior methods: Unlike many assistant-style evaluations (one-shot tasks, static benchmarks), this benchmark emphasizes dynamic co-construction and repair, enabling assessment of whether apparent grounding corresponds to mutually actionable shared state.

Implications for AI Economics

  • Productivity and coordination costs: If AI partners fail to establish durable common ground, human workers may incur substantial additional coordination and repair costs. Economists modeling AI augmentation should account not only for task-level automation gains but also for increased overhead from asymmetric grounding (time spent clarifying, correcting, or supervising AI).
  • Valuation of AI as partner vs. assistant: The benchmark provides operational metrics (turns-to-completion, human repair time, error propagation) that can be used to quantify the marginal productivity of AI systems in collaborative tasks. This helps distinguish complementary (augments human capability) from substitutive (replaces human effort) outcomes in firm-level and labor-market analyses.
  • Design and procurement incentives: Organizations buying or deploying AI for collaborative work should prioritize systems optimized for mutual grounding (e.g., proactive clarification, incremental proposals, shared workspace integration). Procurement frameworks and contracts might tie vendor evaluation to benchmarked reductions in human repair burden and improvements in joint task throughput.
  • Training objectives and model evaluation: From an economic perspective, training objectives that internalize grounding — e.g., incentivizing clarification, shorter iterative turns, and explicit shared-state confirmations — may yield higher downstream labor productivity and lower oversight costs than optimizing only for one-shot accuracy metrics.
  • Externalities and allocation of cognitive labor: The asymmetry observed (humans carrying the burden of repair) can concentrate cognitive strain on certain workers (e.g., managers, domain experts). Economic models of task allocation should incorporate such distributional effects and potential changes in wage premia for roles that require persistent AI supervision or corrective labor.
  • Research priorities for policy and ROI: Policymakers and firms should fund and measure improvements in AI grounding capabilities because small percentage improvements in repair/coordination costs could yield large aggregate productivity gains across many collaborative tasks—especially in high-stakes decision-making, engineering, healthcare, and creative workflows.
  • Empirical economic metrics to adopt: Use benchmark outputs to construct economic indicators such as human time per task saved (or lost), error-correction labor hours, changes in throughput, learning curves (reduction in human repair over repeated interactions), and incidence of catastrophic miscoordination. These can feed cost–benefit and adoption models.

Practical recommendations (brief) - For AI developers: Optimize for interactive, incremental exchanges (encourage clarification, confirmations, concise turns) and integrate shared workspaces / state visualization to lower grounding costs. - For firms/users: Adopt benchmarks like this when selecting AI tools for collaborative roles; measure human repair time and task throughput in pilots; train staff in effective prompting and repair strategies. - For economists and policy analysts: Incorporate grounding-related coordination costs into models of AI-driven productivity change and labor reallocation; use benchmark-derived metrics to estimate real-world complementarities and supervision burdens.

Assessment

Paper Typequasi_experimental Evidence Strengthmedium — The paper validates a theory-grounded benchmark with an empirical user study that reproduces established human–human collaboration findings, providing constructive internal evidence; however, the evidence is limited by likely small/controlled samples, a single task paradigm, and dependence on a particular AI agent, reducing external and ecological validity for broader claims about AI's economic impacts. Methods Rigormedium — The study is theory-driven and confirmatory, using experimental manipulations of situation-awareness and measures of iterative interaction and repair, which suggests careful design; but the summary lacks details on sample size, randomization, pre-registration, statistical power, model robustness checks, and convergence across multiple AI systems, so potential methodological limitations remain. SampleA controlled user study where human participants collaborated with an AI agent on a purpose-built collaborative puzzle task that requires iterative interaction, referential coordination, joint action, and repair; the AI system, participant count, recruitment source, and demographic composition are not specified in the summary. Themeshuman_ai_collab productivity IdentificationControlled manipulation of situation-awareness conditions in a confirmatory user study: participants performed a collaborative puzzle with an AI under varied task/awareness conditions and outcomes were compared across these conditions to attribute differences to the experimental manipulations; causal claims are therefore local to the lab task and the specific AI agent used (randomization, pre-registration, and robustness details not provided in the summary). GeneralizabilitySingle, artificial puzzle task — may not generalize to diverse real-world work tasks or domains., Findings tied to the specific AI agent(s) used — other models/implementations may behave differently., Likely short-term laboratory interactions — unclear if effects persist in long-running or repeated collaborations., Participant pool and cultural context unspecified — limits applicability across populations and workplace settings., Does not measure macroeconomic outcomes (productivity, wages, firm performance) — limits relevance to economy-wide claims.

Claims (9)

ClaimDirectionOutcomeConfidence & EvidenceDetails
AI is becoming increasingly integrated into everyday life, both in professional work environments and in leisure and entertainment contexts. Adoption Rate positive degree of AI integration into everyday life
Reading fidelity high
Study strength medium
not reported
0.48
This integration requires AI to move beyond acting as an assistant for informational or transactional tasks toward a genuine collaborative partner. Adoption Rate positive need for AI capabilities enabling collaboration
Reading fidelity high
Study strength speculative
not reported
0.08
Effective collaboration, whether between humans or between humans and AI, depends on establishing and maintaining common ground: shared beliefs, assumptions, goals, and situational awareness that enable coordinated action and efficient repair of misunderstandings. Team Performance positive quality of collaboration (coordination, repair of misunderstandings)
Reading fidelity high
Study strength high
not reported
0.8
While common ground is a central concept in human collaboration, it has received limited attention in studies of human-AI collaboration. Research Productivity negative research attention to common ground in human-AI studies
Reading fidelity high
Study strength medium
not reported
0.48
We introduce a new benchmark grounded in theories and empirical studies of human-human collaboration. Research Productivity positive availability of a theory-grounded benchmark for human-AI collaboration
Reading fidelity high
Study strength medium
not reported
0.48
The benchmark is based on a collaborative puzzle task that requires iterative interaction, joint action, referential coordination, and repair under varying conditions of situation awareness. Task Allocation positive task design capturing iterative interaction, joint action, referential coordination, repair, and situation awareness
Reading fidelity high
Study strength medium
not reported
0.48
We validate the benchmark through a confirmatory user study in which human participants collaborate with an AI to solve the task. Research Productivity positive benchmark validity via human-AI user study
Reading fidelity high
Study strength medium
not reported
0.48
The results show that the benchmark reproduces established theoretical and empirical findings from human-human collaboration. Team Performance positive similarity of human-AI collaboration patterns to established human-human collaboration findings
Reading fidelity high
Study strength medium
not reported
0.48
The results also reveal clear divergences in human-AI interaction. Team Performance mixed differences/divergences in interaction patterns between human-AI and human-human teams
Reading fidelity high
Study strength medium
not reported
0.48

Notes