1 cumulative citations
View corpus contextA new collaboration benchmark shows humans and AI replicate some human–human coordination patterns but diverge sharply on establishing common ground and repairing misunderstandings. The lab study validates the task’s theoretical grounding but flags limited generality: human-AI teams break down in predictable ways that could constrain real-world productivity gains.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
AI is becoming increasingly integrated into everyday life, both in professional work environments and in leisure and entertainment contexts. This integration requires AI to move beyond acting as an assistant for informational or transactional tasks toward a genuine collaborative partner. Effective collaboration, whether between humans or between humans and AI, depends on establishing and maintaining common ground: shared beliefs, assumptions, goals, and situational awareness that enable coordinated action and efficient repair of misunderstandings. While common ground is a central concept in human collaboration, it has received limited attention in studies of human-AI collaboration. In this paper, we introduce a new benchmark grounded in theories and empirical studies of human-human collaboration. The benchmark is based on a collaborative puzzle task that requires iterative interaction, joint action, referential coordination, and repair under varying conditions of situation awareness. We validate the benchmark through a confirmatory user study in which human participants collaborate with an AI to solve the task. The results show that the benchmark reproduces established theoretical and empirical findings from human-human collaboration, while also revealing clear divergences in human-AI interaction.
Summary
Main Finding
The paper introduces and validates a task-based benchmark for measuring whether and how human–AI pairs develop common ground during collaboration. In a confirmatory online study with one fixed model (GPT-4.1), the benchmark reproduces several known patterns from human–human grounding research (e.g., development of shared referential language and efficiency gains) while also revealing systematic divergences in human–AI interaction — notably AI tendencies toward single-turn or overly long replies, lower rates of clarification-seeking, asymmetric grounding effort (humans bearing most repair work), and superficial grounding cues that do not reliably indicate jointly usable understanding.
Key Points
- Motivation: As AI shifts from one-shot assistants to partners in long-horizon, joint tasks, common ground (shared beliefs, assumptions, goals, situational awareness) becomes essential for coordinated action and efficient repair of misunderstandings.
- Theoretical foundation: Builds on decades of work in grounding and collaborative communication (Clark & Brennan; Krauss; Clark & Wilkes‑Gibbs), and on human–AI literature highlighting failures in shared mental models and situational understanding.
- Benchmark task: A collaborative puzzle requiring iterative dialogue, joint actions, referential coordination, and repair under controlled variations in situation awareness (e.g., shared vs. partial visibility of task state).
- Measured outcomes: Situation awareness, learning effects across repeated interactions, communication efficiency (e.g., message length and turn counts), grounding behaviors (presentation, acceptance, clarification, repair), and task performance.
- Validation study: An online confirmatory experiment where human participants solved the benchmark task with GPT-4.1 as the partner. Results show (a) emergence of some human–human-like grounding patterns, and (b) clear human–AI mismatches in grounding dynamics and burden of repair.
- Diagnostic value: The benchmark surfaces interaction patterns known to hinder grounding (single-turn responses, overly long answers, asymmetric grounding effort, shallow surface cues) and thus can evaluate whether an AI system supports genuine shared understanding rather than merely producing language artifacts that look like grounding.
Data & Methods
- Design rationale: The benchmark is grounded in empirical findings and theory from human–human collaboration. It intentionally requires joint problem-space construction and coordination so that successful performance depends on genuinely shared task representations, not just isolated responses.
- Task structure: Iterative, multi-turn conversational problem solving with referential elements (objects, states) that must be jointly tracked and updated; includes opportunities for clarification and repair; situation-awareness manipulations (e.g., different visibility or reviewability of task state) to vary grounding costs.
- Outcome measures (reported/described): task completion and accuracy, time/turns to completion, textual/linguistic metrics (message length, lexical persistence), counts/types of grounding acts (clarification requests, acknowledgments, repairs), measures of situation awareness and learning across trials.
- Validation sample and model: Confirmatory online user study pairing human participants with GPT-4.1. (The paper reports replication of known empirical grounding patterns and documents qualitative and quantitative divergences; exact sample size and numeric results are reported in the full text.)
- Contrast with prior methods: Unlike many assistant-style evaluations (one-shot tasks, static benchmarks), this benchmark emphasizes dynamic co-construction and repair, enabling assessment of whether apparent grounding corresponds to mutually actionable shared state.
Implications for AI Economics
- Productivity and coordination costs: If AI partners fail to establish durable common ground, human workers may incur substantial additional coordination and repair costs. Economists modeling AI augmentation should account not only for task-level automation gains but also for increased overhead from asymmetric grounding (time spent clarifying, correcting, or supervising AI).
- Valuation of AI as partner vs. assistant: The benchmark provides operational metrics (turns-to-completion, human repair time, error propagation) that can be used to quantify the marginal productivity of AI systems in collaborative tasks. This helps distinguish complementary (augments human capability) from substitutive (replaces human effort) outcomes in firm-level and labor-market analyses.
- Design and procurement incentives: Organizations buying or deploying AI for collaborative work should prioritize systems optimized for mutual grounding (e.g., proactive clarification, incremental proposals, shared workspace integration). Procurement frameworks and contracts might tie vendor evaluation to benchmarked reductions in human repair burden and improvements in joint task throughput.
- Training objectives and model evaluation: From an economic perspective, training objectives that internalize grounding — e.g., incentivizing clarification, shorter iterative turns, and explicit shared-state confirmations — may yield higher downstream labor productivity and lower oversight costs than optimizing only for one-shot accuracy metrics.
- Externalities and allocation of cognitive labor: The asymmetry observed (humans carrying the burden of repair) can concentrate cognitive strain on certain workers (e.g., managers, domain experts). Economic models of task allocation should incorporate such distributional effects and potential changes in wage premia for roles that require persistent AI supervision or corrective labor.
- Research priorities for policy and ROI: Policymakers and firms should fund and measure improvements in AI grounding capabilities because small percentage improvements in repair/coordination costs could yield large aggregate productivity gains across many collaborative tasks—especially in high-stakes decision-making, engineering, healthcare, and creative workflows.
- Empirical economic metrics to adopt: Use benchmark outputs to construct economic indicators such as human time per task saved (or lost), error-correction labor hours, changes in throughput, learning curves (reduction in human repair over repeated interactions), and incidence of catastrophic miscoordination. These can feed cost–benefit and adoption models.
Practical recommendations (brief) - For AI developers: Optimize for interactive, incremental exchanges (encourage clarification, confirmations, concise turns) and integrate shared workspaces / state visualization to lower grounding costs. - For firms/users: Adopt benchmarks like this when selecting AI tools for collaborative roles; measure human repair time and task throughput in pilots; train staff in effective prompting and repair strategies. - For economists and policy analysts: Incorporate grounding-related coordination costs into models of AI-driven productivity change and labor reallocation; use benchmark-derived metrics to estimate real-world complementarities and supervision burdens.
Assessment
Claims (9)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| AI is becoming increasingly integrated into everyday life, both in professional work environments and in leisure and entertainment contexts. Adoption Rate | positive | degree of AI integration into everyday life |
Reading fidelity
high
Study strength
medium
|
not reported
|
| This integration requires AI to move beyond acting as an assistant for informational or transactional tasks toward a genuine collaborative partner. Adoption Rate | positive | need for AI capabilities enabling collaboration |
Reading fidelity
high
Study strength
speculative
|
not reported
|
| Effective collaboration, whether between humans or between humans and AI, depends on establishing and maintaining common ground: shared beliefs, assumptions, goals, and situational awareness that enable coordinated action and efficient repair of misunderstandings. Team Performance | positive | quality of collaboration (coordination, repair of misunderstandings) |
Reading fidelity
high
Study strength
high
|
not reported
|
| While common ground is a central concept in human collaboration, it has received limited attention in studies of human-AI collaboration. Research Productivity | negative | research attention to common ground in human-AI studies |
Reading fidelity
high
Study strength
medium
|
not reported
|
| We introduce a new benchmark grounded in theories and empirical studies of human-human collaboration. Research Productivity | positive | availability of a theory-grounded benchmark for human-AI collaboration |
Reading fidelity
high
Study strength
medium
|
not reported
|
| The benchmark is based on a collaborative puzzle task that requires iterative interaction, joint action, referential coordination, and repair under varying conditions of situation awareness. Task Allocation | positive | task design capturing iterative interaction, joint action, referential coordination, repair, and situation awareness |
Reading fidelity
high
Study strength
medium
|
not reported
|
| We validate the benchmark through a confirmatory user study in which human participants collaborate with an AI to solve the task. Research Productivity | positive | benchmark validity via human-AI user study |
Reading fidelity
high
Study strength
medium
|
not reported
|
| The results show that the benchmark reproduces established theoretical and empirical findings from human-human collaboration. Team Performance | positive | similarity of human-AI collaboration patterns to established human-human collaboration findings |
Reading fidelity
high
Study strength
medium
|
not reported
|
| The results also reveal clear divergences in human-AI interaction. Team Performance | mixed | differences/divergences in interaction patterns between human-AI and human-human teams |
Reading fidelity
high
Study strength
medium
|
not reported
|