The Commonplace
Home Three-study pilot Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

A simple, exact metric reveals humans use implicit conventions to avoid mistakes far more than AI partners do: human pairs outperform literal-hint predictions by about 26 percentage points, AI pairs show near-zero advantage, and human-AI teams fall in between—implying that compatibility with human conventions, not self-play rank, predicts teaming success.

The Convention Gap: Towards Measuring Implicit Communication in Cooperative AI Evaluation
Makoto Fukushima, Hua-Dong Xiong, Ehsan Moradi Pari · September 10, 2026
arxiv descriptive medium evidence 7/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

Arxiv

Latest observation:

  1. Makoto Fukushima unresolved corpus identity
  2. Hua-Dong Xiong unresolved corpus identity
  3. Ehsan Moradi Pari unresolved corpus identity
The paper introduces the 'convention gap'—the difference between failure risk predicted from literal hints and observed failures—and shows humans exploit implicit conventions (gap ≈ +26 pp), AIs do not (≈ −0.7 pp), with human-AI play intermediate (≈ +16 pp), suggesting convention compatibility matters for teaming.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

Cooperative AI agents are evaluated against other AIs, yet human cooperation relies on implicit conventions---shared protocols for reading meaning beyond the literal message---which AI-AI benchmarks may not capture. We propose the \emph{convention gap}, the difference between the failure probability predicted from the literal content of communication and the observed failure rate, as a metric of implicit communication. In the card game Hanabi, the finite deck and deterministic hint constraints make this posterior exactly computable. We replayed about 101,000 play actions from three public datasets of human-human (hanab.live), AI-AI (HOAD), and human-AI (HanabiData) games. The gap was +26.2 percentage points (pp) in human pairs, $-$0.7~pp in AI pairs, and +16.4~pp in human-AI pairs, and was concentrated on plays of cards that had received no hints (+46~pp in human pairs). Within human-AI play, the literal information available to humans was similar across the three AI partners (mean predicted failure 38--41\%), but human failure rates ranged from 14.4\% to 34.4\% and the gap from +24.1 to +6.2~pp; the partner eliciting the largest gap produced the fewest human failures. Game score carried different information: it depended on each corpus's roster composition, whereas the gap separated human from AI play at the agent level. As a known-answer check, Off-Belief Learning agents, whose convention content is controlled by construction, gave a gap of +1.6~pp at the convention-free level, rising monotonically to +21.7~pp. These results suggest that convention compatibility, rather than AI-AI performance, may predict an AI's effectiveness with human partners.

Summary

Main Finding

The paper introduces the "convention gap"—the difference between the failure probability predicted from literal communication alone and the observed failure rate—as an objective metric of implicit (conventional) communication in cooperative AI evaluation. Applied to Hanabi play logs (≈101K plays), the metric shows a clear gradient of implicit-communication use: human-human pairs have a large positive gap (+26.2 percentage points), AI-AI pairs are near-calibrated (−0.7 pp), and human-AI pairs fall in between (+16.4 pp). The gap is concentrated on plays where no literal hint exists (untouched cards), varies by AI partner, correlates inversely with partner loss, and is reproducible on controlled Off-Belief Learning agents.

Key Points

  • Definition: Convention Gap = mean posterior probability of a life loss computed from literal hint information only − observed life-loss rate. Positive gap → players succeed more than literal information predicts (evidence of implicit conventions).
  • Exact posterior: In Hanabi, because of a finite deck and deterministic hint rules, the posterior over card identities (and hence the failure probability for a play) is exactly computable by enumeration—no learning or sampling needed.
  • Summary results (≈100,956 plays across datasets):
    • Human-human (hanab.live): gap = +26.2 pp.
    • AI-AI (HOAD): gap = −0.7 pp (nearly calibrated to literal info).
    • Human-AI (HanabiData): gap = +16.4 pp.
  • Concentration on zero-/one-hint plays: the human gap is largest when the played card received no hints (≈+46 pp), and collapses near zero once two or more hints touch the card.
  • Variation by partner: within human-AI games the literal-information posterior available to humans was similar across AI partners (predicted failure ≈38–41%), but observed human failure rates ranged 14.4%–34.4% and convention gaps varied +6.2 to +24.1 pp; the AI that elicited the largest gap caused the fewest human failures.
  • Validation: Off-Belief Learning (OBL) agents with controlled convention content produced the expected hierarchy of gaps (+1.6 pp with no conventions up to +21.7 pp as conventions are reintroduced), supporting the metric's behavioral interpretation.
  • Controls and robustness: calibration checks with rule-based AIs, binning by posterior, bootstrap CIs and clustering, and analysis by hint count support the interpretation that the gap isolates non-literal (conventional) information.

Data & Methods

  • Task & environment: Standard 2-player Hanabi ("No Variant"), where hints are limited to naming color or rank and pointing to matching cards; misplays cost shared life tokens.
  • Posterior computation: exact weighted enumeration over feasible card identities given the literal information set ht (hints on the card, visible hands, fireworks, discards, tokens). Steps:
  • Apply hard filters from hints to each card’s candidate identities.
  • Weight surviving candidates by remaining unseen copies.
  • Sum weights for playable vs. unplayable candidates to get per-card failure probability.
  • (Optional fast per-card approx omits joint-hand conditioning).
  • Full posterior conditions jointly on the whole hand by enumerating consistent joint assignments with falling-factorial weights; marginalize to the played card.
  • Datasets and scale:
    • hanab.live: 425 completed human 2-player games → 9,017 play records (bot accounts filtered).
    • HOAD: 100 games × 49 AI pairings (7×7) → 62,890 plays.
    • HanabiData (human-AI): 2,040 games from 240 players paired with 3 AI types → 29,049 plays (15,472 human plays).
    • Total ≈100,956 play actions.
  • Failure-rate metrics:
    • Player Loss: fraction of an agent’s own plays that fail.
    • Partner Loss: fraction of partner plays that fail when paired with this agent.
  • Hint-quality diagnostics recorded per hint:
    • Disambiguation power = fraction of candidate identities eliminated by the hint.
    • Playability rate = fraction of hinted cards that are currently playable.
  • Validation & robustness: bootstrapping, agent-level clustering, binning by posterior, comparison to OBL agent hierarchy, and checks that AI behavior is calibrated to the posterior (so deviations reflect human implicit channels).

Implications for AI Economics

  • Evaluation and procurement:
    • Standard AI-AI benchmarks (self-play/cross-play rankings) can misrepresent an agent’s value in human teams. The convention gap is a candidate metric for procurement and evaluation when the AI will operate with humans—organizations should measure convention compatibility, not just self-play score.
  • Product-market fit and adoption:
    • Compatibility (how well an AI elicits and conforms to human conventions) is a distinct product attribute. Buyers will value agents with higher compatibility with their human workforce or user base even if those agents do not top AI-AI leaderboards. This can drive market segmentation—agents optimized for human teaming vs. agents optimized for self-play performance.
  • Network effects and complementarities:
    • Human-AI teaming exhibits complementarities: an AI’s marginal value depends on partner type and conventions. The convention gap quantifies one dimension of complementarity and can be used to price contracts, staffing decisions, or match agents to teams.
  • Investment and R&D incentives:
    • Investors and R&D managers should consider funding human-centered coordination research (e.g., zero-shot coordination, OBL-style grounding) because convention compatibility affects downstream utility. Metrics like the convention gap enable return-on-investment estimates for human-compatible features.
  • Productivity, risk, and externalities:
    • Mis-evaluating an AI’s human compatibility (relying on AI-AI benchmarks) risks deployment failures, reduced productivity, or increased error costs in collaborative tasks. Regulators and firms should account for these coordination externalities when assessing safety, certification, or insurance.
  • Standards, certification, and market signaling:
    • A measurable, comparable statistic (like the convention gap) supports certification regimes and labels for "human-compatible" cooperative AIs. Such signaling can reduce search and matching costs in procurement and increase trust.
  • Limitations & scalability:
    • The convention gap is exactly computable here because Hanabi has a finite known state and deterministic hint rules. Extending the metric to richer real-world settings will require approximations (e.g., learned literal-baseline models, causal identification of literal vs. conventional signals) and careful validation. Economic models and policy should account for estimation error when generalizing beyond structured games.
  • Practical uses for economists and decision-makers:
    • Incorporate convention-gap-like measures into cost–benefit analyses of AI deployment for collaborative tasks.
    • Use the metric in comparative evaluations when choosing AI partners for human teams, and in designing incentive/compensation schemes that reflect complementarities.
    • Factor convention compatibility into pricing/licensing and product-differentiation strategies (agents optimized for human compatibility can command premiums or serve niche markets).

Overall, the convention gap offers a practical, interpretable diagnostic separating literal-information performance from convention-enabled human coordination—an attribute with direct implications for evaluation, procurement, market structure, and policy when AI systems are deployed to collaborate with humans.

Assessment

Paper Typedescriptive Evidence Strengthmedium — Large sample (~101k play actions) across three public datasets, exact closed-form posterior computation, and multiple robustness/validation checks (AI calibration, OBL known-answer test) provide convincing descriptive evidence that a non-literal information channel exists and differs across settings; however, the analysis is observational in a toy domain (Hanabi) and cannot establish causal mechanisms or generalize directly to real-world team productivity. Methods Rigorhigh — The paper leverages Hanabi's finite combinatorics to compute an exact posterior without approximation, uses a custom replay engine to reconstruct information states, applies bootstrapping/clustering robustness checks, calibrates against multiple AI baselines, and includes a known-answer validation (OBL agents); remaining limitations are dataset selection/exclusion rules and the observational nature of the comparisons. SampleThree public Hanabi corpora: hanab.live (human-human) — 425 completed 2-player games, 9,017 play records (after bot exclusions); HOAD (AI-AI) — 100 games for each of 49 AI pairings (7 AIs × 7), 62,890 play records; HanabiData (human-AI) — 2,040 games from 240 players paired with 3 AI types, 29,049 play records (15,472 human plays); total ≈100,956 play actions analyzed, all standard 2-player 'No Variant' Hanabi. Themeshuman_ai_collab adoption IdentificationCompute an exact Bayesian posterior probability of a play causing a life loss using only literal hint information (enumeration over feasible card identities given Hanabi's deterministic hint rules), then define the 'convention gap' as the difference between that posterior-predicted failure rate and the observed failure rate; validate interpretation with controls (rule-based AI calibration, concentration on hint-free plays, Off-Belief Learning known-answer hierarchy) but no exogenous variation or causal identification. GeneralizabilityDomain is a stylized, small-state cooperative game (Hanabi) — results may not generalize to complex real-world human-AI tasks., Analysis limited to 2-player standard configuration; multiagent settings or different rule variants may behave differently., Public datasets may have selection biases (player skill distribution, bot-filtering choices, cultural/behavioral idiosyncrasies)., Studied AI architectures are a subset (HOAD, OBL, specific agents); other training regimes or multimodal communication channels might produce different convention behavior., Metric isolates implicit communication in a setting with deterministic hint semantics; settings without analogous literal channels may not permit the same exact computation., Observational replay of logged games — no randomized assignment of partners or interventions to identify causal mechanisms.

Claims (9)

ClaimDirectionOutcomeConfidence & EvidenceDetails
The convention gap was +26.2 percentage points in human-human play, −0.7 percentage points in AI-AI play, and +16.4 percentage points in human-AI play. Error Rate positive Difference between literal-information-predicted life-loss probability and observed life-loss rate
Reading fidelity high
Study strength medium
n=100956
+26.2 pp in human pairs; −0.7 pp in AI pairs; +16.4 pp in human-AI pairs
0.18
Human pairs succeeded more often than literal hint information predicted, whereas AI-AI play was approximately calibrated to the literal-information posterior. Error Rate mixed Observed life-loss rate relative to posterior-predicted life-loss probability
Reading fidelity high
Study strength medium
n=71807
+26.2 pp for human pairs; −0.7 pp for AI pairs
0.18
The human convention gap was concentrated in plays of cards that had received no hints, with a gap of approximately +46 percentage points. Error Rate positive Convention gap for plays involving cards with no literal hint information
Reading fidelity high
Study strength medium
+46 pp
0.18
Within human-AI play, human failure rates differed substantially across the three AI partners despite similar literal-information predictions: mean predicted failure was 38–41%, while human failure rates ranged from 14.4% to 34.4%. Team Performance mixed Human life-loss/failure rate when partnered with different AI agents
Reading fidelity high
Study strength medium
n=15472
Mean predicted failure 38–41%; observed human failure 14.4–34.4%
0.18
Among the three human-AI partners, the AI that elicited the largest convention gap produced the fewest human failures. Team Performance negative Human partner failure rate as a function of AI partner’s convention gap
Reading fidelity high
Study strength medium
n=15472
Convention gaps ranged from +24.1 to +6.2 pp
0.18
The convention gap separated human from AI play at the agent level more effectively than mean game score, while game score depended on the composition of each corpus’s roster. Team Performance mixed Ability of convention gap versus game score to distinguish cooperation settings and agent types
Reading fidelity high
Study strength low
n=100956
0.09
Off-Belief Learning agents produced a convention gap of +1.6 percentage points at the convention-free level, increasing monotonically to +21.7 percentage points as convention content increased. Error Rate positive Convention gap under controlled levels of implicit convention content
Reading fidelity high
Study strength medium
+1.6 pp to +21.7 pp
0.18
The literal-information posterior for Hanabi play outcomes can be computed exactly by enumeration without sampling, parameter estimation, or approximation. Other positive Exact computability of the posterior probability of life loss
Reading fidelity high
Study strength high
not reported
0.3
All seven rule-based AI agents had own-play convention gaps between −5.8 and +0.7 percentage points, and the pooled AI calibration curve was within 3 percentage points of the diagonal for posterior failure probabilities of 0.5 or lower. Error Rate null_result AI own-play failure rate relative to literal-posterior predictions
Reading fidelity high
Study strength medium
Individual gaps from −5.8 to +0.7 pp; pooled calibration within 3 pp
0.18

Notes