The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

A shop-floor conversational robot draws more attention but depresses sales-ready engagement: the bot (and a visibility-boosting fixture) prompts more passersby to stop yet lowers staff approaches, entries and purchases as clerks avoid interrupting and interactions cluster at the threshold.

From Metrics to Meaning: Insights from a Mixed-Methods Field Experiment on Retail Robot Deployment
Sichao Song, Yuki Okafuji, Takuya Iwamoto, Jun Baba, Hiroshi Ishiguro · January 05, 2026
arxiv quasi_experimental medium evidence 7/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

Arxiv

Latest observation:

  1. Sichao Song unresolved corpus identity
  2. Yuki Okafuji unresolved corpus identity
  3. Takuya Iwamoto unresolved corpus identity
  4. Jun Baba unresolved corpus identity
  5. Hiroshi Ishiguro unresolved corpus identity

Semantic Scholar

Latest observation:

  1. Sichao Song provider ID
  2. Yuki Okafuji provider ID
  3. Takuya Iwamoto provider ID
  4. Jun Baba provider ID
  5. Hiroshi Ishiguro provider ID
A conversational service robot increased passerby stopping (especially with a fixture) but reduced clerk-led approaches, store entries, assisted experiences, and purchases because clerks deferred to the robot and many interactions remained anchored at the doorway.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

We report a mixed-methods field experiment of a conversational service robot deployed under everyday staffing discretion in a live bedding store. Over 12 days we alternated three conditions--Baseline (no robot), Robot-only, and Robot+Fixture--and video-annotated the service funnel from passersby to purchase. An explanatory sequential design then used six post-experiment staff interviews to interpret the quantitative patterns. Quantitatively, the robot increased stopping per passerby (highest with the fixture), yet clerk-led downstream steps per stopper--clerk approach, store entry, assisted experience, and purchase--decreased. Interviews explained this divergence: clerks avoided interrupting ongoing robot-customer talk, struggled with ambiguous timing amid conversational latency, and noted child-centered attraction that often satisfied curiosity at the doorway. The fixture amplified visibility but also anchored encounters at the threshold, creating a well-defined micro-space where needs could ``close'' without moving inside. We synthesize these strands into an integrative account from the initial show of interest on the part of a customer to their entering the store and derive actionable guidance. The results advance the understanding of interactions between customers, staff members, and the robot and offer practical recommendations for deploying service robots in high-touch retail.

Summary

Main Finding

A small conversational retail robot significantly increased attention (stopping) at a store entrance—especially when paired with a purpose-built fixture—but reduced clerk-initiated downstream engagement (approach, entry, assisted experience) and lowered purchase rates per stopper. Net purchases per passerby did not change significantly. Qualitative interviews attribute this divergence to entrance anchoring (often child-focused), clerk non-interruption norms, ambiguous handoff timing due to conversational latency, and fixture-anchored “micro-spaces” where curiosity closed without interior engagement.

Key Points

  • Quantitative headline effects (Baseline → Robot-only → Robot+Fixture):
    • Stopping per passerby: 0.73% → 1.50% → 2.48% (Robot+Fixture > Robot-only > Baseline; p < .001).
    • Clerk approach per stopper: 19.73% → 13.41% → 4.83% (decline; p < .001).
    • Store entry per stopper: 31.97% → 20.12% → 9.46% (decline; p < .001).
    • Assisted in-store experience per stopper: 18.37% → 6.10% → 4.02% (decline; p < .001).
    • Purchase per stopper: 11.56% → 5.49% → 2.62% (decline; p < .001 for omnibus; Baseline > Robot+Fixture).
    • Purchase per passerby: 0.084% → 0.082% → 0.065% (no significant difference).
    • Sample product touch: increased most in Robot-only (63.72%) compared to Baseline (47.62%); Robot+Fixture was lower (46.68%).
  • Qualitative themes from six staff interviews:
    • T1 Entrance anchoring: robot captures gaze, especially children, increasing stops but not purchase intent.
    • T2 Conversion barriers: high-ticket items (bedding) require clerk negotiation; robot interest often insufficient.
    • T3 Timing/priority: clerks avoid interrupting robot-customer interaction and face ambiguous timing due to conversational latency.
    • T4 Role positioning: norms developed where robot attracts and clerks defer; handoffs were rarely explicit.
    • T5 Fixture effects: fixture increased visibility and stopping but anchored interactions at the threshold, reducing interior movement.
  • Interpretation: Robots can be powerful attention generators but may inadvertently reduce human-led selling if staging, handoff protocols, and staff coordination are not explicitly designed.

Data & Methods

  • Setting: live 12-day field experiment in a bedding retail store in Japan during normal operations.
  • Conditions: Baseline (no robot), Robot-only (robot at entrance), Robot+Fixture (robot mounted on an entrance fixture). Each condition ran 4 non-consecutive days, balanced for weekday/weekend.
  • Robot system: small humanoid (“Sota”), speech via Google Speech APIs, conversational/dialogue state managed with OpenAI GPT-4 family, front display for prompts; fixture provided platform and sample product placement (dimensions ~0.70m W × 0.58m D × 1.68m H).
  • Measures:
    • Primary funnel metrics: passerby counts; stopping per passerby; per-stopper outcomes — clerk approach, store entry, staff-led entry, assisted in-store experience (>30s), sample touch, and purchase.
  • Data collection/analysis:
    • Video annotation from entrance webcam and store 360° camera; two researchers logged observations.
    • Statistical tests: Pearson chi-square omnibus tests with Bonferroni-adjusted pairwise comparisons; Cochran–Armitage trend tests reported.
  • Qualitative component:
    • Six semi-structured staff interviews (1 manager + 5 clerks), 30–60 minutes each, audio-recorded and transcribed; reflexive thematic analysis with iterative coding.
  • Ethics/data handling: Institutional review approved; video stored encrypted; transcripts pseudonymized.

Implications for AI Economics

  • On returns to robot deployment
    • Attention ≠ conversion. Investments that raise passersby attention (stopping) do not guarantee higher sales for high-ticket, advice-dependent goods. ROI should be estimated with downstream conversion metrics (entry, assisted engagement, purchase per passerby), not only impressions or stopping.
    • The fixture was a low-cost staging investment that amplified visibility, but it changed the spatial and conversational ecology in ways that reduced conversion—so staging costs can have ambiguous marginal returns.
  • Labor complementarity vs substitution
    • Short-term effect: robot reduced clerk-initiated approaches (substitutive effect on the act of initiating contact), but that substitution was not productivity-enhancing because it lowered assisted selling (a complementary service for high-price items).
    • Policy/design implication: robots that attract but do not hand off effectively can weaken the complementarity that drives sales. Explicit, low-friction handoff protocols (visual cues, short turn-taking scripts, incentives for timely staff intervention) are needed to preserve complementarity.
  • Organizational externalities and behavioral frictions
    • Non-intrusion norms and ambiguous conversational latency created coordination frictions. Economists and managers should model not just customer responses but also endogenous staff behavior (priority setting, social norms) when predicting impacts of automation in service contexts.
  • Segmentation and product fit
    • Robot attraction skewed to children and accompanying adults; product-market fit matters. For child-oriented products, robots may improve conversion; for high-consideration items requiring consultation, robots may generate low-value traffic.
  • Measurement and evaluation recommendations
    • Use funnel-based KPIs: stopping per passerby, entry per stopper, assisted interactions, purchases per stopper and per passerby. Track sample touch and staff interventions separately.
    • Include staff behavior metrics and qualitative feedback in evaluations to capture coordination costs and latent effects.
    • Run randomized or alternating-condition deployments and measure heterogeneity by customer segment and time-of-day to estimate net revenue impacts.
  • Design & incentive suggestions for practitioners
    • Make handoffs explicit and low-cost: short robot prompts like “A staff member will be right with you” plus a visual signal to staff.
    • Spatially separate “attract” (robot/fixture) and “recommend” (clerk engagement) zones so curiosity does not close at the threshold.
    • Train staff on timing and minimal interruption norms; consider scheduling or incentive mechanisms to reduce hesitation in intervening.
    • Tailor robot behavior and staging to product type and target customer segment.
  • Research opportunities for AI economics
    • Model the multi-agent interaction: customers, robot, and staff as strategic agents with incomplete information and coordination frictions—analyse when robot deployment increases or decreases expected revenue.
    • Dynamic/adaptive policies: evaluate learning algorithms that adjust robot behavior or handoff timing to maximize conversions while minimizing staff coordination costs.
    • Long-run equilibrium: study whether norms and staff behaviors evolve (e.g., staff learn to intervene) and how that changes welfare and firm-level returns over time.

Bottom line: Service robots can reliably draw attention but may depress revenue-generating human interactions unless deployment design explicitly integrates staging, handoffs, staff incentives, and product-market fit. Economic evaluations must therefore include downstream conversion metrics and staff-behavior effects, not only attention metrics.

Assessment

Paper Typequasi_experimental Evidence Strengthmedium — The field experiment provides direct behavioral measures (video-coded stopping, entry, purchase) and consistent patterns across conditions, giving reasonably credible evidence for effects on early-stage attention; however, short duration (12 days), single-site deployment, potential novelty effects, non-randomized condition sequencing, and staff discretion weaken causal certainty for downstream outcomes and limit external validity. Methods Rigormedium — Strengths include a mixed-methods design, systematic video annotation of the customer funnel, and follow-up interviews to unpack mechanisms; weaknesses include lack of randomized assignment, small temporal/sample scope (one store, 12 days), limited interviewer sample (six staff), and no reported robustness checks or statistical controls for time-varying confounders. SampleObservational data from passersby and in-store customers at a single bedding retail store across 12 days under three alternating conditions (Baseline, Robot-only, Robot+Fixture), with video annotations of stages from passerby to purchase (stopping, clerk approach, store entry, assisted experience, purchase); six post-experiment staff interviews used for qualitative interpretation. Exact counts and demographics are not reported in the abstract. Themeshuman_ai_collab org_design IdentificationAlternating three conditions (Baseline, Robot-only, Robot+Fixture) across 12 days in a single live bedding store, comparing video-annotated customer funnel outcomes within-store across time blocks; accompanying staff interviews used to interpret mechanisms. Causal claims rely on within-store contrasts over time rather than randomized assignment, and staff discretion was not controlled. GeneralizabilitySingle-store setting limits transferability to other retail types or scales, Short deployment (12 days) risks novelty effects and limits long-run inference, Staff behavior and discretion are context-specific and may not generalize, Customer demographics/local foot-traffic patterns likely idiosyncratic, Findings tied to the specific robot model and fixture design used, Sequencing and time-of-day/season effects not fully controlled

Claims (7)

ClaimDirectionOutcomeConfidence & EvidenceDetails
The robot increased stopping per passerby (highest with the fixture). Adoption Rate positive stopping per passerby (passersby who stopped to engage with the robot/service)
Reading fidelity high
Study strength medium
not reported
0.48
Clerk-led downstream steps per stopper—clerk approach, store entry, assisted experience, and purchase—decreased when the robot was present. Firm Revenue negative rates of clerk approach, store entry, assisted experience, and purchase per stopper
Reading fidelity high
Study strength medium
not reported
0.48
Clerks avoided interrupting ongoing robot-customer talk, which contributed to reduced clerk-led downstream engagement. Task Allocation negative frequency or likelihood of clerk interruption/approach during robot-customer interactions
Reading fidelity high
Study strength medium
n=6
0.48
Clerks struggled with ambiguous timing amid conversational latency from the robot, complicating decisions about when to intervene. Task Allocation negative clerks' decision-making/timing for intervention during interactions
Reading fidelity high
Study strength low
n=6
0.24
The robot produced a child-centered attraction that often satisfied curiosity at the doorway, reducing customers' movement into the store. Adoption Rate negative store entry rate (share of stoppers who enter the store)
Reading fidelity high
Study strength medium
not reported
0.48
The additional fixture amplified robot visibility (increasing stopping) but anchored encounters at the threshold, creating a micro-space where needs could 'close' without moving inside. Adoption Rate mixed visibility/stopping per passerby and store entry (threshold anchoring reducing entry)
Reading fidelity high
Study strength medium
not reported
0.48
Study design: a mixed-methods field experiment alternating three conditions over 12 days with video annotation of the service funnel and six post-experiment staff interviews. Research Productivity null_result research design and data collection procedures (duration, conditions, video annotation, interview count)
Reading fidelity high
Study strength high
n=12
0.8

Notes