The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

AI-written and human-edited AI marketing emails substantially boost short-run sales: randomized trials at one online retailer show LLM and hybrid emails roughly double gross order profits versus no-email, and a profit-based decision rule prefers AI options to conventional human writers after accounting for labor and software costs.

Large Language Models and Creative Content Design: a case study of email marketing at Wine Access
Jean-Pierre Dubé, Ariel Xu · January 13, 2026 · Quantitative Marketing and Economics
openalex rct medium evidence 8/10 relevance Full text usable extracted full text DOI Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

OpenAlex

Latest observation:

  1. Jean-Pierre Dubé provider ID
  2. Ariel Xu provider ID

Semantic Scholar

Latest observation:

  1. Jean-Pierre Dubé provider ID
  2. A. Xu provider ID
Three RCTs at a single online retailer find that LLM-generated and hybrid human-edited LLM email content roughly double gross profits relative to a no-email control, and a decision-theoretic calculation that accounts for labor and software costs selects an AI-based policy over traditional salaried writers.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

Abstract A sequence of three randomized controlled trials (RCTs) is conducted to support a small online business’ decision of whether and how to implement AI in the creation of email marketing content. Recent developments in frequentist statistical decision theory are used to accommodate small samples available for testing in the small-business setting. The RCTs comprise three test policy cells with email content created by (i) salaried writers (“human”), (ii) a large-language model (“LLM”), and (iii) a “hybrid” combination of a human editing the content created by the LLM, respectively. When a “no email” control policy is included, all three test cells approximately double the gross profits from orders relative to the control cell. The RCTs vary whether the hybrid cell is edited by a salaried writer or the marketing team. The LLM cells vary whether the AI is pre-trained using historic emails or uses a prompt-based generative pre-trained transformer (“GPT”). Decision theory always selects one of the AI cells over the standard human policy on the basis of total annual profit net of related labor and software overhead.

Summary

Paper: Dubé, J.-P. & Xu, A. (2026). Large Language Models and Creative Content Design: a case study of email marketing at Wine Access. Quantitative Marketing and Economics (24:1). https://doi.org/10.1007/s11129-025-09303-9

Main Finding

In a three-stage randomized controlled case study at a small online wine retailer, LLM-generated email newsletters—either used directly or in a hybrid (LLM + human editing) workflow—produce roughly the same gross profits per campaign as newsletters written by a salaried writing team, while materially reducing annual operating costs. Under a frequentist minimax-regret decision rule (asymptotic minimax regret), one of the AI policies is chosen over the full-human policy in each experiment, implying that AI implementation (automation or augmentation) is the profit-maximizing choice for this firm.

Key Points

  • Experimental design: three randomized field experiments comparing four cells (no-email control, Human writers, LLM-generated, Hybrid LLM→human edits). Later experiments varied who edited the hybrid output (professional writer vs. marketing head) and switched from a pre-trained LLM to a prompt-based GPT (Claude).
  • Main outcomes: all three active email cells (Human, LLM, Hybrid) approximately doubled purchase incidence, revenues, and gross profits relative to the no-email control. Differences in gross profit between Human and AI cells were statistically insignificant at pre-registered levels.
  • Cost effect: because AI policies substantially reduce writer labor costs (with modest software license fees), decision-theoretic analysis selects at least one AI policy over the Human policy in every RCT when considering annual net profit.
  • Productivity heterogeneity: evidence suggests salaried writers’ productivity declined under AI automation, while marketers receptive to AI experienced productivity gains under hybrid workflows.
  • Contribution: extends prior LLM email-subject-line RCTs by testing LLMs on full email creative and evaluating financial outcomes (gross profits) and normative firm decisions using frequentist decision theory.

Data & Methods

  • Setting: Wine Access (WA), a Napa-based online wine retailer that sends twice-daily newsletter campaigns (NC). Unit of analysis: a "campaign" (twice-daily NCs over two weeks).
  • Samples & power: WA allocated a sampling frame of ~27,500 regular customers per experiment. Decision-theoretic (AMMR) framework implies much smaller sample requirements than conventional hypothesis testing; several thousand consumers per cell suffice for useful decisions.
  • RCTs:
    • RCT1 (early 2024): 4 cells — Control (no NC), Human (3-person salaried writing team), LLM (pre-trained LLM via GUI), Hybrid (LLM output edited by a salaried writer not in the Human cell).
    • RCT2: Hybrid editing shifted to the marketing team (more AI-positive editors).
    • RCT3 (2025): switched from pre-trained LLM to prompt-based GPT (Claude); prompts included wine attributes plus an example historical offer; no historic-email training dataset used in this trial.
  • LLMs & implementation:
    • Pre-trained LLM: Mistral7B trained on five years of WA historic newsletters (3,148 unique NCs, 2019–2023) via a third-party consultant and delivered via a GUI that generated draft NCs in ~5 seconds.
    • GPT (Claude): prompt-based generation requiring more prompt effort; chosen for readability and lower hallucination rate; licensing fee differences noted.
  • Economic inputs / costs:
    • Current human-writing cost: three salaried writers at $125,000 each ($375,000/year).
    • Pre-trained LLM license: ~$1,000/year; adopting LLM automation could reduce writer labor by ~91.67% (retaining one part-time writer).
    • Claude license: ~$1,200/year; different labor savings under automation/hybrid (83.33% or 75% depending on configuration).
    • Typical human NC writing time: ~3 hours; LLM GUI produced drafts in seconds (but GPT prompts required more editing).
  • Decision theory:
    • Frequentist minimax-regret approach (asymptotic minimax regret, AMMR) used to choose policies based on estimated treatment effects and annual net profits (gross profits from orders minus labor and license overhead).
    • The optimal test decision reduces to threshold rules on estimated treatment effects; threshold examples computed relative to break-even annual costs (e.g., a per-consumer breakpoint of ~$0.35 in an illustrative comparison).
    • Advantage: actionable policy recommendations with realistic small-sample sizes for small firms; does not depend on p-values or standard errors in the same way null-hypothesis tests do.
  • Outcomes:
    • Treatment effects measured over 2-week campaigns; all three active cells produced statistically and economically significant increases over no-email control.
    • No robust statistical difference between Human vs. AI cells in gross profits across experiments.
    • Decision-theoretic selection favored AI in all experiments after accounting for license and labor cost differences.

Implications for AI Economics

  • Automation vs. augmentation: This case illustrates scenarios where AI can substitute for routine creative labor (email copywriting) without loss of revenue—creating strong incentives for labor displacement in small digital-content tasks. Hybrid workflows can yield productivity gains for staff who are receptive and adept at working with AI.
  • Cost-effectiveness for small firms: With modest license fees and substantial labor cost reductions, LLMs can be profit-increasing for small firms even when performance (gross profit per campaign) is roughly equal to human-generated content. Decision-theoretic methods make rigorous adoption decisions feasible with small samples.
  • Heterogeneous effects and labor markets: The heterogeneous productivity effects (decline for salaried writers; increase for receptive marketers) imply that AI adoption will have distributional consequences—differential demand for editing/supervisory skills and possible job reallocation toward monitoring, prompt design, and fact-checking.
  • Research and policy implications:
    • External validity caution: results are from a single firm, product category (wine), and unpersonalized newsletter creative; outcomes may differ for other product types, highly personalized targeting, or firms that already employ individualized creative strategies.
    • Quality control and hallucinations: even with good empirical outcomes, LLMs occasionally hallucinate or make factual errors; firms must invest in editing/fact-checking protocols, which affects net labor savings.
    • Dynamic effects: as LLMs and prompt-based GPTs evolve, the comparative advantages, editing needs, and cost structures will change; continuous evaluation (via small RCTs + decision theory) is advisable.
    • Labor policy: evidence of substitutability in routine creative tasks reinforces arguments for worker reskilling, safety nets, and policies to smooth transitions for displaced creative workers.
  • Methodological implication for applied micro studies: the paper showcases AMMR/frequentist decision-theory as a practical framework for small firms to make data-driven adoption choices when conventional large-sample hypothesis testing is infeasible.

Limitations to note: single-firm case study (generalizability limits), some decision-theoretic approximations (AMMR multi-arm extension not solved), assumed symmetric weight on Type I/II errors, and the campaign-level (not per-email) treatment effect assumption.

Assessment

Paper Typerct Evidence Strengthmedium — Internal validity is strong because effects are estimated via randomized assignment, and decision-theory methods address small-sample inference; however strength is limited by small sample sizes, a single small-business setting, possible unreported implementation details (e.g., balance checks, pre-registration), short-run outcomes, and sensitivity to specific LLM versions and training choices. Methods Rigormedium — Use of RCTs and recent frequentist decision-theory adjustments for small samples indicates careful methodology, but the manuscript appears to lack information on sample sizes, randomization checks, pre-registration, long-run follow-up, and robustness to alternative outcome definitions, which lowers overall rigor. SampleThree RCTs run in a single small online business comparing email content created by (i) salaried human writers, (ii) a large-language model (LLM) — with variants including historic-email pre-training or prompt-based GPT — and (iii) a hybrid where humans edit LLM output (editing done by either a salaried writer or the marketing team); some trials include a no-email control; primary outcomes are gross profits from orders and a decision-theoretic evaluation of annual profit net of labor and software overhead; exact sample sizes not reported in the abstract. Themesproductivity human_ai_collab IdentificationRandomized controlled trials that randomly assign email recipients (or equivalent marketing units) to one of three test policies (human-written, LLM-generated, hybrid human-edited LLM) and a no-email control in some specifications; causal effects identified by random assignment, with frequentist decision-theory methods applied to accommodate small sample inference and to select the profit-maximizing policy net of labor and software costs. GeneralizabilitySingle small online business — limited external validity to other firms or industries, Specific to email marketing — may not generalize to other marketing channels or product types, Results depend on specific LLM implementations and training (historic-email pre-training vs prompt-based GPT) and may not hold for other models or future model updates, Short-run sales/profit outcomes observed; long-run effects (customer lifetime value, brand effects, churn) not assessed, Small sample sizes reduce precision and may amplify idiosyncratic firm-level effects, Organizational factors (marketing skill, editing quality) and local market demographics may limit applicability

Claims (5)

ClaimDirectionOutcomeConfidence & EvidenceDetails
A sequence of three randomized controlled trials (RCTs) is conducted to support a small online business’ decision of whether and how to implement AI in the creation of email marketing content. Other null_result implementation decision support (whether/how to implement AI for email marketing)
Reading fidelity high
Study strength medium
not reported
0.6
The RCTs comprise three test policy cells with email content created by (i) salaried writers ("human"), (ii) a large-language model ("LLM"), and (iii) a "hybrid" combination of a human editing the content created by the LLM. Other null_result type of email content generation policy (human vs LLM vs hybrid)
Reading fidelity high
Study strength high
not reported
1.0
When a "no email" control policy is included, all three test cells approximately double the gross profits from orders relative to the control cell. Firm Revenue positive gross profits from orders
Reading fidelity high
Study strength medium
approximately double
0.6
The RCTs vary whether the hybrid cell is edited by a salaried writer or the marketing team, and LLM cells vary whether the AI is pre-trained using historic emails or uses a prompt-based generative pre-trained transformer ("GPT"). Other null_result variation in hybrid editing workflow and LLM training/prompting approach
Reading fidelity high
Study strength high
not reported
1.0
Applying recent developments in frequentist statistical decision theory to the trial data always selects one of the AI cells over the standard human policy on the basis of total annual profit net of related labor and software overhead. Firm Revenue positive total annual profit net of related labor and software overhead
Reading fidelity high
Study strength medium
not reported
0.6

Notes