Apparent welfare gains from simulated marketplace guardrails collapse when the evaluation scaffold and agent incentives are controlled, revealing fragile construct validity in LLM agent commerce studies; simple reporting of buyer surplus can mask opposite welfare effects.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
No provider observation is available for this paper.
Missing data, not a zero citation count.
Interactive simulations increasingly evaluate policies in markets populated by language-model agents. Their outputs can look economic---prices, profits, consumer surplus, and welfare---without instantiating the behavior named in the claim. We audit this risk in a multi-turn buyer--seller testbed for configurable hotel transactions. An initial implementation reported welfare gains from two marketplace guardrails of +87.4, +35.0, and +28.8 across a Qwen2.5 1.5B--14B ladder. It also gave guarded and unguarded agents different offer schemas and choice procedures. Holding the schema and buyer chooser fixed changes the paired contrasts to +7.2, -13.9, and +23.8. The four largest 14B single-generation effects averaged +229; after three generations per profile-condition, they averaged +37.6 (95% bootstrap interval [-34.2, 109.3]), while generation residuals account for 49.9% of variation in this post-hoc probe. A seller-incentive check is non-monotone: increasing profit pressure produces less profit than the default seller prompt. Scripted positive controls show why this matters. A profit-maximizing seller already attains first-best welfare, so guardrails mostly redistribute and reduce welfare; they create welfare only when the seller is explicitly programmed to force inefficient bundles. We contribute a construct-validity contract separating incentive validity, protocol isolation, stochastic stability, and welfare accounting, and returning INVALID or INCONCLUSIVE before substantive policy claims. In our case, the original estimate is INVALID under protocol isolation, while the controlled study remains INCONCLUSIVE under incentive validity and stochastic stability. The case does not show that guardrails are ineffective; it shows their apparent value is unidentified until the simulated agents and protocol pass these checks.
Summary
Main Finding
Apparent welfare gains from simple “guardrails” in LLM-agent buyer–seller simulations can be artefacts of implementation and sampling choices. When treatment scaffolds (offer schema and buyer decision procedure) are unified and generation noise is accounted for, the originally large positive effects largely disappear or become statistically and substantively unresolved. Construct-validity checks (incentives, protocol isolation, stochastic stability, accounting) are necessary before making causal policy claims from A2A market simulations.
Key Points
- Headline vs controlled result:
- Original asymmetric implementation (different offer schema/chooser across cells) reported Both−None welfare gains of +87.4, +35.0, +28.8 (Qwen2.5 1.5B, 3B, 14B).
- After holding offer schema and buyer chooser constant, the paired Both−None contrasts became +7.2 (95% CI [−8.1, 23.8]), −13.9 ([−26.2, −5.2]), and +23.8 ([−1.5, 56.6]) — large shrinkage and a sign reversal at 3B.
- Single-generation instability:
- Four largest 14B single-generation gains averaged +229 welfare units.
- After collecting three generations per profile-condition, those four profile effects averaged +37.6 (95% bootstrap CI [−34.2, 109.3]).
- Generation residuals explained 49.9% of sum-of-squares in this post-hoc probe; exact profile-level sign-flip placebo p = 0.50 (two-sided).
- Incentive validity failure (non-monotone response):
- A seller prompt designed to increase profit-seeking produced more probing but less profit than the standard prompt (example: standard profit 57.5 vs profit-pressure 33.8 on a small 3-profile test).
- C1 (incentive validity) therefore remained Inconclusive — LLM prompts did not reliably produce monotone strategic responses.
- Scripted positive controls clarify mechanism:
- Deterministic profit-maximizing seller attains first-best welfare; guardrails mainly redistribute surplus (transfer) rather than create welfare.
- Deterministic “inefficient-bundling” seller (which forces inefficient components when ungoverned) shows guardrails can increase welfare in settings with pathological seller behavior. Thus guardrails’ welfare effect depends critically on the assumed seller technology.
- Evaluation contract (four predicates):
- C1 Incentive validity — seller responds to stated objective.
- C2 Protocol isolation — only intended constraints vary (schema/chooser fixed).
- C3 Stochastic stability — treatment effects exceed generation noise; reps nested within profile-condition.
- C4 Accounting completeness — report completion, buyer surplus, seller profit, welfare from ground truth.
- Decision rule: fail C1 or C2 → Invalid causal claim; fail to resolve C1/C3 → Inconclusive; all pass → substantive interpretation allowed.
- Applied to this study:
- Original headline: Invalid (C2 violated).
- Controlled study: Inconclusive overall (C1 and C3 unresolved).
- C4 passed (welfare accounting was reported).
Data & Methods
- Testbed: configurable hotel transactions (mandatory base S, optional add-ons). Buyer accepts the utility-maximizing feasible offer if it weakly beats an assigned outside option. Ground-truth metrics:
- Buyer surplus BSθ(S,p) = Vθ(S) − p − oθ
- Seller profit SPθ(S,p) = p − C(S)
- Welfare Wθ = Vθ(S) − C(S) − oθ
- Welfare reported as fraction of profile-specific first-best (first-best selects optional components where private value > cost).
- Models: Qwen2.5-Instruct at 1.5B, 3B, 14B; same model for buyer and seller per run; temperature = 0.2.
- Profiles: 30 held-out synthetic profiles used for main paired analyses; 60-profile panel used for scripted deterministic controls.
- Experiments:
- E1 Scaffold control: unified offer schema and buyer chooser vs original asymmetric scaffold.
- E2 Repeated-generation forensic: for four profiles with largest single-generation gains, collected 3 generations per profile-condition to probe winner’s-curse and generation variance.
- E3 Incentive manipulation check: three seller prompts (compliance, standard “self-interested”, stronger profit-pressure) at 14B on 3 profiles with 2 generations each.
- E4 Scripted positive controls: deterministic profit-maximizing and inefficient-bundling sellers across 60 profiles with the same schema/chooser.
- Statistics:
- Profile-paired contrasts (Both−None), profile-level bootstrapping (20,000 draws) for CIs.
- Repetitions averaged within profile-condition (no pseudo-replication).
- Exact sign-flip placebo enumeration for probe; decomposition of sums-of-squares to quantify generation residuals.
- Reported approximate MDE to highlight low precision when relevant.
Implications for AI Economics
- Construct validity is essential: labeling an LLM as “self-interested” or a prompt as a “guardrail” is not enough. Researchers must operationalize and test behavioral implications (e.g., monotone response to profit coefficient or revealed-preference consistency).
- Keep the transaction-construction scaffold fixed and isolate the intended policy variable. Offer schemas, parsers, choosers, and other protocol scaffolds can generate large apparent policy effects if they vary with treatment.
- Treat stochastic generations correctly: replicate within profile-condition, nest repetitions, and avoid treating single stochastic rollouts as independent observations. Report generation variance and perform profile-level inference.
- Jointly report completion rate, buyer surplus, seller profit, and welfare from ground truth. Buyer-surplus gains alone can be pure transfers, not welfare creation.
- Use scripted positive/negative controls to probe whether metrics and protocols would detect expected effects under known seller technologies. This helps separate metric failure from absence of the economic mechanism.
- Be cautious in extrapolating policy conclusions from A2A simulations without passing explicit validity checks (incentives, isolation, stability, accounting). When checks fail or are inconclusive, restrict claims to descriptive behavior rather than causal policy effects.
- Practical checklist for future studies: (1) Define behavioral tests for roles, (2) Fix protocol scaffolds, (3) pre-register/nest replications per profile-condition, (4) perform manipulation checks and scripted controls, (5) report full welfare accounting and variance decomposition.
Summary takeaway: Market-simulation outcomes with LLM agents can be dominated by scaffolding choices and generation noise. Robust policy claims require explicit construct-validity tests, nested replication, and positive controls to identify whether observed welfare changes reflect real economic mechanisms or implementation artefacts.
Assessment
Claims (9)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| Under the original asymmetric schema and buyer-choice scaffold, the mean Both-minus-None welfare contrasts were +87.4, +35.0, and +28.8 for Qwen2.5 1.5B, 3B, and 14B, respectively. Consumer Welfare | positive | Mean welfare difference between the Both-guardrail and None conditions |
Reading fidelity
high
Study strength
medium
|
n=30
+87.4, +35.0, and +28.8 welfare units for 1.5B, 3B, and 14B
|
| When the guarded and unguarded conditions use a unified offer schema and buyer chooser, the Both-minus-None welfare contrasts change to +7.2 for 1.5B, −13.9 for 3B, and +23.8 for 14B. Consumer Welfare | mixed | Welfare difference between Both and None guardrail conditions under the unified protocol |
Reading fidelity
high
Study strength
medium
|
n=30
+7.2 (95% CI [−8.1, 23.8]), −13.9 ([−26.2, −5.2]), and +23.8 ([−1.5, 56.6]) welfare units
|
| The original positive welfare result cannot be interpreted as a clean economic guardrail effect because the treatment conditions changed both the policy and the transaction-construction scaffold. Governance And Regulation | negative | Validity of the causal interpretation of the welfare estimate |
Reading fidelity
high
Study strength
high
|
n=30
Effect shrank by 92% at 1.5B and reversed sign at 3B
|
| For the four 14B profiles selected because they had the largest single-generation welfare effects, repeating each profile-condition three times reduced the mean Both-minus-None effect from +229 to +37.6 welfare units. Consumer Welfare | negative | Both-minus-None welfare effect across selected 14B profiles |
Reading fidelity
high
Study strength
medium
|
n=4
Single-generation mean +229; three-generation mean +37.6; 95% bootstrap interval [−34.2, 109.3]
|
| Generation residuals accounted for 49.9% of the sum-of-squares variation in the repeated-generation 14B probe. Consumer Welfare | negative | Variation in repeated welfare-effect measurements |
Reading fidelity
high
Study strength
medium
|
n=24
49.9% of sum-of-squares variation
|
| The stronger profit-pressure seller prompt increased rent-extraction questions but produced less seller profit and a lower profit share than the standard self-interest prompt. Organizational Efficiency | mixed | Seller profit, seller profit share, and number of rent-extraction questions |
Reading fidelity
high
Study strength
low
|
n=3
Rent probes 0.83 vs. 0.17; profit 33.8 vs. 57.5; profit share 0.37 vs. 0.56
|
| A scripted profit-maximizing seller achieved first-best welfare in the None condition, while applying both guardrails reduced welfare. Consumer Welfare | negative | Welfare and first-best attainment under Both versus None guardrails |
Reading fidelity
high
Study strength
high
|
n=60
Welfare decreased by 24.5 units; first-best attainment decreased by 15.4 percentage points; None welfare was 169.1 and 100% of first best
|
| When a scripted seller is forced to bundle optional components whose buyer value is below seller cost, both guardrails increase welfare and first-best attainment. Consumer Welfare | positive | Welfare and first-best attainment under Both versus None guardrails |
Reading fidelity
high
Study strength
high
|
n=60
Welfare increased by 18.7 units; first-best attainment increased by 11.1 percentage points
|
| The study classifies the original headline estimate as Invalid because the treatment changes the transaction scaffold, while the controlled study remains Inconclusive because incentive validity and generation stability are unresolved. Governance And Regulation | null_result | Admissibility and validity of the causal policy interpretation |
Reading fidelity
high
Study strength
high
|
MDE 180.9 versus repeated point estimate +37.6; exact sign-flip p = .50; generation residuals 49.9%
|