0 cumulative citations
View corpus contextA new verifiable simulated marketplace lets buyer and merchant AI agents transact under an auditable protocol; benchmark results across ten models (capability-track means ~66–86%, large-catalog ~56–91%) reveal process-level errors that final-state checks can miss, highlighting bottlenecks exposed by large catalogs.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
In vibe coding, people describe software in natural language and delegate implementation to AI agents. By analogy, vibe commerce allows people to express buying or selling goals in natural language and delegate the corresponding tasks to agents. Commerce, however, requires independently controlled Buyer and Merchant agents to interact in a shared market while preserving their private objectives and distinct authority. We introduce Agentic Commerce World (ACWorld), an environment for evaluating such agents across ongoing transactions. Through its Vibe Commerce Protocol (VCP), ACWorld validates agent actions before updating shared transaction state and records the resulting interactions, making agent behavior auditable and evaluation reproducible. The ACWorld Benchmark contains a 200-task capability-coverage track and a 60-task large-catalog track that searches 785,022 transactable listings. Across ten models, mean scores range from 65.9% to 85.6% and from 56.1% to 91.4%, respectively. Our analysis shows that process-level evidence is necessary: final state alone can miss evaluated errors, incomplete trajectories still retain useful process signals, and large-catalog tasks expose bottlenecks across stages.
Summary
Main Finding
Agentic Commerce World (ACWORLD) is a stateful, auditable environment and benchmark that evaluates independently controlled Buyer and Merchant agents interacting in a shared market under a verification protocol (Vibe Commerce Protocol, VCP). By separating agent decisions from the authority to change shared state and recording validated commits with evidence, ACWORLD makes agent behavior reconstructable and rescoring reproducible. Empirical evaluation across two complementary tracks (capability-coverage and large-catalog) shows substantial variation across models (capability-coverage means ~65.9–85.6%; large-catalog means ~56.1–91.4%), and highlights that process-level evidence matters: final-world state alone can mask errors, intermediate execution traces provide useful learning signals, and large catalogs expose stage-specific bottlenecks (search, evidence, validation, execution).
Key Points
- Vibe commerce: people express buying/selling goals in natural language and delegate execution to Buyer and Merchant agents that retain private objectives and authority.
- Vibe Commerce Protocol (VCP): binds each proposed action to an authenticated actor and enforces four invariants:
- Each action is attributed to an authenticated actor.
- Messages cannot directly change World state.
- Every commit cites an accepted Commerce Intelligence Platform (CIP) validation.
- Reapplying ordered commits from the same initial state reproduces the final state and event digest.
- Architecture: Buyer agents, Merchant agents, Commerce Intelligence Platform (policy/validation layer), and a deterministic World (W_{j+1} = F(W_j, u_j)). Agents propose actions via VCP; CIP validates; World commits authorized effects.
- Execution traces: ACWORLD stores, for each decision t, a linked record Lt = (actor, decision, validation result, execution evidence, ∆World). Traces can be reconstructed and rescored deterministically.
- Scoring: Tasks declare r boolean/fractional predicates with weights; verified predicate values c_i ∈ [0,1] are aggregated as weighted average S = (Σ w_i c_i) / (Σ w_i). Full credit requires all required predicates; partial credit captures intermediate verified progress.
- Benchmark structure:
- Capability-coverage track: 200 tasks (116 Buyer, 84 Merchant), 10 families, 80 capabilities; uses 1,082 product snapshots.
- Large-catalog track: 60 tasks (45 Buyer, 15 Merchant), 4 families, 18 capabilities; searches 785,022 transactable listings drawn from ~791k source records.
- Interfaces/tools: role-specific skill manifests (Buyer 11, Merchant 25), shared typed business tools (55) and read-only World queries (18).
- Empirical results (high-level):
- Top-performing models reach ~83–86% on capability-coverage; large-catalog results vary more, with some models >90% on the track.
- Main sources of lost credit in capability-coverage runs without full credit are often evidence (grounding), choice (poor selection), and execution/authority (invalid or unauthorized actions).
- Key empirical observation: 99 of 861 runs that did not get full credit nonetheless reached a final World state also produced by a full-credit execution—i.e., matching final state can hide intermediate protocol violations or missing evidence.
Data & Methods
- Environment:
- ACWORLD implements VCP and provides tools for search, offers, evaluation, settlement, evidence collection, and governance.
- World persists inventory, orders, payments, trust records, receipts; commits are atomic and validated.
- Data:
- Capability-coverage uses curated snapshots (1,082) assembled into 200 executable test cases across 80 capabilities.
- Large-catalog uses a read-only catalog of 785,022 transactable listings (from ~791,431 public records), enabling global-oracle checks for search tasks.
- Task construction:
- Families cover transaction lifecycle and adjacent behaviors (Discovery, Grounding, Preference, Negotiation, Multi-item, Inventory/Lifecycle, Governance, Adversarial, Timing, etc.).
- Each capability is converted into multiple executable cases; reference policies constrained to role-visible info are used to guarantee solvability and to define expected behavior for deterministic rescoring.
- Evaluation protocol:
- Agents receive goals and role-limited observations, produce VCP messages/actions.
- CIP validates messages; approved events are committed to World and recorded with evidence.
- Evaluator reconstructs World from initial state and ordered commits to verify event digests and apply predicate checks.
- Aggregated predicate score S is reported per task; both full- and partial-credit signals are used.
- Models evaluated: ten models (examples in paper: GPT-5.6 Sol, Gemini 3.6 Flash, Gemini 3.5 Flash, Kimi K3, GPT-5.6 Terra, Claude Sonnet 5, Qwen3.5 Plus, etc.). Reported ranges: capability-coverage mean scores 65.9%–85.6%; large-catalog 56.1%–91.4% (paper gives precise per-model means and dominant failure sources).
- Reproducibility: commits and evidence permit deterministic reconstruction and rescoring without re-running the agents.
Implications for AI Economics
- Measuring agent-level market interactions:
- ACWORLD provides a principled way to attribute state changes to agents and to audit multi-agent transactions. This enables rigorous measurement of agent-level behaviors (pricing strategies, bargaining dynamics, information disclosure) in controlled experimental markets.
- Studying private information and strategic behavior:
- The setting models bilateral bargaining with private ceilings/floors and repeated interactions across a market. Researchers can study signaling, strategic misreporting, bargaining equilibria, and welfare consequences when agents follow ML-based policies.
- Market design and protocol effects:
- By separating decision from commit and validating through CIP, ACWORLD lets economists test how different protocol rules (validation strictness, evidence requirements, settlement timing, fee structures) affect outcomes: efficiency, latency, manipulation risk, and incentives to reveal truthful information.
- Auditability and accountability in automated commerce:
- The reconstructable trace model supports post hoc audit and liability assignment—important for regulatory compliance, dispute resolution, and trust mechanisms. Economists can quantify how auditability affects agent behavior, market trust, and transaction costs.
- Intermediate signals for learning and mechanism design:
- Verified intermediate predicates provide candidate rewards for RL-based agent training (e.g., reward shaping from evidence/validation signals). This opens avenues to study how learning dynamics interact with market mechanisms and whether process-level incentives lead to more robust market outcomes than terminal-only rewards.
- Empirical exploration of market frictions and bottlenecks:
- Large-catalog experiments show stage-specific bottlenecks (search/recall, evidence grounding, validation authority). Economists can use ACWORLD to measure how such frictions influence price dispersion, search costs, and consumer surplus when agents optimize automated procurement or selling.
- Policy and regulatory experiments:
- The framework enables controlled simulations of policy interventions (mandatory evidence standards, seller verification rules, dispute-resolution protocols) to assess trade-offs between fraud reduction, compliance costs, and market efficiency before real-world deployment.
- Limitations to bear in mind:
- “Verifiable” here means reconstructable under the declared ACWORLD contract, not cryptographic or production-grade security. The environment abstracts many real-world features (e.g., payment settlement latency, external shocks) and does not itself validate the learning utility of intermediate signals.
- Agents’ private reasoning is not observable; only their actions and the evidence they submit are recorded—this is realistic but constrains inference about internal incentives.
- Research directions enabled:
- Comparative studies of agent strategies (truthful vs. manipulative) and welfare implications.
- Mechanism design for automated marketplaces with ML agents (designing rules that induce desirable equilibria).
- Learning algorithms that incorporate validated process signals for improved strategic behavior.
- Empirical analysis of auditability’s impact on trust, fraud, and compliance costs.
Takeaway: ACWORLD gives a reproducible, auditable testbed to study the economics of automated commerce—how protocol design, evidence requirements, and agent policies jointly shape market outcomes—while providing traced, verifiable execution records that make it possible to evaluate process-level behaviors as well as terminal outcomes.
Assessment
Claims (8)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| The capability-coverage track contains 200 tasks spanning 10 commerce families and 80 capabilities. Other | positive | Breadth of evaluated agent commerce capabilities |
Reading fidelity
high
Study strength
high
|
n=200
200 tasks across 10 families and 80 capabilities
|
| The large-catalog track evaluates 60 tasks over 785,022 transactable listings. Other | positive | Scale of catalog-based agent evaluation |
Reading fidelity
high
Study strength
high
|
n=60
785,022 transactable listings
|
| Across ten evaluated models, mean scores on the capability-coverage track range from 65.9% to 85.6%, while mean scores on the large-catalog track range from 56.1% to 91.4%. Other | mixed | Agent benchmark score |
Reading fidelity
high
Study strength
medium
|
n=10
65.9% to 85.6% on capability coverage; 56.1% to 91.4% on large catalog
|
| Final transaction state alone can miss evaluated errors: 99 of 861 capability-coverage runs without full credit reached a state also produced by a full-credit execution of the same task. Error Rate | mixed | Ability of final-state evaluation to detect process errors |
Reading fidelity
high
Study strength
medium
|
n=861
99 of 861 runs
|
| ACWORLD's recorded commits support reconstruction and rescoring of transaction outcomes without another model call. Governance And Regulation | positive | Auditability and reproducibility of agent transaction evaluation |
Reading fidelity
high
Study strength
high
|
not reported
|
| VCP enforces four transaction-integrity invariants: each action belongs to an authenticated actor, messages cannot directly update the World, every commit cites accepted platform validation, and replaying ordered commits from the same initial state reproduces the final state and event digest. Ai Safety And Ethics | positive | Transaction integrity and auditability |
Reading fidelity
high
Study strength
high
|
four invariants
|
| Reference policies achieved full credit on all 60 large-catalog tasks. Other | positive | Task evaluation score |
Reading fidelity
high
Study strength
medium
|
n=60
full credit on all 60 tasks
|
| Making all scoring predicates equally weighted changes model means by at most 1.8 percentage points and preserves the model ordering. Other | null_result | Sensitivity of benchmark model means and rankings to scoring weights |
Reading fidelity
high
Study strength
medium
|
n=10
at most 1.8 percentage points
|