6 cumulative citations
View corpus contextLLM-based agents can perform many multi-step workplace tasks but still miss roughly 40% of assignments; failures cluster in tool use, planning and contextual inference, with weaker models failing basic steps and stronger ones stumbling on tasks requiring inference beyond explicit directions.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
The advancement of large language model (LLM) based agents has shifted AI evaluation from single-turn response assessment to multi-step task completion in interactive environments. We present an empirical study evaluating frontier AI models on 150 workplace tasks within a realistic e-commerce RL environment from Surge. Our analysis reveals an empirically-derived \emph{hierarchy of agentic capabilities} that models must master for real-world deployment: (1) tool use, (2) planning and goal formation, (3) adaptability, (4) groundedness, and (5) common-sense reasoning. Even the best-performing models fail approximately 40\% of the tasks, with failures clustering predictably along this hierarchy. Weaker models struggle with fundamental tool use and planning, whereas stronger models primarily fail on tasks requiring contextual inference beyond explicit instructions. We introduce a task-centric design methodology for RL environments that emphasizes diversity and domain expert contributions, provide detailed failure analysis, and discuss implications for agent development. Our findings suggest that while current frontier models can demonstrate coherent multi-step behavior, substantial capability gaps remain before achieving human-level task completion in realistic workplace settings.
Summary
Main Finding
Frontier LLM-based agents tested in a realistic e-commerce RL environment (CORECRAFT) exhibit a predictable hierarchy of capability failures. Even the best models (GPT-5.2, Claude Opus 4.5, Gemini 3 Pro) fail roughly 40% of 150 realistic workplace tasks. Failures cluster along five empirically-derived capability levels — tool use, planning & goal formation, adaptability, groundedness, and common-sense reasoning — with weaker models failing at lower levels and stronger models failing mainly on groundedness and contextual/common-sense inferences.
Key Points
- Evaluation scale and result: 150 fully autonomous customer-support style tasks in an e-commerce simulator; top models still fail ~40% of tasks (best-performing models ≈ 60% pass rate).
- Capability hierarchy (diagnostic):
- Tool use — correct API/tool invocation and argument mapping.
- Planning & goal formation — decomposing multi-step tasks and sequencing actions.
- Adaptability — recognizing and recovering from unexpected or empty results.
- Groundedness — maintaining state and avoiding hallucinations (e.g., temporal/state confusion).
- Common-sense reasoning — drawing contextually appropriate inferences beyond explicit instructions.
- Failure patterns:
- Weak models predominantly fail at Levels 1–2 (tool use, planning).
- Mid-tier models show mixed failures; targeted training yields measurable gains.
- Strong models mostly fail at Levels 4–5 (grounding and common-sense/contextual inference).
- Task design: tasks were realistic, designed by domain experts, and ranged from single-step lookups to complex multi-system workflows (e.g., product compatibility + cheapest fix).
- Practical observation: production agents’ practice of limiting autonomy (e.g., ~10 steps before human intervention) aligns with observed capability limitations and reliability risk.
Data & Methods
- Environment: CORECRAFT — a sandboxed RL environment simulating an online retailer with interconnected entities (customers, employees, product catalog, orders, tickets), a tool API implemented via a Model Context Protocol (MCP), telemetry, and task management.
- Task set: 150 fully autonomous customer-support tasks crafted by domain experts; success criteria encoded in prompts and rubrics to permit automated (LLM-judge) and human-evaluated scoring.
- Tool/interface: Structured MCP schemas for search/retrieve/create/update operations; agents could make unlimited tool calls.
- Models evaluated: broad set of frontier and legacy models (examples: GPT-5.2, GPT-5, GPT-4o, Claude Opus/Sonnet 4.5, Gemini 3 Pro/2.5 Pro, Nova 2/1 Pro, Kimi K2 Turbo, Qwen3-Max, Mistral Medium).
- Protocol: identical system prompt + tool schemas + task prompt per run; complete action trajectories logged; failures were categorized into the capability hierarchy via manual/LLM-based failure analysis.
- Key quantitative result: best model ~60% pass rate; failure modes concentrated by capability level as above.
Implications for AI Economics
- Realizable economic value is constrained by reliability:
- A ~40% failure rate on realistic workplace tasks means substantial oversight, error-handling, or limited autonomy is still required, reducing net automation gains and increasing operational costs.
- Organizations will likely continue hybrid human-agent workflows; human intervention and review remain necessary and costly.
- Heterogeneous returns by task complexity:
- Simple lookup and well-defined tool-invocation tasks (Level 1) are already tractable and can be automated with modest investment, delivering near-term productivity gains.
- Multi-system workflows that require planning or adaptability (Levels 2–3) yield intermediate returns if targeted system- and tool-level improvements are made.
- High-value tasks demanding grounding and contextual/common-sense reasoning (Levels 4–5) are the hardest barriers to full automation—these will limit displacement of higher-skill workers and preserve demand for human judgment.
- Deployment economics and product design:
- Firms will need to budget for monitoring, human-in-the-loop review, rollback/compensation procedures, and liability insurance when using agents in customer-facing workflows.
- Constraining agent autonomy (e.g., step caps, approval gates) is rational given failure clustering; this affects throughput and ROI calculations.
- Development prioritization for economic impact:
- Investing in improvements to Levels 1–3 (tooling, planner architectures, robustness to empty/dirty results) likely yields the largest near-term increase in usable automation.
- Marginal improvements addressing Levels 4–5 are more technically challenging but necessary for high-value, low-supervision applications; these improvements will unlock disproportionate economic value once achieved.
- Evaluation and procurement advice:
- Buyers and regulators should use task-centric, domain-expert-designed interactive benchmarks (like CORECRAFT) rather than static benchmarks to estimate deployed performance and costs.
- Economic models for adoption should incorporate empirical pass rates, expected human oversight time per failed task, and the cost of incorrect actions (reputation, refunds, regulatory fines).
- Labor market and policy considerations:
- The current capability profile favors augmentation (assistive agents for routine tasks) over wholesale substitution of skilled workers; retraining and role redesign will be central to capture productivity gains.
- Policymakers and firms should anticipate transitional costs, design standards for reliability testing, and rules for liability and disclosure given nontrivial failure risks.
Overall takeaway: frontier agents can perform coherent multi-step workplace behavior, but systematic capability gaps — especially in grounding and contextual/common-sense reasoning — meaningfully limit safe, fully autonomous deployment and should be explicitly accounted for in economic impact assessments and deployment plans.
Assessment
Claims (10)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| The advancement of large language model (LLM) based agents has shifted AI evaluation from single-turn response assessment to multi-step task completion in interactive environments. Other | positive | evaluation paradigm (single-turn vs multi-step task completion) |
Reading fidelity
high
Study strength
medium
|
not reported
|
| We present an empirical study evaluating frontier AI models on 150 workplace tasks within a realistic e-commerce RL environment from Surge. Other | positive | evaluation coverage (number of tasks evaluated) |
Reading fidelity
high
Study strength
high
|
n=150
|
| Our analysis reveals an empirically-derived hierarchy of agentic capabilities that models must master for real-world deployment: (1) tool use, (2) planning and goal formation, (3) adaptability, (4) groundedness, and (5) common-sense reasoning. Other | positive | capability taxonomy / ordering of capabilities |
Reading fidelity
high
Study strength
medium
|
n=150
|
| Even the best-performing models fail approximately 40% of the tasks. Error Rate | negative | task failure rate |
Reading fidelity
high
Study strength
medium
|
n=150
≈40% failure rate
|
| Failures cluster predictably along the hierarchy of agentic capabilities. Error Rate | negative | distribution of failures across capability categories |
Reading fidelity
high
Study strength
medium
|
n=150
|
| Weaker models struggle with fundamental tool use and planning. Error Rate | negative | task-specific failure rate for tool use and planning |
Reading fidelity
high
Study strength
medium
|
n=150
|
| Stronger models primarily fail on tasks requiring contextual inference beyond explicit instructions. Error Rate | negative | failure rate on contextual-inference tasks |
Reading fidelity
high
Study strength
medium
|
n=150
|
| We introduce a task-centric design methodology for RL environments that emphasizes diversity and domain expert contributions. Other | positive | design methodology introduction |
Reading fidelity
high
Study strength
low
|
not reported
|
| We provide detailed failure analysis and discuss implications for agent development. Other | positive | qualitative failure analysis |
Reading fidelity
high
Study strength
medium
|
n=150
|
| Current frontier models can demonstrate coherent multi-step behavior, but substantial capability gaps remain before achieving human-level task completion in realistic workplace settings. Error Rate | mixed | multi-step task completion ability vs gap to human-level performance |
Reading fidelity
high
Study strength
medium
|
n=150
≈40% failure rate (implying gaps vs human-level completion)
|