0 cumulative citations
View corpus contextLanguage models now reach into browsers, APIs, simulations and robots, but scope has outpaced trustworthy delegation; engineers and policymakers should expand action authority only where provenance, failure detection, recovery, and calibrated human control are demonstrably supported.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
Large language models become consequential agents when surrounding systems let outputs change external state. Models now call tools, operate interfaces, delegate work, retain state, inhabit generated worlds, and control robots or laboratory equipment. Such advances are often narrated as one march toward autonomy, conflating model competence, system integration, persistence, and safe authority. This critical review synthesizes primary research and official technical specifications available by 31 August 2026. We organize the evidence along delegated authority, temporal persistence, and environmental coupling, while separating model, harness, and environment. Within the evidence examined, action-interface expansion is documented more convincingly than robust completion, recovery, authorization, or independent verification. Model Context Protocol and Agent2Agent improve interoperability but do not establish trustworthy delegation; multi-agent organization adds specialization alongside cost and correlated failure. Persistent simulations and world models support training and planning but do not themselves demonstrate agency; robotics and self-driving laboratories establish bounded feasibility rather than unattended open-world reliability. We propose justified delegation as an analytical and normative heuristic, not an observed law or certified score: expand action scope only where evidence supports provenance, bounded authority, failure detection, safe recovery, and calibrated human control. This framing yields a research agenda for coupled model-harness evaluation, capability-based permissions, durable state, cross-agent accountability, and staged physical validation.
Summary
Main Finding
Within the public evidence reviewed through 31 August 2026, language models have driven a clear and broad expansion of action interfaces (tools, APIs, UI automation, multi-agent protocols, world models, robot/instrument drivers), but that expansion outpaces credible evidence for verifiable, unattended autonomy. In short: we are better at letting models act; we are less well able to justify wide delegation of authority to them. The paper frames justified delegation (provenance, bounded authority, failure detection, safe recovery, calibrated human control) as the appropriate normative and evaluative heuristic.
Key Points
-
Conceptual framing
- Agency is a property of a configured system S = (M, H, E, U): model M, harness H (prompts, memory, tools, policies, verification), environment E (digital/social/virtual/physical), and delegator U (objective + authority).
- Three orthogonal dimensions matter: delegated authority (what state changes are permitted), temporal persistence (how long credentials/objectives/state survive), and environmental coupling (resettable simulator → live digital services → shared physical/social systems).
- Distinguishes three meanings of “better”: policy competence (model picks better actions under fixed harness), action coverage (harness exposes more interfaces/devices), and assurance (authorization, verification, containment, recovery).
-
Empirical synthesis
- Foundational research (ReAct, Toolformer, Reflexion, cognitive-architecture work) established the reasoning–action loop but does not by itself establish trustworthy external effects.
- Evidence categories are separated: peer‑reviewed benchmarks/systems, controlled physical experiments, preprints/technical reports, open specs, and first‑party previews. Each supports different claims and has distinct blind spots.
- Representative benchmark results (WebArena, OSWorld, SWE-agent, τ-bench, AgentBench) show capability gains but persistent reliability gaps relative to humans and brittleness under repeated trials or distribution shift.
- Multi-agent protocols (e.g., Agent2Agent, Model Context Protocol) increase interoperability and specialization but do not solve delegation trust; they introduce organizational costs and correlated-failure risk.
- Persistent simulations and world models aid training and planning but are not by themselves evidence of goal-directed agency in the real world.
- Robotics and automated laboratories demonstrate bounded feasibility (RT-2, Coscientist, ChemCrow, Gemini Robotics, etc.) but typically operate under narrow, resettable conditions and with human oversight; unattended open‑world reliability remains unproven.
-
Key normative claim
- "Justified delegation" should guide expansion of action scope: allow model actions only where provenance is available, authority is bounded, failures are detectable, recovery is safe, and human control is calibrated.
-
Research agenda (high level)
- Coupled model–harness evaluation (not model-only benchmarking).
- Capability-based permissions and principled authorization surfaces.
- Durable state and accountable provenance for long-lived tasks.
- Cross-agent accountability and verification protocols.
- Staged, physical validation pathways before granting higher authority in real-world settings.
Data & Methods
- Review type: critical, question-driven synthesis (not a systematic or quantitative meta‑analysis).
- Literature cutoff: 31 August 2026.
- Evidence selection principle: prioritize peer‑reviewed benchmarks and controlled experiments for capability claims; preprints and institutional reports for architectures/releases; open specifications for protocol semantics; first‑party system cards/previews only for existence, stated design, and disclosed failure modes.
- Evidence categories (and their interpretive boundaries):
- Peer‑reviewed benchmark/systems papers — measure task performance under published settings; blind to live credentials, longitudinal incidents.
- Peer‑reviewed controlled physical experiments — measure feasibility in disclosed apparatus; blind to long-term unattended operation and rare hazards.
- Preprints/technical reports — useful for architectures/artifacts; lack peer review/independent replication.
- Open technical specifications — define messages/roles/state transitions; blind to implementations and deployment reliability.
- First‑party previews/system cards — show existence and claimed constraints; blind to independent replication and full denominators.
- Analytical unit: configured action system S = (M, H, E, U). The paper evaluates systems across delegated authority, temporal persistence, and environmental coupling while keeping model, harness, and environment analytically separate.
Implications for AI Economics
-
Value creation vs. deployable automation
- Economic value from model advances will often accrue through better harnesses, protocols, and integrations (H and E) rather than model weights alone. Investment returns depend on system engineering (interface design, verification, credentials) as much as on model capability.
- Tasks with structured, typed interfaces (APIs, authenticated services, idempotent operations) are cheaper to automate credibly than tasks requiring high-fidelity perception, irreversible physical effects, or complex social judgments.
-
Labor markets and task composition
- Expect faster automation of digitally mediated, well‑specifiable tasks (data entry, API orchestration, code editing with test harnesses) than of open‑ended, high‑autonomy tasks (field robotics, unsupervised lab work).
- Multi-agent specialization enables decomposition of complex workflows but may increase coordination costs and create correlated failure modes; firms will trade off specialization gains against organizational and assurance costs.
-
Risk pricing, adoption, and investment timing
- Assurance gaps imply a risk premium on deploying agentic systems in live environments; firms will invest in monitoring, provenance, insurance, and human‑in‑the‑loop procedures. Under‑estimating these costs can produce overoptimistic ROI estimates based on capability benchmarks.
- Regulatory or procurement requirements that demand verifiable provenance, bounded authority, and recoverability will favor platforms that provide those features. Markets for third‑party verification, auditing, and policy enforcement services are likely to grow.
-
Platformization and market structure
- Open protocols (e.g., Model Context Protocol, agent-to-agent specs) can reduce integration costs and create multi‑sided platforms, but they also shift the locus of economic value toward governance, credential brokers, and assurance tooling.
- Firms that control robust harnesses (credential management, verification logs, safe sandboxes) may capture outsized rents compared with pure model providers.
-
Insurance, liability, and externalities
- Externalities from poorly constrained action interfaces (misinformation propagation, privacy leaks, physical harm) will create demand for liability frameworks and insurance products. Economic models of deployment should incorporate expected liability and remediation costs.
- Markets will form for capability‑based permissions (fine‑grained access control tied to evaluated competence) — these reduce systemic tail risk and influence pricing of agent services.
-
Measurement and forecasting recommendations for economists
- Do not infer deployable automation from benchmark/task‑completion scores alone; adjust forecasts for assurance, verification, and persistence gaps.
- Evaluate the marginal economic returns to investments in harness and verification infrastructure (H), not just model improvements (M).
- Model heterogeneity matters: tasks with high environmental coupling (shared physical/social systems) require different capital and organizational investments than low‑coupling digital automation.
-
Policy and organizational takeaways
- Procurement and regulation should require staged validation (sandbox → bounded live deployment → broader authority) and insist on externally verifiable provenance and recovery mechanisms before permitting high‑stakes delegation.
- Firms should adopt “capability‑based permissioning” internally: grant actions incrementally based on documented evidence and monitoring capability, rather than broad, opaque credentials.
Overall, the paper suggests that economic effects from agentic AI will be mediated strongly by system integration, governance, and verification choices. Economists and decision‑makers should treat action‑interface expansion and verifiable autonomy as separate variables when modeling adoption, productivity, labor impact, and regulatory responses.
Assessment
Claims (11)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| In the original WebArena study, the strongest reported GPT-4 agent completed 14.41% of tasks, compared with 78.24% for humans. Task Completion Time | negative | WebArena task completion rate |
Reading fidelity
high
Study strength
high
|
14.41% of tasks completed by GPT-4 versus 78.24% by humans
|
| In the original OSWorld study, the best reported agent configuration achieved below 12.2% performance, compared with 72.4% for humans. Task Completion Time | negative | OSWorld task performance |
Reading fidelity
high
Study strength
high
|
below 12.2% for the best agent configuration versus 72.4% for humans
|
| The evidence reviewed supports the conclusion that action-interface expansion has been documented more convincingly than robust completion, recovery, authorization, or independent verification. Organizational Efficiency | mixed | Relative evidence for action coverage versus reliable, verifiable autonomy |
Reading fidelity
high
Study strength
medium
|
not reported
|
| A longer benchmark task horizon supports a capability claim but does not establish durable autonomy. Task Completion Time | mixed | Task-completion horizon and sustained unattended operation |
Reading fidelity
high
Study strength
medium
|
not reported
|
| The SWE-agent agent-computer interface materially affected performance, with the original NeurIPS study reporting a 12.5% pass@1 rate on SWE-bench. Developer Productivity | positive | SWE-bench pass@1 rate |
Reading fidelity
high
Study strength
high
|
12.5% pass@1
|
| A benchmark score cannot be attributed to model weights alone when changing the interaction surface changes what the same class of model can accomplish. Developer Productivity | positive | Effect of interface design on agent task performance |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Model Context Protocol and Agent2Agent improve interoperability, but their specifications do not establish trustworthy delegation, implementation compliance, semantic correctness, or end-task reliability. Governance And Regulation | mixed | Interoperability and trustworthy delegation |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Multi-agent organization adds role specialization but also introduces additional cost and correlated failure; some reported gains may reflect additional inference budget rather than organization itself. Team Performance | mixed | Multi-agent performance, cost, and failure behavior |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Persistent simulations and world models can support training and planning, but they do not by themselves demonstrate agency. Other | null_result | Evidence of agency from persistent simulations and world models |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Robotics and self-driving laboratories establish bounded feasibility rather than unattended open-world reliability. Organizational Efficiency | mixed | Reliability of autonomous physical operation outside bounded procedures |
Reading fidelity
high
Study strength
medium
|
not reported
|
| The appropriate research objective is not to maximize action coverage in isolation, but to expand action coverage at a rate supported by competence and assurance. Governance And Regulation | positive | Safe and justified expansion of delegated action scope |
Reading fidelity
high
Study strength
speculative
|
not reported
|