0 cumulative citations
View corpus contextGenerative AI accelerates and broadens enterprise decision-making but rarely translates into large-scale financial returns unless firms overhaul workflows and enforce human-in-the-loop governance; quick wins in speed and ideation coexist with persistent risks from hallucination, bias, and weak integration.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
Generative artificial intelligence (GenAI) is moving from a novelty confined to chatbots and content drafting into something enterprises are beginning to fold into how they actually decide things: pricing, hiring, supply chain routing, capital allocation. This paper examines that shift through the lens of decision intelligence, the discipline concerned with engineering better organizational decisions by combining data, models, and human judgment. Using a PRISMA-informed narrative review of academic and industry literature published mainly between 2019 and 2026, the paper traces how large language models and related generative systems are being embedded into enterprise decision workflows, what measurable value they are producing, and where they fall short. The review finds genuine opportunities: compressed analysis cycles, wider access to sophisticated reasoning for non-specialist decision-makers, and new forms of scenario generation once reserved for expert analysts. At the same time, the literature converges on a stubborn set of challenges, including hallucinated or unreliable outputs, algorithmic bias, unclear governance accountability, and a persistent gap between pilot-stage enthusiasm and enterprise-level financial return. The paper argues that organizations capturing durable value are not necessarily those with the most advanced models, but those that have redesigned decision workflows, built human-in-the-loop verification into high-stakes processes, and treated GenAI as a collaborator rather than an oracle. It closes with practical implications and a short research agenda.
Summary
Main Finding
Generative AI (GenAI) is reshaping enterprise decision intelligence by compressing analysis cycles, widening access to analytic reasoning for non-specialists, and generating richer option sets upstream of formal decisions. However, durable economic value is concentrated in a small minority of organizations. The literature shows that model sophistication alone does not determine returns; organizational practices — workflow redesign, human-in-the-loop controls, governance, and measurement discipline — mediate whether GenAI use translates into measurable enterprise-level financial gains.
Key Points
-
Adoption vs. value gap
- Regular GenAI use rose rapidly (survey evidence: from ~2/3 of firms to ~9/10), but few firms have scaled GenAI into production; fewer than 1 in 10 report scaling and only ~5–6% report enterprise-level financial returns >5% of EBIT (McKinsey, 2025).
- A tiny subset of “high performers” capture outsized returns (e.g., 46 firms among 876 in McKinsey’s sample reporting >10% EBIT attributable to AI).
-
Principal opportunities
- Compressed analysis cycles: faster synthesis of unstructured information reduces decision latency.
- Democratized reasoning: conversational interfaces lower the technical barrier for managers to engage directly with data-grounded analysis.
- Scenario and option generation: GenAI expands upstream ideation and counterfactual framing, changing the menu of strategic options.
- Productivity and workforce effects: large task-level exposure (Eloundou et al., 2023)—~80% of U.S. workforce have at least 10% of tasks affected; ~20% could see over half of tasks affected—implies reallocation of time toward higher-judgment activities.
- Macro sizing: consultancy estimates of aggregate annual value range roughly $2.6–$4.4 trillion across many use cases.
-
Core challenges and risks
- Hallucination/unreliable outputs: fluent but fabricated content remains a key technical risk; grounding and guardrails reduce but do not eliminate the problem (Airia, 2026).
- Bias and fairness: model and data biases persist and can propagate into organizational decisions.
- Governance and accountability: enterprise governance is often underdeveloped relative to adoption; unclear accountability for AI-driven decisions is a significant barrier.
- Integration and workflow fit: value accrues to firms that integrate GenAI into decision architectures (DS layer) and redesign workflows, not simply those with better models.
- Pilot-to-scale gap: many pilots do not translate into sustained value capture; organizational practices explain differences more than vendor/model choice.
-
Human–AI interaction findings
- Observational and experimental studies report a “Human+” pattern: users alternate between eliciting and steering AI outputs and retain decisive judgment especially under high uncertainty.
- Distrust of unchallengeable outputs leads users to prefer systems that support interrogation and verification.
Data & Methods
- Review type: PRISMA-informed narrative synthesis (structured selection and screening but not quantitative meta-analysis).
- Time window: literature published 2019–mid-2026 (captures pre- and post-ChatGPT shifts).
- Search scope: academic databases (Scopus, Web of Science, arXiv, ACM/IEEE) and industry reports (McKinsey, Deloitte, Gartner); search terms combined GenAI/LLMs with decision-making/enterprise terms.
- Screening and inclusion criteria:
- Must address GenAI/LLMs specifically;
- Must engage organizational/enterprise decision-making;
- Must be verifiable (publication/preprint).
- Final corpus: 16 core sources (peer-reviewed articles, PRISMA-guided systematic review, conference papers, arXiv, SSRN, industry surveys). Key referenced works include Duan et al. (2019), Feuerriegel et al. (2024), Eloundou et al. (2023), López-Solís et al. (2025), Bastone et al. (2026), McKinsey & Company (2025), and others.
- Limitations: modest number of core sources; heterogeneous evidence (qualitative case studies, experiments, industry surveys); findings are thematic and conceptual rather than pooled effect estimates.
Implications for AI Economics
-
Modeling returns to GenAI must include organizational frictions and complementarities
- Economic models of AI productivity should treat workflow redesign, governance, and managerial practices as mediators — not just model quality or task automation potential.
- Returns are heterogeneous: firm-level complementarities (skill mix, leadership, measurement) explain much of cross-firm variance in ROI.
-
Measurement and valuation
- Standard productivity accounting should incorporate decision-latency gains, option-value effects from expanded scenario generation, and risk-adjusted expected losses from hallucination/failure modes.
- Valuation of AI investments should include costs for governance, human-in-the-loop verification, integration into decision systems, and ongoing monitoring.
-
Labor and distributional effects
- Task-level exposure implies substantial reallocation rather than simple displacement: time saved on synthesis/drafting may raise demand for higher-order judgment tasks.
- Policy/firm-level planning should anticipate retraining, new role definitions (AI verifiers, prompt engineers), and potential short-run adjustment costs.
-
Risk, externalities, and policy
- Hallucination and biased outputs create negative externalities (misinformation cascades, regulatory non-compliance risk) that justify firm investment in guardrails and may motivate sectoral regulation or standards.
- Governance failure risks can have systemic consequences in high-stakes domains (finance, healthcare, legal) — economic models should factor in regulatory uncertainty and compliance costs.
-
Research priorities for AI economics
- Empirical estimates of risk-adjusted ROI across sectors and firm types, distinguishing pilots vs. scaled deployments.
- Microdata studies linking specific organizational practices (workflow redesign, metrics, leadership) to realized financial outcomes.
- Measurement frameworks for decision-quality improvements (not just time savings) and their translation into firm value.
- Labor market analyses of reallocation effects, wage dynamics for AI-complementary tasks, and firm-level hiring/training responses.
- Modeling multi-agent and systemic impacts where erroneous GenAI-driven decisions can propagate across markets.
Practical takeaway for economists and managers: when assessing GenAI’s economic value, explicitly model—and where possible measure—the organizational investments (integration, governance, human verification) required to convert technical capability into durable financial return.
Assessment
Claims (13)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| Regular generative AI use inside organizations increased from roughly two-thirds of firms to nearly nine in ten firms reporting AI use in at least one business function within about eighteen months. Adoption Rate | positive | Organizational use of generative AI in at least one business function |
Reading fidelity
high
Study strength
medium
|
increase from roughly two-thirds to nearly nine in ten firms
|
| Fewer than one in ten organizations reported scaling AI agents into full production. Adoption Rate | negative | Scaling of AI agents into full production |
Reading fidelity
high
Study strength
medium
|
fewer than 10% of organizations
|
| Only approximately 5% to 6% of organizations reported measurable enterprise-level financial returns exceeding 5% of EBIT from AI. Firm Productivity | negative | Enterprise-level financial return attributable to AI, measured relative to EBIT |
Reading fidelity
high
Study strength
medium
|
approximately 5%–6% of organizations reporting returns exceeding 5% of EBIT
|
| Around 80% of the U.S. workforce could have at least 10% of its tasks affected by large language models, while nearly 20% of workers might have more than half of their tasks affected. Automation Exposure | positive | Share of workers and tasks exposed to potential LLM effects |
Reading fidelity
high
Study strength
medium
|
80% with at least 10% of tasks affected; nearly 20% with over half of tasks affected
|
| In a systematic review of generative AI and strategic decision-making in entrepreneurial initiatives, 27 of 30 synthesized studies reported improvements in operational efficiency attributable to GenAI. Organizational Efficiency | positive | Operational efficiency associated with generative AI use |
Reading fidelity
high
Study strength
medium
|
n=30
27 of 30 studies reported improvements
|
| The same systematic review found that human judgment remained decisive under high uncertainty, while several included studies identified bias and diminished trust as unresolved concerns even when efficiency gains were present. Decision Quality | mixed | Role of human judgment, perceived trust, and bias in AI-assisted strategic decision-making |
Reading fidelity
high
Study strength
medium
|
n=30
|
| Among 876 companies providing sufficiently detailed data, 46 high performers attributed more than 10% of EBIT to AI deployment and reported returns exceeding $10 for every $1 invested, approximately three times the average return across the full sample. Firm Productivity | positive | AI-attributable EBIT and return on AI investment |
Reading fidelity
high
Study strength
medium
|
n=876
46 high performers; more than 10% of EBIT; exceeding $10 return per $1 invested; roughly three times the average return
|
| Workflow redesign, senior leadership engagement, and measurement discipline were more strongly associated with enterprise AI value capture than the particular vendor model used. Organizational Efficiency | positive | Enterprise value capture from generative AI |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Across the reviewed literature, generative AI consistently compressed analysis cycles by reducing the time required to process unstructured information and produce readable summaries for strategic decisions. Task Completion Time | positive | Time required to synthesize information and conduct due diligence on strategic options |
Reading fidelity
high
Study strength
medium
|
n=30
|
| Generative AI enabled non-specialist participants without prior entrepreneurial or advanced analytical experience to interrogate the viability of a business idea through natural-language dialogue. Decision Quality | positive | Ability of non-specialists to engage in data-intensive business reasoning |
Reading fidelity
high
Study strength
low
|
not reported
|
| Generative AI can expand the set of strategic options considered by decision-makers by producing alternative framings, counterfactual scenarios, and draft strategic options before formal evaluation. Decision Quality | positive | Breadth of strategic options and scenarios considered |
Reading fidelity
high
Study strength
low
|
not reported
|
| The literature identifies hallucination—fluent, confident, and sometimes fabricated outputs—as a major limitation for enterprise decision intelligence because it can cause decision-makers to act on false information. Ai Safety And Ethics | negative | Reliability and factual accuracy of AI-generated decision-support outputs |
Reading fidelity
high
Study strength
medium
|
not reported
|
| McKinsey & Company (2025) estimated the aggregate annual opportunity from generative AI at $2.6 trillion to $4.4 trillion across 63 identified use cases. Firm Productivity | positive | Estimated aggregate annual economic value from generative AI use cases |
Reading fidelity
high
Study strength
low
|
n=63
$2.6 to $4.4 trillion in annual value across 63 use cases
|