0 cumulative citations
View corpus contextGenerative AI helps managers make better decisions by synthesizing and surfacing information, but firms rarely convert adoption into clear profit gains because poor workflow integration and governance blunt value capture; ensembling models or multi-run evaluations reduce bias and better match expert judgments.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
0 cumulative citations
View corpus contextGenerative artificial intelligence (GenAI) has moved from experimental novelty to a central input in organizational decision-making, with McKinsey's Q1 2026 Global AI Survey finding that 65 percent of organizations now use generative AI in at least one business function, roughly double the adoption rate reported ten months earlier, and 72 percent report at least one AI workload in production. Despite this scale of adoption, the empirical record on business value remains sharply divided. This paper synthesizes recent academic and industry evidence, drawing on 24 sources published primarily between 2023 and 2026, to examine three interlocking questions: what theoretical frameworks currently explain GenAI's role in managerial and strategic decision-making, how enterprises are applying GenAI in practice across functional areas, and what structural challenges limit the translation of GenAI adoption into measurable business value. The paper synthesizes evidence from strategic management research on AI-assisted evaluation of business alternatives, organizational theory on GenAI's emerging roles in decision processes, and empirical field studies on ambiguity handling and sycophantic behavior in AI-generated business advice, alongside a widely cited 2025 MIT study finding that 95 percent of enterprise generative AI pilots fail to deliver measurable profit-and-loss impact. Findings indicate that GenAI functions most reliably as an augmentation tool that aggregates and structures diverse inputs for human judgment, rather than as an autonomous decision-maker, that single-model evaluations of strategic alternatives are frequently inconsistent and biased while aggregated multi-model evaluations approximate expert human judgment, and that the primary barrier to enterprise value is organizational and workflow integration rather than model capability. The paper concludes with a proposed decision-integration framework and implications for executives, AI governance functions, and researchers.
Summary
Main Finding
Generative AI (GenAI) is delivering most value as a human-augmentation technology that aggregates, structures, and surfaces information for managerial judgment rather than as an autonomous strategic decision-maker. Although adoption has scaled rapidly (e.g., McKinsey Q1 2026: 65% of organizations use GenAI in ≥1 business function; 72% report ≥1 AI workload in production), measurable P&L impact is elusive — driven less by model capability and more by failures of organizational and workflow integration. Single-model strategic evaluations are often inconsistent and biased; aggregating multiple models or model runs yields evaluations that better approximate expert human judgment.
Key Points
- Adoption vs value gap: Adoption has surged, yet empirical evidence on firm-level profit impact is mixed. A widely cited 2025 MIT study reports 95% of enterprise GenAI pilots fail to deliver measurable P&L effects.
- Best-supported role: GenAI reliably augments human decision-making by structuring inputs, surfacing alternatives, and synthesizing evidence rather than autonomously deciding.
- Evaluation behavior:
- Single-model outputs for strategic choices are frequently inconsistent and prone to bias.
- Aggregated multi-model or multi-run evaluations reduce variance and bias, approximating expert judgments more closely.
- Behavioral risks: Field studies document GenAI weaknesses in handling ambiguity and a tendency toward sycophantic/adaptive responses that can mislead decision-makers.
- Primary bottleneck: Organizational integration — embedding GenAI into decision workflows, governance, incentives, and change management — is the main obstacle to translating adoption into measurable business value, more so than current model limitations.
- Practical recommendation: Firms see larger gains from investments in decision integration, model ensembling, and governance than from marginally better model capabilities alone.
Data & Methods
- Evidence base: Synthesis of 24 sources published primarily between 2023–2026, combining academic literature and industry reports.
- Key included sources/types (as cited in paper):
- Large-scale surveys (e.g., McKinsey Q1 2026 Global AI Survey) documenting adoption and productionization rates.
- The 2025 MIT enterprise study reporting pilot outcomes and P&L impact.
- Strategic management research on AI-assisted evaluation of alternatives.
- Organizational theory work on decision processes and technology-mediated cognition.
- Empirical field studies and controlled experiments examining ambiguity handling, sycophancy, and comparative model performance.
- Methods of synthesis:
- Cross-study comparative analysis to identify consistent patterns (e.g., augmentation role, aggregation benefits).
- Triangulation across quantitative surveys, qualitative case studies, and field experiments.
- Development of a conceptual decision-integration framework grounded in the empirical regularities.
Implications for AI Economics
- Measurement and valuation
- Rethink metrics: Move beyond adoption and “AI in production” counts to decision-quality and workflow-integration metrics (e.g., time-to-decision improvement, error-reduction in judged outcomes, realized P&L per integrated workflow).
- Attribution challenges: Economists should model and empirically separate model contribution from organizational implementation effects when estimating returns to GenAI.
- Models of diffusion and returns
- Incorporate organizational frictions (integration costs, governance, managerial incentives, training) into models of GenAI diffusion and firm-level productivity gains.
- Estimate the relative returns to improving models versus reducing integration frictions; early evidence suggests diminishing returns to model-only improvements absent integration.
- Policy and governance
- Regulation and guidance should focus not only on model safety but also on deployment governance (decision thresholds, human-in-loop requirements, audit trails) to reduce sycophancy and miscalibration risks.
- Managerial decisions and investment priorities
- Firms should prioritize: (1) workflow redesign and change management, (2) ensemble/aggregation approaches for strategic evaluation, (3) human-in-loop interfaces and calibration training, and (4) governance and measurement systems to capture decision-level impacts.
- Research agenda
- Causal studies measuring P&L impact conditional on integration investments (randomized or quasi-experimental).
- Structural estimation of the cost-benefit frontier: integration spending vs model improvement.
- Longitudinal firm-level analyses to capture learning effects and path dependence in GenAI value capture.
- Microdata studies of decision outcomes comparing single-model vs aggregated-model interventions.
- Behavioral studies quantifying harms from sycophancy and ambiguity, and testing mitigation strategies (e.g., model uncertainty quantification, adversarial prompts).
Suggested operationalization (decision-integration framework summary) - Inputs: Data aggregation, provenance and quality controls. - Model layer: Multi-model ensembles and structured multi-run evaluations to reduce bias/variance. - Human layer: Clear roles, decision-preservation of human judgment, calibration training, and feedback loops. - Workflow embedding: Integration into existing processes with mapped incentives and measurable KPIs. - Governance & measurement: Audit trails, outcome attribution methods, and continuous evaluation of decision-quality and economic impact.
Overall, the economics of GenAI value creation requires shifting focus from model performance improvements alone to the institutional and transactional factors that enable firms to capture generated value.
Assessment
Claims (9)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| By Q1 2026, 65% of organizations used generative AI in at least one business function. Adoption Rate | positive | Organizational adoption of generative AI |
Reading fidelity
high
Study strength
medium
|
65%
|
| By Q1 2026, 72% of organizations reported at least one AI workload in production. Adoption Rate | positive | Productionization of AI workloads |
Reading fidelity
high
Study strength
medium
|
72%
|
| Enterprise generative-AI pilots frequently fail to produce measurable profit-and-loss effects; a cited 2025 MIT study reports that 95% of pilots failed to deliver measurable P&L effects. Firm Revenue | negative | Measurable enterprise profit-and-loss impact from generative-AI pilots |
Reading fidelity
high
Study strength
low
|
95% failed to deliver measurable P&L effects
|
| Generative AI currently delivers most value as a human-augmentation technology that aggregates, structures, and surfaces information for managerial judgment rather than acting as an autonomous strategic decision-maker. Decision Quality | positive | Support for human managerial judgment and decision-making |
Reading fidelity
high
Study strength
medium
|
n=24
|
| Single-model outputs for strategic choices are frequently inconsistent and prone to bias. Decision Quality | negative | Consistency and bias of strategic evaluations |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Aggregating multiple models or multiple model runs reduces evaluation variance and bias and produces judgments that more closely approximate expert human judgments. Decision Quality | positive | Accuracy, variance, and bias of strategic evaluations relative to expert judgment |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Field studies document that generative AI performs poorly in handling ambiguity and can produce sycophantic or adaptive responses that mislead decision-makers. Ai Safety And Ethics | negative | Reliability of AI-assisted decisions under ambiguity and susceptibility to misleading responses |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Organizational and workflow integration—including decision workflows, governance, incentives, and change management—is the primary obstacle to converting generative-AI adoption into measurable business value. Organizational Efficiency | negative | Conversion of AI adoption into measurable business value |
Reading fidelity
high
Study strength
medium
|
n=24
|
| Early evidence suggests that marginal improvements in model capabilities produce diminishing returns when organizational integration is absent. Firm Productivity | negative | Returns to model capability improvements conditional on organizational integration |
Reading fidelity
medium
Study strength
low
|
not reported
|