1 cumulative citations
View corpus contextFactSet's AI boosts the richness and speed of analyst reports—more sources, broader topics and advanced methods—but raises forecast errors by nearly 60% as more mixed signals become harder to synthesize, especially for overburdened analysts.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
We study how generative artificial intelligence (AI) transforms the work of financial analysts. Using the 2023 launch of FactSet's AI platform as a natural experiment, we find that adoption produces markedly richer and more comprehensive reports -- featuring 40% more distinct information sources, 34% broader topical coverage, and 25% greater use of advanced analytical methods -- while also improving timeliness. However, forecast errors rise by 59% as AI-assisted reports convey a more balanced mix of positive and negative information that is harder to synthesize, particularly for analysts facing heavier cognitive demands. Placebo tests using other data vendors confirm that these effects are unique to FactSet's AI integration. Overall, our findings reveal both the productivity gains and cognitive limits of generative AI in financial information production.
Summary
Main Finding
FactSet’s December 2023 launch of a domain-specific generative-AI platform (Mercury) caused analysts who used FactSet to produce materially richer and faster research but also substantially less precise forecasts. AI-assisted reports used many more distinct sources, covered broader topics, and applied more advanced methods; timeliness improved, yet forecast errors rose importantly. The rise in errors is best explained by higher information-synthesis costs (more balanced and conflicting signals), not by hallucination, boilerplate, or simple speed-over-review.
Key Points
- Treatment and effect sizes (post-Mercury, FactSet users vs. others):
- +40% more distinct information sources (relative to sample mean)
- +34% broader topical coverage
- +25% greater use of advanced analytical methods
- Figures showed the largest source gain (+61% relative to mean); industry and macro topics rose +48% and +41%
- Forecast issuance became timelier (22% of one standard deviation improvement)
- Forecast errors increased by ≈$0.44 (about +59% of sample average)
- Mechanism:
- Evidence rejects hallucination/low-quality-text and speed-over-review explanations.
- AI increased the balance of positive and negative signals; individually informative signals become harder to synthesize into single precise forecasts — a synthesis-cost channel.
- Forecast degradation concentrated for analysts with heavier cognitive loads (covering more firms, integrating more sources).
- Investors reacted less to AI-assisted reports, consistent with attenuated signal synthesis by market participants.
- Heterogeneous adoption:
- More likely among analysts with technical/IT backgrounds, elite education, higher forecast frequency, less prior covering experience, or larger teams.
- Robustness and falsification:
- Propensity-score matching + entropy balancing; difference-in-differences framework with month, broker, industry fixed effects and many controls.
- Placebo tests using Bloomberg, Refinitiv, Capital IQ as pseudo‑treatments → no effects.
- Results robust to alternative matchings, exclusion of individual brokerages, broker-year-month fixed effects, Poisson ML, and controls for report length/readability.
Data & Methods
- Data:
- 52,428 U.S. sell‑side equity research reports spanning 2022–2024.
- Report-level citations allowed identification of platform usage (explicit FactSet citations → treated reports).
- Measurement and extracted variables:
- Used GPT-4o-mini to parse full report text and embedded objects (tables, figures, appendices) to extract:
- Number/type of distinct information sources (textual, tabular, visual)
- Topical breadth (firm-, industry-, macro-level topics)
- Analytical methods used (historical analysis, valuation, forecasting, etc.)
- Outcome variables: forecast timeliness, forecast error (forecast accuracy), and market reaction measures.
- Used GPT-4o-mini to parse full report text and embedded objects (tables, figures, appendices) to extract:
- Identification strategy:
- Natural experiment: FactSet’s Mercury launch (Dec 14, 2023) as quasi-exogenous shock; Mercury provided AI at no extra cost to existing FactSet subscribers.
- Main econometric approach: difference-in-differences comparing FactSet users (treated) vs. users of other vendors (control), augmented by propensity-score matching and entropy balancing to construct balanced samples.
- Fixed effects for month, broker, industry; controls for firm, report, and analyst characteristics.
- Placebo tests with other vendors, multiple robustness checks (matching variants, broker-level exclusions, alternative estimators).
- Tests of mechanisms:
- Predictive power of positive/negative signals checked — remained individually predictive (arguing against information deterioration).
- Measures of textual redundancy/readability checked — no increase (arguing against boilerplate).
- Timing analysis: accuracy decline not concentrated among fastest reports (arguing against speed-over-review).
- Heterogeneity analysis to show synthesis cost is greater where cognitive burden is higher.
Implications for AI Economics
- Productivity paradox in knowledge work: GenAI can simultaneously expand information supply and processing capacity (more sources, topics, methods, faster issuance) while increasing cognitive frictions that reduce the precision of high-level judgments (forecast accuracy). Productivity gains are not purely monotone.
- Attention and information-processing constraints become central frictions when upstream AI raises the volume and balance of signals. Economic models of AI adoption should incorporate human synthesis costs and bounded-rational aggregation, not only data retrieval or prediction improvements.
- Market efficiency and asset-pricing effects: richer analyst outputs do not necessarily translate to clearer price discovery; investor reactions to AI-assisted reports attenuate, suggesting downstream welfare and efficiency effects may be ambiguous.
- Heterogeneous returns to AI investment:
- Human capital complementarities matter: analysts with stronger technical skills and team resources adopt earlier and likely extract more net benefit.
- Firms and brokers should consider upskilling and team structure adjustments to capture AI gains while mitigating synthesis costs.
- Platform design and policy:
- Tight integration of LLMs with audited, domain-specific data (as with Mercury) materially changes outputs; vendor design choices matter.
- Mitigants that reduce synthesis cost — e.g., AI summarization that produces conflict-resolving syntheses, clearer signal aggregation tools, decision aids for prioritization — could improve overall quality.
- Regulators and institutions should monitor vendor concentration and accountability (model lineage, provenance) because domain-specific GenAI materially alters information architecture in capital markets.
- Research and measurement implications:
- Future empirical and theoretical work on AI in economics should measure structural dimensions of information production (sources, topical breadth, methods) and incorporate human processing constraints.
- Generalizability caveats: findings are for sell-side equity research in the U.S. around a domain-specific platform rollout; other domains or more mature AI–human workflows may display different trade-offs.
- Aggregate welfare: net effects ambiguous. AI increases information access and speed (potentially welfare-enhancing) but can weaken decision precision and market responses unless coupled with tools or organizational changes that reduce synthesis costs.
Assessment
Claims (9)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| Adoption produces markedly richer and more comprehensive reports featuring 40% more distinct information sources. Output Quality | positive | number of distinct information sources cited in analyst reports |
Reading fidelity
high
Study strength
medium
|
40% more distinct information sources
|
| Adoption leads to 34% broader topical coverage in analyst reports. Output Quality | positive | topical coverage breadth of analyst reports |
Reading fidelity
high
Study strength
medium
|
34% broader topical coverage
|
| Adoption increases the use of advanced analytical methods in reports by 25%. Output Quality | positive | use of advanced analytical methods in analyst reports |
Reading fidelity
high
Study strength
medium
|
25% greater use of advanced analytical methods
|
| AI adoption improves the timeliness of analyst reports. Task Completion Time | positive | timeliness (release timing) of analyst reports |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Forecast errors rise by 59% following AI-assisted report adoption. Error Rate | negative | forecast error magnitude |
Reading fidelity
high
Study strength
high
|
59% increase in forecast errors
|
| AI-assisted reports convey a more balanced mix of positive and negative information that is harder to synthesize. Output Quality | mixed | balance of positive and negative information in reports and ease/difficulty of synthesizing information |
Reading fidelity
high
Study strength
medium
|
not reported
|
| The increase in forecast errors is particularly pronounced for analysts facing heavier cognitive demands. Error Rate | negative | forecast error magnitude (heterogeneous effect by analyst cognitive demand) |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Placebo tests using other data vendors confirm that these effects are unique to FactSet's AI integration. Adoption Rate | positive | presence/absence of the reported content, timeliness, and error effects in reports from other data vendors |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Overall, generative AI produces productivity gains (richer reports, timeliness) but also reveals cognitive limits (higher forecast errors) in financial information production. Organizational Efficiency | mixed | trade-off between productivity-related outputs (richness, timeliness) and error rates in forecasting |
Reading fidelity
high
Study strength
medium
|
not reported
|