The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

FactSet's AI boosts the richness and speed of analyst reports—more sources, broader topics and advanced methods—but raises forecast errors by nearly 60% as more mixed signals become harder to synthesize, especially for overburdened analysts.

Generative AI for Analysts
Jian Xue, Qian Zhang, Wu Zhu · December 12, 2025
arxiv quasi_experimental medium evidence 8/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

Arxiv

Latest observation:

  1. Jian Xue unresolved corpus identity
  2. Qian Zhang unresolved corpus identity
  3. Wu Zhu unresolved corpus identity

Semantic Scholar

Latest observation:

  1. Jian Xue provider ID
  2. Qian Zhang provider ID
  3. Wufang Zhu provider ID
Using FactSet's 2023 AI launch as a natural experiment, the paper finds that AI adoption made analyst reports richer, broader, and timelier but increased forecast errors by 59% because AI-produced outputs conveyed more balanced and harder-to-synthesize information, with larger error increases for analysts under heavier cognitive load.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

We study how generative artificial intelligence (AI) transforms the work of financial analysts. Using the 2023 launch of FactSet's AI platform as a natural experiment, we find that adoption produces markedly richer and more comprehensive reports -- featuring 40% more distinct information sources, 34% broader topical coverage, and 25% greater use of advanced analytical methods -- while also improving timeliness. However, forecast errors rise by 59% as AI-assisted reports convey a more balanced mix of positive and negative information that is harder to synthesize, particularly for analysts facing heavier cognitive demands. Placebo tests using other data vendors confirm that these effects are unique to FactSet's AI integration. Overall, our findings reveal both the productivity gains and cognitive limits of generative AI in financial information production.

Summary

Main Finding

FactSet’s December 2023 launch of a domain-specific generative-AI platform (Mercury) caused analysts who used FactSet to produce materially richer and faster research but also substantially less precise forecasts. AI-assisted reports used many more distinct sources, covered broader topics, and applied more advanced methods; timeliness improved, yet forecast errors rose importantly. The rise in errors is best explained by higher information-synthesis costs (more balanced and conflicting signals), not by hallucination, boilerplate, or simple speed-over-review.

Key Points

  • Treatment and effect sizes (post-Mercury, FactSet users vs. others):
    • +40% more distinct information sources (relative to sample mean)
    • +34% broader topical coverage
    • +25% greater use of advanced analytical methods
    • Figures showed the largest source gain (+61% relative to mean); industry and macro topics rose +48% and +41%
    • Forecast issuance became timelier (22% of one standard deviation improvement)
    • Forecast errors increased by ≈$0.44 (about +59% of sample average)
  • Mechanism:
    • Evidence rejects hallucination/low-quality-text and speed-over-review explanations.
    • AI increased the balance of positive and negative signals; individually informative signals become harder to synthesize into single precise forecasts — a synthesis-cost channel.
    • Forecast degradation concentrated for analysts with heavier cognitive loads (covering more firms, integrating more sources).
    • Investors reacted less to AI-assisted reports, consistent with attenuated signal synthesis by market participants.
  • Heterogeneous adoption:
    • More likely among analysts with technical/IT backgrounds, elite education, higher forecast frequency, less prior covering experience, or larger teams.
  • Robustness and falsification:
    • Propensity-score matching + entropy balancing; difference-in-differences framework with month, broker, industry fixed effects and many controls.
    • Placebo tests using Bloomberg, Refinitiv, Capital IQ as pseudo‑treatments → no effects.
    • Results robust to alternative matchings, exclusion of individual brokerages, broker-year-month fixed effects, Poisson ML, and controls for report length/readability.

Data & Methods

  • Data:
    • 52,428 U.S. sell‑side equity research reports spanning 2022–2024.
    • Report-level citations allowed identification of platform usage (explicit FactSet citations → treated reports).
  • Measurement and extracted variables:
    • Used GPT-4o-mini to parse full report text and embedded objects (tables, figures, appendices) to extract:
      • Number/type of distinct information sources (textual, tabular, visual)
      • Topical breadth (firm-, industry-, macro-level topics)
      • Analytical methods used (historical analysis, valuation, forecasting, etc.)
    • Outcome variables: forecast timeliness, forecast error (forecast accuracy), and market reaction measures.
  • Identification strategy:
    • Natural experiment: FactSet’s Mercury launch (Dec 14, 2023) as quasi-exogenous shock; Mercury provided AI at no extra cost to existing FactSet subscribers.
    • Main econometric approach: difference-in-differences comparing FactSet users (treated) vs. users of other vendors (control), augmented by propensity-score matching and entropy balancing to construct balanced samples.
    • Fixed effects for month, broker, industry; controls for firm, report, and analyst characteristics.
    • Placebo tests with other vendors, multiple robustness checks (matching variants, broker-level exclusions, alternative estimators).
  • Tests of mechanisms:
    • Predictive power of positive/negative signals checked — remained individually predictive (arguing against information deterioration).
    • Measures of textual redundancy/readability checked — no increase (arguing against boilerplate).
    • Timing analysis: accuracy decline not concentrated among fastest reports (arguing against speed-over-review).
    • Heterogeneity analysis to show synthesis cost is greater where cognitive burden is higher.

Implications for AI Economics

  • Productivity paradox in knowledge work: GenAI can simultaneously expand information supply and processing capacity (more sources, topics, methods, faster issuance) while increasing cognitive frictions that reduce the precision of high-level judgments (forecast accuracy). Productivity gains are not purely monotone.
  • Attention and information-processing constraints become central frictions when upstream AI raises the volume and balance of signals. Economic models of AI adoption should incorporate human synthesis costs and bounded-rational aggregation, not only data retrieval or prediction improvements.
  • Market efficiency and asset-pricing effects: richer analyst outputs do not necessarily translate to clearer price discovery; investor reactions to AI-assisted reports attenuate, suggesting downstream welfare and efficiency effects may be ambiguous.
  • Heterogeneous returns to AI investment:
    • Human capital complementarities matter: analysts with stronger technical skills and team resources adopt earlier and likely extract more net benefit.
    • Firms and brokers should consider upskilling and team structure adjustments to capture AI gains while mitigating synthesis costs.
  • Platform design and policy:
    • Tight integration of LLMs with audited, domain-specific data (as with Mercury) materially changes outputs; vendor design choices matter.
    • Mitigants that reduce synthesis cost — e.g., AI summarization that produces conflict-resolving syntheses, clearer signal aggregation tools, decision aids for prioritization — could improve overall quality.
    • Regulators and institutions should monitor vendor concentration and accountability (model lineage, provenance) because domain-specific GenAI materially alters information architecture in capital markets.
  • Research and measurement implications:
    • Future empirical and theoretical work on AI in economics should measure structural dimensions of information production (sources, topical breadth, methods) and incorporate human processing constraints.
    • Generalizability caveats: findings are for sell-side equity research in the U.S. around a domain-specific platform rollout; other domains or more mature AI–human workflows may display different trade-offs.
  • Aggregate welfare: net effects ambiguous. AI increases information access and speed (potentially welfare-enhancing) but can weaken decision precision and market responses unless coupled with tools or organizational changes that reduce synthesis costs.

Assessment

Paper Typequasi_experimental Evidence Strengthmedium — The paper leverages a plausibly exogenous product launch and implements difference-in-differences/event-study comparisons plus placebo checks, which provide credible quasi-experimental variation; however, lack of random assignment leaves open potential selection or concurrent confounders (e.g., differential adoption timing, unobserved analyst or firm-level shocks), and short-term effects may not reflect longer-run adaptation. Methods Rigormedium — Multiple text-derived outcome measures, forecast-error analysis, and placebo/vendor comparisons suggest careful empirical work and robustness checks, but potential measurement error in text-coded variables, unobserved time-varying confounders, and limited detail on handling heterogeneous timing/selection reduce methodological certainty. SampleA dataset of professional financial analysts' research reports and forecasts around the 2023 launch of FactSet's AI platform, including metadata (analyst identity, timestamps, firm coverage), text of reports (used to measure number of information sources, topical breadth, analytical methods, and timeliness), forecast outcomes to compute errors, and a control sample of reports tied to other major data vendors for placebo comparisons. Themesproductivity human_ai_collab IdentificationUses the 2023 rollout of FactSet's AI platform as an exogenous shock and compares analyst outputs before and after the launch for FactSet users versus analysts tied to other data vendors (difference-in-differences / event-study style natural experiment), supplemented by placebo tests using other vendors to rule out concurrent industry-wide changes. GeneralizabilityFindings are specific to financial analysts and the institutional setting of sell-side/buy-side research and may not generalize to other occupations or industries., Results pertain to FactSet's specific AI integration and contemporaneous tool design; different generative-AI products could yield different effects., Short-term post-launch effects may not reflect longer-run adaptation, learning, or changes in incentives., Potential geographic or market-structure limits if sample is concentrated in particular regions or market segments., Measured outputs (reports and forecast errors) capture certain productivity dimensions but may miss other value-added tasks (client interactions, trading signals).

Claims (9)

ClaimDirectionOutcomeConfidence & EvidenceDetails
Adoption produces markedly richer and more comprehensive reports featuring 40% more distinct information sources. Output Quality positive number of distinct information sources cited in analyst reports
Reading fidelity high
Study strength medium
40% more distinct information sources
0.48
Adoption leads to 34% broader topical coverage in analyst reports. Output Quality positive topical coverage breadth of analyst reports
Reading fidelity high
Study strength medium
34% broader topical coverage
0.48
Adoption increases the use of advanced analytical methods in reports by 25%. Output Quality positive use of advanced analytical methods in analyst reports
Reading fidelity high
Study strength medium
25% greater use of advanced analytical methods
0.48
AI adoption improves the timeliness of analyst reports. Task Completion Time positive timeliness (release timing) of analyst reports
Reading fidelity high
Study strength medium
not reported
0.48
Forecast errors rise by 59% following AI-assisted report adoption. Error Rate negative forecast error magnitude
Reading fidelity high
Study strength high
59% increase in forecast errors
0.8
AI-assisted reports convey a more balanced mix of positive and negative information that is harder to synthesize. Output Quality mixed balance of positive and negative information in reports and ease/difficulty of synthesizing information
Reading fidelity high
Study strength medium
not reported
0.48
The increase in forecast errors is particularly pronounced for analysts facing heavier cognitive demands. Error Rate negative forecast error magnitude (heterogeneous effect by analyst cognitive demand)
Reading fidelity high
Study strength medium
not reported
0.48
Placebo tests using other data vendors confirm that these effects are unique to FactSet's AI integration. Adoption Rate positive presence/absence of the reported content, timeliness, and error effects in reports from other data vendors
Reading fidelity high
Study strength medium
not reported
0.48
Overall, generative AI produces productivity gains (richer reports, timeliness) but also reveals cognitive limits (higher forecast errors) in financial information production. Organizational Efficiency mixed trade-off between productivity-related outputs (richness, timeliness) and error rates in forecasting
Reading fidelity high
Study strength medium
not reported
0.48

Notes