Does AI-enriched analyst reports improve forecast accuracy?
When generative AI is integrated into analyst platforms, does the richer information and broader coverage it produces translate into better forecasts? This matters because output quality and decision accuracy may diverge.
"Generative AI for Analysts" measures what happens to sell-side research when a domain-specific AI platform is bundled into an existing data product. The authors treat the December 2023 introduction of MERCURY, FactSet's GenAI platform, as a plausibly exogenous change in AI access. Across 46,853 report-firm pairs issued by U.S. analysts, FactSet-associated reports after the integration show 26% more distinct information sources, 24% broader topical coverage and 21% more analytical methods, and the abstract adds that they became more timely. The decision-quality result runs the other way: "relative forecast accuracy declines when analysts face greater information-processing demands." A machine-learning benchmark fed the same observable inputs shows no analogous decline, which the authors read as "a human processing constraint rather than poorer underlying information." Every richness figure is the authors' own measurement. They used GPT-4o-mini to code each report's text, so the numbers are coded from the reports, not reported by FactSet or by the analysts.
The mechanism the excerpt gives is attention rather than capability. The authors write that AI "does not magically resolve attention constraints"; when workload and information demands are high, "AI-generated information compounds their cognitive burden." The same tool that enriches reports under routine conditions "can become a source of distraction when attention is strained." The paper also says the strain reaches investors, who show weaker market reactions to AI-assisted analyst reports. The treatment is defined narrowly too. The authors observe FactSet citations, not analysts' prompts or clicks, so the estimate is the "incremental effect of exposure to FACTSET after its GenAI integration" rather than the effect of verified MERCURY use. They note this makes the estimate conservative if control reports also use other GenAI tools.
Against the nearest notes, this excerpt is a case where the output that looks better is not the output that predicts better. Can self-ratings replace objective performance scores for AI competence? argues that self-report cannot stand in for performance; here the coded richness measures rise while the objective forecast-accuracy measure falls, which is the same warning in a different form. Where have workers actually delegated tasks to AI? also places AI use in information-intensive work, but it measures adoption from agent skills, whereas this excerpt measures what happens to output once a tool sits on proprietary data. The bottleneck argument echoes What makes accountable judgment scarce when AI cognition is cheap?: when first-pass information is cheap, the scarce input is the analyst's processing capacity. Does generative AI shift knowledge workers away from communication? reports an individual-documentation tilt from workplace trace data; this excerpt measures the reports themselves.
The excerpt does not establish the decision-quality result on its own. The forecast-accuracy specification, the benchmark design and the placebo tests that make a platform-wide trend unlikely appear only as claims in the abstract and introduction; the excerpt ends before any accuracy table. The investor-reaction finding is asserted in the related-work discussion without its evidence. The design covers one platform and one population of analysts, and "relative" in "relative forecast accuracy" is never defined in the text provided. So the strength the evidence allows is narrower than the title's wording. The study supports that report richness and forecast accuracy diverge under heavier demands, and it makes human processing a candidate explanation, because the benchmark did not decline. The "human constraint" reading is an inference from that contrast, which the authors state as "pointing to." It does not show that analysts would be more accurate with less information.
Inquiring lines that read this note 2
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
Does disclosing AI authorship change how audiences evaluate the writing? How do models learn from self-generated outputs without cascading failures?Related concepts in this collection 4
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Can self-ratings replace objective performance scores for AI competence?
Do people's perceptions of their own AI competence match what they can actually do? This matters because assessment systems might rely on the wrong type of measure to evaluate workplace readiness.
same divergence: an output or self-rated feature moves while objective performance does not follow
-
Where have workers actually delegated tasks to AI?
Existing AI-exposure measures predict where AI could work, not where workers have actually adopted it. This research asks which occupations have embedded AI into real workflows, and whether that pattern matches technical capability or conversational tool use.
both place AI use in information-intensive work; this one tracks outcomes after integration, not adoption
-
What makes accountable judgment scarce when AI cognition is cheap?
When AI systems can perform cognitive tasks cheaply and at scale, what human capabilities become most valuable? This explores whether judgment, verification, and accountability are the true bottlenecks in labor markets shaped by generative AI.
echoes the bottleneck argument with a measured case where cheap information leaves processing capacity scarce
-
Does generative AI shift knowledge workers away from communication?
When knowledge workers adopt generative AI heavily, do they spend proportionally more time on individual documentation and less on coordination with colleagues? Understanding this matters because it suggests AI may reshape not just productivity but the social fabric of how teams work together.
same individual-documentation tilt, seen in workplace trace data rather than in the reports
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Generative AI for Analysts
- GenAI as a Power Persuader: How Professionals Get Persuasion Bombed When They Attempt to Validate LLMs
- Generative AI at Work
- AI Now Writes as Many Online Articles as Humans
- The Short-Term Effects of Generative Artificial Intelligence on Employment: Evidence from an Online Labor Market
- Research: Gen AI Makes People More Productive—and Less Motivated
- Can AI Do Strategy?
- AI-Powered (Finance) Scholarship
Original note title
FactSet's GenAI integration enriched analyst reports but lowered forecast accuracy when processing demands rose — a human constraint, not poorer information