Can AI forecasters beat expert humans at venture evaluation?
Do frontier large language models outperform experienced managers and investors at predicting fundraising success? This matters because venture assessment is a genuinely uncertain, ill-structured judgment task where human expertise is assumed essential.
In a fully prospective tournament using Kickstarter campaigns launched after the training cutoffs of all tested models, frontier LLMs outperformed human forecasters at strategic foresight. Thirty live U.S. technology ventures were evaluated "while fundraising remained in progress and outcomes were unknown," and a battery of frontier and open-weight LLMs completed 870 pairwise comparisons predicting fundraising success. Benchmarked against 346 Prolific-recruited experienced managers and three MBA-trained investors working under monitored conditions, "human evaluators achieved rank correlations with actual outcomes between 0.04 and 0.45, while several frontier LLMs exceeded 0.60, with the best (Gemini 2.5 Pro) reaching 0.74 — correctly ordering nearly four of every five venture pairs." Neither wisdom-of-the-crowd ensembles of humans nor human-AI hybrid teams outperformed the best standalone model.
The paper frames the result as evidence for what it calls "unbounding rationality" (Csaszar 2025) — AI freed from the bounded-rationality constraints (limited information processing, inconsistent attention, computational limits) that decades of research (Simon 1947; Kahneman et al. 1982, 2016; Tetlock 2005) show degrade human judgment under uncertainty. The design is built specifically to rule out training-data leakage and lookup: campaigns post-date every model's training cutoff, a roughly 500-word anonymized summary strips the venture's name, the platform, funds raised to date, and consumer comments, and testing confirmed the anonymized summaries were difficult to identify via search. Outcomes had not yet been determined when forecasts were made and were pre-registered to Zenodo before results were known. The authors treat strategic foresight — judging "wicked," "ill-structured," irreducibly uncertain problems (Churchman 1967; Simon 1973; Knight 1921) — as categorically harder than domains like chess or protein folding, where ground truth is fixed and verifiable.
This sits alongside Can language models beat human venture capital experts?, but makes a stronger claim: VCBench's human benchmarks (tier-1 VC precision of 5.6%) were already near chance, so a modest LLM edge sufficed to win. Here the human range (0.04–0.45) is wider and sometimes substantial, yet the best LLM still clears it by a wide margin (0.74), and the fully prospective design forecloses the leakage objection a historical dataset like VCBench cannot. It also extends Can retrieval-augmented language models forecast like human experts? — where that system only neared the crowd aggregate, this tournament's LLMs surpass individual expert forecasters outright on a genuinely novel strategic-judgment task, without retrieval augmentation. The structure recalls Can machines learn to predict which research ideas will work?: both use pairwise-comparison tournaments to show an LLM beating domain experts, though this venture tournament finds off-the-shelf frontier models already ahead, with no fine-tuning or retrieval needed to win.
The excerpt does not establish why LLMs forecast better — whether from broader information synthesis, freedom from managerial overconfidence or anchoring, or some artifact of how the 500-word summaries were written and scored. The sample is narrow: 30 technology-category campaigns on one platform over one 72-hour window, evaluated by Prolific-recruited managers (not necessarily trained investors) alongside only three MBA investors. Whether the result generalizes to higher-stakes, longer-horizon strategic decisions — M&A, market entry, CEO succession — where outcomes take years to resolve and cannot be reduced to a single crowdfunding total, is untested here. The finding that hybrid human-AI teams and crowd ensembles underperform the best standalone model is reported but not mechanistically explained, which leaves open whether human input is actively harmful to blend in or merely redundant.
Inquiring lines that read this note 2
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
How do real-world evaluations reveal AI capabilities that benchmarks hide? What prevents LLMs from applying their reasoning knowledge to improve outputs?Related concepts in this collection 5
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Can language models beat human venture capital experts?
Explores whether LLMs can outperform top investors at founder success prediction in a domain where even experts show only modest accuracy. Tests whether AI forecasting is competitive in sparse-signal, high-uncertainty settings.
contrasts a near-chance human bar (VCBench) with this tournament's wider, sometimes-substantial human range that LLMs still clear
-
Can retrieval-augmented language models forecast like human experts?
Can language models augmented with search and reasoning match or exceed the forecasting accuracy of competitive human crowd forecasters on events beyond their training data? This tests whether AI can handle genuine predictive judgment.
extends nearing-the-crowd forecasting to outright surpassing individual expert forecasters, without retrieval augmentation
-
Can machines learn to predict which research ideas will work?
Can a fine-tuned language model with access to published papers predict which unimplemented AI ideas will succeed empirically, and would it outperform human researchers making the same judgment?
shares the pairwise-tournament-beats-experts structure, but needs no fine-tuning or retrieval to win
-
Does using LLMs actually improve strategic decision making?
An experiment tested whether LLM assistance changes how people think through strategic choices and whether those changes lead to better predictions. Understanding this matters because organizations increasingly rely on AI to augment human decision-making.
evidence for A's no-help-from-pooling finding: LLM use doesn't improve humans' own forecast accuracy, only reshapes mental models and raises overload
-
Do newer frontier LLMs actually make better strategic decisions?
A strategy simulation benchmark tests whether the latest LLMs can balance short-term profit against long-term growth investment—a core challenge in real strategic reasoning.
qualifies A's scope: in a strategy-simulation benchmark, frontier LLMs score below MBA students, unlike their edge in Kickstarter forecasting ranking
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- The Strategic Foresight of LLMs: Evidence from a Fully Prospective Venture Tournament
- How Well Can AI Do Strategy? Empirical Benchmarking Using Strategy Simulations
- AI-Augmented Strategic Decision-Making Under Time Constraints: An Experimental Study on Mental Representations and Strategic Foresight
- VCBench: Benchmarking LLMs in Venture Capital
- The Ideation-Execution Gap: Execution Outcomes of LLM-Generated versus Human Research Ideas
- Model Swarms: Collaborative Search to Adapt LLM Experts via Swarm Intelligence
- Large Language Models Often Know When They Are Being Evaluated
- FrontierChallenge: Evaluating Scientific Workflow Completion
Original note title
a fully prospective venture tournament finds LLMs outforecast human forecasters at strategic foresight — pooling humans and AI does not help