The Strategic Foresight of LLMs: Evidence from a Fully Prospective Venture Tournament

Paper · arXiv 2602.01684 · Published February 2, 2026
AI at Work

Abstract Can artificial intelligence outperform humans at strategic foresight—the capacity to form accurate judgments about uncertain, high-stakes outcomes before they unfold? We address this question through a fully prospective prediction tournament using live Kickstarter crowdfunding projects. Thirty U.S.-based technology ventures, launched after the training cutoffs of all models studied, were evaluated while fundraising remained in progress and outcomes were unknown. A diverse suite of frontier and open-weight large language models (LLMs) completed 870 pairwise comparisons, producing complete rankings of predicted fundraising success. We benchmarked these forecasts against 346 experienced managers recruited via Prolific and three MBA-trained investors working under monitored conditions. The results are striking: human evaluators achieved rank correlations with actual outcomes between 0.04 and 0.45, while several frontier LLMs exceeded 0.60, with the best (Gemini 2.5 Pro) reaching 0.74—correctly ordering nearly four of every five venture pairs. These differences persist across multiple performance metrics and robustness checks. Neither wisdom-of-the-crowd ensembles nor human-AI hybrid teams outperformed the best standalone model.

Keywords: artificial intelligence; strategic foresight; strategic decision-making; opportunity evaluation; strategic uncertainty

Introduction. 1.1 Can AI Make Strategic Predictions?

In 1943, over lunch at Bell Labs, Alan Turing and Claude Shannon debated what kind of artificial intelligence they hoped to build. Shannon envisioned a machine of vast intellectual power—what we might now call an AI scientist. Turing countered with a different ambition.

“No, I’m not interested in developing a powerful brain,” he quipped. “All I am after is just a mediocre brain, something like the President of the American Telephone and Telegraph Company” (Hodges 1983:251). He imagined a machine that could digest facts about commodity prices and stocks, then answer the question: “Do I buy or sell?” In modern terms, Turing was describing an AI CEO.

Eight decades later, the asymmetry between these two visions is stark. AI scientists now autonomously discover physical laws, design proteins, and generate novel hypotheses (Jumper et al. 2021, Fang et al. 2025, Ghafarollahi and Buehler 2025). Yet the AI CEO remains unrealized. Sam Altman captures this gap in a thought experiment: “What would have to happen for an AI CEO to be able to do a much, much better job of running OpenAI than me, which clearly will happen someday. But how can we accelerate that? What’s in the way of that?” (Altman 2025). Turing’s question has thus become a frontier challenge of our time: whether machines can navigate the deep and often irreducible uncertainties of strategic choice.

Central to this challenge is strategic foresight: the capacity to form accurate, forwardlooking judgments about uncertain, high-stakes business outcomes before they unfold. Unlike chess or protein folding, where objective functions are fixed and solutions can be verified against independent ground truth, strategic decisions unfold in environments characterized by ambiguity, incomplete information, and endogenous change. Decision makers must anticipate technological shocks, interpret noisy signals about customer adoption, and act on assessments that cannot be validated until long after the moment of choice. Strategic decision making is thus an archetype of what scholars have called wicked (Churchman 1967), ill-structured (Simon 1973), or irreducibly uncertain (Knight 1921) problems—domains where prediction requires judgment rather than deduction.

Strategic foresight matters precisely because it is so difficult. Many core strategy theories hold that above-normal profits and competitive advantage derive from a superior ability to predict the future value of resources, the attractiveness of industries, and value-creation opportunities (Csaszar and Laureiro-Mart ́ınez 2018). Since strategic decisions pay off in the future, and since advantage stems from either luck or foresight, only foresight is controllable by the firm. Improving foresight thus gives managers practical levers to shape decision quality and, by extension, firm performance. Empirical evidence supports this logic: firms exhibiting superior foresight display systematically higher productivity and performance (e.g., Bloom et al. 2026).

Yet human foresight is notoriously limited. Decades of research show that managers are boundedly rational: they are limited in information processing, constrained in computational capacity, and inconsistent in attention (Simon 1947, Kahneman et al. 1982, 2016). Even seasoned experts often perform no better than chance (Tetlock 2005). Efforts to improve foresight through broader cognitive representations, teaming, and structured updating have shown promise (Tetlock and Gardner 2015, Csaszar and Laureiro-Mart ́ınez 2018, Csaszar and Rhee 2026, Peterson and Wu 2021, Kapoor and Wilde 2023). Yet organizational processes often amplify rather than attenuate individual biases, making disciplined judgment difficult to achieve (Powell et al. 2011).

These limitations of human cognition suggest a natural question: might artificial intelligence offer a path forward? AI systems do not share the same constraints as humans—they have access to vastly more information, possess greater computational capacity, and exhibit unwavering consistency. Recent work has characterized this potential as “unbounding rationality” (Csaszar 2025): relaxing the cognitive bottlenecks that have long defined the boundaries of human judgment. Given the importance of strategic foresight and recent advances in AI, a consequential question emerges: Can AI outperform humans at strategic foresight?

Answering this question empirically faces a critical obstacle: most real-world outcomes are already known to both researchers and models, making it difficult to distinguish genuine prediction from pattern retrieval. Recent work has begun addressing this training-data leakage problem by evaluating LLMs on events occurring after their training cutoffs or on synthetic data not present in training corpora (Townsend et al. 2025, Schoenegger et al. 2024, Doshi et al. 2025). Dedicated forecasting benchmarks have also been developed (Karger et al.

Method. 3.1 Empirical Setting and Sample Kickstarter, a prominent online crowdfunding platform where entrepreneurs solicit financial contributions to bring proposed products to market, provides a natural setting for such evaluation. Each Kickstarter project launches publicly and remains open for several weeks before its fundraising outcome is determined. By selecting projects initiated after the latest model training cutoffs and capturing their content while funding is still in progress, we ensure that (i) no model could have seen these data during training, and (ii) the outcome variable itself (the ultimate amount raised) has not yet been determined. This design cleanly avoids any training data leakage, providing a genuinely forward-looking test of predictive capability.

To anonymize and standardize the information set, we provided both LLMs and human forecasters with an approximately 500-word summary of each campaign describing the product, team, key features, and risks, while excluding the venture’s name, the platform name (Kickstarter), funds raised to date, consumer comments, and other dynamic signals that might reveal campaign performance. Testing confirmed that identifying projects via search engines using these summaries is difficult. This design also guards against cheating: in addition to making search difficult, outcomes could not be looked up because they had not yet been determined.

We sampled the 30 live projects from Kickstarter’s Technology category with the closest 3.2 Prediction Task and Rankings Our empirical design evaluates how various LLMs and humans rank these projects based on their ex ante prediction of fundraising success compared to the actual funds raised ex post. All evaluations and rankings were completed during a 72-hour window (October 31 – November 2, 2025) while project fundraising was still live and outcomes were not yet known.

We also uploaded the predictions by the LLMs, Prolific participants, and experts to Zenodo in advance of actual outcomes becoming known.2

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

How do real-world evaluations reveal AI capabilities that benchmarks hide? What prevents LLMs from applying their reasoning knowledge to improve outputs? How does fine-tuning trade off accuracy against reasoning quality? What gaps exist between benchmark performance and real deployment outcomes? When should retrieval systems decide to fetch new information? What limits language model accuracy in evaluating ideas? Can artificial systems establish authority in domains requiring expert judgment? How can we maintain privacy when agents prioritize task completion? How does AI adoption reshape collaboration patterns in knowledge work? Does disclosing AI authorship change how audiences evaluate the writing? How do models learn from self-generated outputs without cascading failures? How do AI systems determine and balance multiple competing objectives? How can AI systems reliably guide voters without introducing political bias? What governance mechanisms can effectively constrain widely deployed AI systems? How should humans and AI agents share control and decision-making?