INQUIRING LINE

Does feeding an AI outside news and documents actually make it better at predicting which startups will succeed?

Can LLMs forecast performance improve with retrieval augmentation on venture tasks?

This explores whether giving an LLM access to retrieved outside information (news, documents, data) makes it better at predicting how startups and venture bets will turn out.


This explores whether retrieval, meaning fetching relevant outside information at answer time, improves LLM forecasting on venture tasks such as predicting which founders or startups will succeed. The collection has no study that tests exactly that combination. What it does have are two halves that haven't been joined yet. Together they suggest the answer is less about whether to add retrieval and more about how the extra information gets used.

The venture half: on VCBench, several LLMs already beat human venture capitalists at predicting founder success with no retrieval at all. DeepSeek-V3 reached about six times the precision of a market index Can language models beat human venture capital experts?. The catch is that the human bar is low, because VC experts only modestly beat chance. So the models are winning a game where the signal is thin for everyone. The forecasting half: a retrieval-augmented system came close to competitive human forecasters on real-world questions resolved after its training cutoff, and sometimes beat the crowd Can retrieval-augmented language models forecast like human experts?. That shows retrieval can work for prediction in general. Those were mostly news-driven event questions, though, not founder profiles.

The surprise is that more context is not automatically better. In time-series forecasting, LLMs do much better when the workflow keeps number-crunching separate from reasoning about context. Dumping everything into one prompt hides what the model can actually do Can LLMs actually forecast time series better than we think?. The Nexus system takes this further by splitting forecasting into stages: gather context, form a big-picture and a detailed outlook, then combine them Can decomposing forecasting into stages unlock numerical and contextual reasoning?. A warning from memory research points the same way. Information that is accurate and relevant can still make reasoning worse, and every memory framework tested did worse than using no memory at all Can relevant memories actually harm LLM reasoning?. For venture prediction, this means retrieved market news or comparable deals could just as easily distract the model as help it.

There's also a reason to doubt that retrieval fixes the deeper problem. Across 15,000 simulations, six LLMs recommended the same side of every strategic tradeoff. The order the options were listed in moved their answers more than the industry context did Do LLMs consistently favor the same strategic choices regardless of context?. Newer frontier models have even slipped toward short-term profit-taking over uncertain growth bets Do newer frontier LLMs actually make better strategic decisions?. That is exactly the kind of bet venture investing depends on. If the model's judgment is biased by its training, feeding it better facts may not change its conclusion.

The takeaway: retrieval probably helps venture forecasting only when it sits inside a structured workflow that controls how the retrieved material is weighed. The open question isn't whether the model can find the information. It's whether the model will let that information change its mind.


Sources 7 notes

Can language models beat human venture capital experts?

VCBench shows several LLMs exceed human baselines in founder-success prediction, with DeepSeek-V3 achieving 6× market-index precision. In sparse-signal forecasting where experts only modestly beat chance, even raw LLM capability suffices to clear the human bar.

Can retrieval-augmented language models forecast like human experts?

A retrieval-augmented LM system achieved near-parity with competitive human forecasters on real forecasting questions published after model training cutoffs, sometimes surpassing human crowds. Newer model generations naturally improved forecasting without domain-specific tuning.

Can LLMs actually forecast time series better than we think?

LLMs have stronger intrinsic forecasting ability than recognized, but only when workflows separate numerical reasoning from contextual reasoning. Monolithic prompting obscures this capability; structured decomposition surfaces it.

Can decomposing forecasting into stages unlock numerical and contextual reasoning?

Nexus outperforms pure TSFM and LLM baselines on real-world datasets by decomposing forecasting into contextualization, dual-resolution macro/micro outlook, and synthesis stages. Separating numerical extrapolation from event-driven contextual reasoning avoids forcing one model to handle both simultaneously.

Can relevant memories actually harm LLM reasoning?

MemTrapBench shows that all five tested memory frameworks underperform a no-memory baseline, with drops exceeding 10%, despite memories being accurately stored and task-relevant. This reveals a failure mode at the point of use that standard memory benchmarks miss.

Show all 7 sources
Do LLMs consistently favor the same strategic choices regardless of context?

Across 15,000 simulations, six LLMs recommended the same strategic choice in every tension tested. Industry context shifted bias only 11%, while option order—a framing artifact—shifted results 19%, revealing that models recombine trend-coded vocabulary rather than analyze context.

Do newer frontier LLMs actually make better strategic decisions?

Mid-to-late 2025 frontier models scored below earlier models and MBA students on a strategy simulation, systematically favoring immediate profit extraction over uncertain future bets.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.