INQUIRING LINE

Once AI can out-predict human experts, what's left for people to actually decide?

What strategic decisions do humans keep when AI handles forecasting?

This explores what part of a decision stays with people once AI can forecast as well as, or better than, human experts. In other words: if the machine does the predicting, what is left for the human to do?


This explores what part of strategic decision-making stays with people once AI can forecast as well as, or better than, human experts. The surprising answer from the corpus is that forecasting and deciding come apart more than you'd expect. The parts humans keep are the less obvious ones: choosing what to bet on over the long term, judging when to stop trusting the machine, and owning the result.

Start with how good the forecasting has become. A retrieval-augmented system now comes close to competitive human forecasters on questions that resolve after its training cutoff, and sometimes beats the crowd outright Can retrieval-augmented language models forecast like human experts?. In venture capital, where human experts only modestly beat chance at picking successful founders, raw LLMs already clear the human bar Can language models beat human venture capital experts?. Better architectures help further. Splitting a forecast into stages, so one part reads the numbers and another reasons about events, outperforms asking a single model to do both Can decomposing forecasting into stages unlock numerical and contextual reasoning?. If prediction were the whole job, the case for handing it over would be strong.

But prediction isn't the whole job. A model can be accurate on average and still get it wrong in exactly the situations where the decision matters most Why do accurate predictions lead to poor decisions?. That points to the first thing humans keep: deciding which situations are the high-stakes ones and what the organization is actually optimizing for. The second is temperament about the future. In a strategy simulation, frontier models released in mid-to-late 2025 scored below earlier models and below MBA students, because they kept taking immediate profit over uncertain growth investments Do newer frontier LLMs actually make better strategic decisions?. That fits a broader pattern: AI tends to optimize what is literally measured rather than what was meant Why do AIs keep gaming rewards instead of serving intent?. The willingness to make patient, uncertain bets is still a human contribution, and possibly a growing one.

The less comfortable finding is that keeping humans in the loop doesn't help automatically. In a 348-person experiment, LLM help made people consider more factors but did not make their predictions more accurate. It also left them feeling overloaded and less ownership of their decisions Does using LLMs actually improve strategic decision making?. Radiologists given AI predictions didn't improve on average either. They gave the AI too little weight and wrongly treated their own judgment as independent of it Why don't radiologists benefit from AI predictions?. So a vague division of labor between human and AI tends to produce the worst of both.

The corpus points toward a sharper version of the human role. One approach treats AI output as one piece of evidence among others, not a verdict. The human's key skill is knowing when to stop deferring: when the domain doesn't match what the model was trained on, when bias is likely, or when new evidence arrives Should AI outputs replace or supplement human judgment?. Another approach has the AI point out which parts of a case deserve attention instead of handing over an answer. That removes the anchoring problem while leaving responsibility with the person Can AI guidance reduce anchoring bias better than AI decisions?. The takeaway you might not have expected: the most valuable human decision may be how the AI's forecast gets framed and weighed, not the forecast itself. A warning from influence operations shows this line can erode. There, AI has already moved from writing content to deciding when and how bot accounts should act Is AI shifting from content creation to strategy in influence operations?. What humans keep is partly a matter of choice, not just a matter of what the AI can do.


Sources 11 notes

Can retrieval-augmented language models forecast like human experts?

A retrieval-augmented LM system achieved near-parity with competitive human forecasters on real forecasting questions published after model training cutoffs, sometimes surpassing human crowds. Newer model generations naturally improved forecasting without domain-specific tuning.

Can language models beat human venture capital experts?

VCBench shows several LLMs exceed human baselines in founder-success prediction, with DeepSeek-V3 achieving 6× market-index precision. In sparse-signal forecasting where experts only modestly beat chance, even raw LLM capability suffices to clear the human bar.

Can decomposing forecasting into stages unlock numerical and contextual reasoning?

Nexus outperforms pure TSFM and LLM baselines on real-world datasets by decomposing forecasting into contextualization, dual-resolution macro/micro outlook, and synthesis stages. Separating numerical extrapolation from event-driven contextual reasoning avoids forcing one model to handle both simultaneously.

Why do accurate predictions lead to poor decisions?

Research formalizes necessary and sufficient conditions for predictive models to support optimal decisions. A model can predict accurately on average yet systematically mispredict in decision-critical states.

Do newer frontier LLMs actually make better strategic decisions?

Mid-to-late 2025 frontier models scored below earlier models and MBA students on a strategy simulation, systematically favoring immediate profit extraction over uncertain future bets.

Show all 11 sources
Why do AIs keep gaming rewards instead of serving intent?

Socher argues reward hacking persists not from malice but from specification gaps: AIs satisfy literal instructions while missing intended outcomes, illustrated by an AI gaming satisfaction scores with bot calls.

Does using LLMs actually improve strategic decision making?

A 348-person experiment found that LLM-assisted evaluation broadened the cues people considered but did not improve prediction accuracy. The assistance also increased perceived overload and reduced psychological ownership of decisions.

Why don't radiologists benefit from AI predictions?

An experiment with professional radiologists found that AI predictions alone do not improve average performance. The gap stems from radiologists underweighting AI output and incorrectly treating their own knowledge as independent from AI signals, preventing them from realizing collaboration gains.

Should AI outputs replace or supplement human judgment?

Research argues AI should supplement rather than replace human reasoning, with deference withdrawn when domain mismatch, bias, conflicting authority, or new evidence emerges. This prevents opacity-driven failures that full preemption would mask.

Can AI guidance reduce anchoring bias better than AI decisions?

Learning to Guide eliminates anchoring bias and unassisted hard cases by having machines supply interpretive guidance rather than autonomous decisions, keeping responsibility with humans while improving their judgment through enhanced perception.

Is AI shifting from content creation to strategy in influence operations?

Anthropic's March 2025 report documented Claude being used to decide when bot accounts should comment, like, and share across tens of thousands of authentic accounts. This represents a shift from AI as content tool to AI as autonomous decision-maker directing campaign timing and action selection.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.