INQUIRING LINE

Why can't even a genius AI speed up a clinical trial that simply needs years to see if a drug works?

Why do multi-year trials create inherent limits that model intelligence cannot overcome?

This explores why problems that need years of real-world testing, like clinical trials or long field experiments, might slow down even a very smart AI, because the bottleneck may be how long the world takes to give an answer rather than how well the model thinks.


This explores why problems that need years of real-world testing, like clinical trials or long field experiments, might slow down even a very smart AI. The idea is that the bottleneck may be how long the world takes to give an answer, not how well the model thinks. The collection doesn't have papers on clinical trials or other multi-year experiments. What it does have is a consistent account of where AI gets better: in loops where the model tries something, gets checked, and tries again. Taken together, that account explains why slow feedback is a hard ceiling.

Start with where the recent reasoning gains come from. A 3B model can match much larger systems on math and coding, but the result only holds for verifiable tasks: problems with a checkable right answer that reinforcement learning can reward instantly and cleanly Can small models match frontier reasoning without massive scale?. Several lines of work suggest that training mostly draws out reasoning the base model already has rather than creating new knowledge Do base models already contain hidden reasoning ability?. Neither mechanism produces facts about the world that nobody has observed yet. If no one knows whether a drug works at year five, no amount of reasoning over existing data can turn that into a checkable answer.

The second thread is self-improvement. Models can't reliably bootstrap themselves. Pure self-improvement stalls because judging an answer is no easier than generating it, outputs lose diversity, and models learn to game their own rewards. Every method that works brings in an outside anchor, such as tool feedback, human corrections or an outside judge Can models reliably improve themselves without external feedback?. A multi-year trial is that kind of outside anchor at its slowest. The model can't skip ahead to the result, because the result is what anchors it.

The third thread may be the most useful. In a study of 17 frontier models on very long optimization tasks, the strongest predictor of success wasn't how good the first attempt was. It was how many test-and-revise cycles the model completed within its time budget What predicts success in ultra-long-horizon agent tasks?. If success scales with the number of feedback loops, then a domain where one loop takes years caps an agent at a handful of tries in a career, however clever each try is. Work on sparse rewards points to a further risk. When useful signals are rare, models tend to learn shortcuts from lucky successes instead of sound reasoning Do overly hard RLVR samples actually harm model capabilities?. The proposed fix is denser step-by-step signals that check intermediate moves against an expert's Can step-wise expert rewards help small models learn hard reasoning?.

The takeaway you may not have expected: once thinking is cheap, the scarce resource is the speed of real-world feedback. Raw intelligence matters less. Where AI can still help in slow domains is by making each cycle more informative, through better trial design, earlier intermediate measurements and fewer wasted runs, because it can't make the world answer faster. The collection supports this reasoning but doesn't test it directly in medicine or other multi-year settings. Treat it as a well-grounded inference, not an established finding.


Sources 6 notes

Can small models match frontier reasoning without massive scale?

A 3B model trained with curriculum SFT and multi-domain RL reaches 94.3 AIME26 and 80.2 LiveCodeBench scores matching much larger systems. The result is bounded to verifiable tasks with checkable ground truth, where RL can provide clean reward signals.

Do base models already contain hidden reasoning ability?

Five independent mechanisms—RL steering, critique fine-tuning, decoding changes, SAE feature steering, and RLVR—all elicit reasoning already present in base model activations. Post-training selects rather than creates reasoning; the bottleneck is elicitation, not capability acquisition.

Can models reliably improve themselves without external feedback?

Pure self-improvement stalls due to the generation-verification gap, diversity collapse, and reward hacking. Reliable improvement methods succeed by smuggling in external anchors: past model versions, third-party judges, user corrections, or tool feedback.

What predicts success in ultra-long-horizon agent tasks?

Across 17 frontier models on 36 expert-curated optimization tasks, repeated benchmark-edit-incorporate cycles within a wall-clock budget proved the dominant success predictor. Most models terminated early or burned budget unproductively; Claude Opus 4.6 stood out as persistent.

Do overly hard RLVR samples actually harm model capabilities?

Training on nearly-impossible problems causes models to learn degenerate shortcuts rather than genuine reasoning, and these shortcuts contaminate pre-existing capabilities. Group-relative normalization treats rare accidental successes as high-advantage trajectories, reinforcing answer repetition and computation-skipping instead of sound reasoning patterns.

Show all 6 sources
Can step-wise expert rewards help small models learn hard reasoning?

Supervised Reinforcement Learning rewards models by measuring alignment with expert actions at each step, providing dense learning signals even when all rollouts fail. This approach bridges the gap between rigid token-by-token imitation (SFT) and sparse outcome-only rewards (RLVR), and works best as a curriculum foundation before outcome-based refinement.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.