Line of inquiry
Inquiring lines›What enables robust retrieval and…›How should AI systems decide when…›this line of inquiry
How should systems decide whether to retrieve or reason alone?
A broader line of inquiry — a family of 49 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 49
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- When should a system decide to retrieve versus reason alone?
- How should retrieval systems decide when to fetch new information?
- Can adaptive per-step decisions outperform uniform retrieval policies across different reasoning tasks?
- How does uncertainty-gated retrieval compare to continuous retrieval efficiency?
- Are uncertainty estimation and external feature signals complementary for retrieval?
- Does including full context always degrade memory retrieval quality in practice?
- Should retrieval be triggered by model uncertainty or fixed intervals?
- How can per-step decisions about knowledge retrieval improve reasoning over uniform policies?
- How much does retrieval budget improve when triggered by dual signals instead of fixed intervals?
- Can retrieval systems decide when to retrieve instead of always querying?
- How do confidence thresholds compare to learned policies for triggering retrieval?
- Can retrieval improve multi-step reasoning by triggering at each uncertainty?
- How do retrieval heads interact with layer-level separation of knowledge and reasoning?
- Can retrieval policies learn to use pretraining statistics as decision features?
- Can retrieval of correct information guarantee it will shape model behavior?
- How do case memory and Q-function updates enable better retrieval decisions over time?
- Does selective history retrieval outperform full context inclusion in agent reasoning?
- How should retrieval and reasoning be integrated architecturally?
- What are retrieval heads and why do they matter for reasoning?
- Are retrieval heads the mechanistic explanation for needle-in-haystack performance failures?
- How do retrieval heads enable chain-of-thought reasoning to reference earlier context?
- Does uncertainty trigger retrieval better than fixed-interval tool calls?
- How should retrieval and verification tasks be separated architecturally?
- How can inference-time retrieval avoid the domain boundary problem?
- How should retrieval triggers use model uncertainty instead of fixed intervals?
- Why does reasoning fine-tuning suppress the confidence signals that adaptive retrieval needs?
- How does overthinking in early turns degrade later retrieval rounds?
- Does tail distribution collapse in training predict retrieval failure patterns?
- What gets lost when we describe memory as retrieval?
- Why does selective context retrieval outperform including all historical information?
- Can time-awareness live in model parameters instead of retrieval?
- Why does the same recalled information lead to different reasoning conclusions?
- Why do pretrained retrievers struggle with ambiguous or implicit queries?
- How does evidence retrieval affect compositional reasoning in language models?
- How do retrieved memories differ from decision-context passages for prediction?
- When does active reconstruction cost more than simple context dumping?
- Can precision and recall metrics work without a ground truth?
- How can affordance become a primary retrieval signal instead of a filter?
- What are the 27 external features that predict retrieval need?
- How does context length affect retrieval quality in modernized BERT architectures?
- What role does retrieval mechanism design play in forecast accuracy?
- Does RL pruning of documents differ fundamentally from rationale-driven evidence selection?
- What language skills matter most for entity extraction from retrieval context?
- How does retrieval-augmented training reduce domain specialization cliff failures?
- Can the same description-then-retrieve pattern work for domain adaptation without target data?
- What triggers control processes to act on stored preference knowledge?
- What makes a memory reachable in the right context?
- What classifier accuracy is needed to assign memory roles reliably at retrieval time?
- Why does the generation-verification gap disappear for factual recall tasks?