INQUIRING LINE

Does letting an AI look things up fix its trouble judging which ideas are actually good or worth pursuing?

Do retrieval systems help language models overcome their inability to estimate idea value?

This explores whether giving language models access to outside information (retrieval) can fix a deeper weakness: their trouble judging which ideas are actually good, novel, or worth pursuing. The corpus has no paper that tests this directly, but it has nearby evidence on what retrieval does and doesn't fix.


This explores whether retrieval, meaning letting a model look things up instead of relying only on what it memorized, can help it judge which ideas are actually good, novel, or worth pursuing. To be upfront: no paper in this collection tests retrieval as a fix for judging idea value. What the corpus does have is evidence on a narrower question: can retrieval correct the habits that make models poor judges in the first place? The answer is mostly no, with one interesting exception.

Start with why models judge poorly. Models often rate a claim as true because it sounds familiar from training, not because the evidence supports it. In entailment tests, they kept saying a conclusion followed from a premise even when the premise was random, as long as the conclusion was something they'd seen before Do LLMs predict entailment based on what they memorized?. If a model judges ideas the same way, it will favor ideas that resemble its training data, which undercuts any attempt to spot what's new. That pull is collective too: because so many people use the same models, writing converges and users absorb the model's framings without noticing Do large language models narrow human expression and thought?. A judge that rewards the familiar adds to that sameness.

Retrieval might seem like the fix, but there's a catch. When retrieved information conflicts with strong prior associations, models often ignore it. Prompting alone can't override those priors. It takes direct intervention in the model's internal representations Why do language models ignore information in their context?. A related failure: models can correctly explain a standard, fail to apply it, and even notice that they failed Can LLMs understand concepts they cannot apply?. So handing a model good evaluation criteria, or strong comparison cases, doesn't mean it will use them. Most retrieval research also targets a different problem: getting the right facts at the right time Can retrieval systems ground answers in the right time? or deciding when to look something up at all When should language models retrieve external knowledge versus use internal knowledge? Can simple uncertainty estimates beat complex adaptive retrieval?. Better facts are not the same thing as better judgment.

The exception points somewhere unexpected. XConf improves a model's confidence estimates not by retrieving outside knowledge but by retrieving the model's own past attempts and checking how often similar ones turned out right. The gain comes entirely from the stored outcomes. Remove them and it disappears Can past performance predict when a model will be right?. That suggests a different kind of retrieval for judging ideas. Instead of fetching related papers and asking the model to reflect, you would fetch a record of what happened to similar ideas: which ones replicated, got adopted, or went nowhere. The model's own reasoning in the moment may simply not be enough to judge value, but a track record of outcomes could be.

The takeaway is that retrieval probably can't teach a model taste. It might be able to stand in for taste by bringing in outcome history the model can't produce on its own. That reframes the problem as building a memory of what happened to past ideas, not just a better search engine. Whether that works for something as open-ended as research ideas, rather than tasks with checkable answers, is a gap in this collection.


Sources 8 notes

Do LLMs predict entailment based on what they memorized?

McKenna et al. (2023) identified attestation bias: LLMs predict entailment based on whether the hypothesis appears in training data, not whether the premise actually supports it. Random premise experiments show models maintain high entailment predictions when hypotheses are attested, proving they respond to memorized propositions rather than premise-hypothesis relationships.

Do large language models narrow human expression and thought?

LLMs mirror skewed slices of human experience shaped by training data regularities, and widespread reliance on identical models amplifies convergence. Co-writing studies show users unconsciously adopt model stances and framings.

Why do language models ignore information in their context?

Research demonstrates that LMs generate outputs inconsistent with their context because parametric knowledge from training dominates over in-context information. Textual prompting alone cannot override strong priors; causal intervention in representations is required.

Can LLMs understand concepts they cannot apply?

Models can explain concepts accurately, fail to apply them, and recognize the failure—a triple pattern incompatible with human cognition. This indicates functionally disconnected explanation and execution pathways rather than simple knowledge gaps.

Can retrieval systems ground answers in the right time?

TempRALM adds a temporal term to retrieval scoring alongside semantic similarity, achieving up to 74% improvement over baseline systems when documents have multiple time-stamped versions. The approach requires no model retraining or index changes.

Show all 8 sources
When should language models retrieve external knowledge versus use internal knowledge?

DeepRAG models each reasoning step as a Markov Decision Process where the model learns when to retrieve versus rely on parametric knowledge. The 21.99% improvement comes from better-targeted retrieval and elimination of noise from unnecessary external knowledge.

Can simple uncertainty estimates beat complex adaptive retrieval?

Calibrated token-probability uncertainty consistently beats multi-call adaptive retrieval on single-hop tasks and matches performance on multi-hop, using a fraction of the LM and retriever calls. The model's self-knowledge proves more reliable than external heuristics for deciding when to retrieve.

Can past performance predict when a model will be right?

XConf matches ten-sample self-consistency at a tenth of the cost by retrieving the model's past episodes with similar confidence levels and reading their historical success rates. Ablations show the signal depends entirely on stored outcomes, not on the retrieval prompt itself.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.