INQUIRING LINE

Does searching a paper's full text beat searching just its abstract, and when does that extra text actually help?

How much does using full PDF text improve over abstract-only retrieval?

This explores whether searching over a paper's full text (methods, results, tables) finds better material than searching over abstracts alone, and how large that gain is.


This explores whether retrieving over a paper's full text finds better material than retrieving over abstracts alone, and by how much. The short answer: none of these notes runs that head-to-head comparison, so the corpus can't give you a number. What it does have is more useful than a number. It explains *when* more text helps, when it doesn't, and why the answer depends on how you search rather than how much you index.

The case for full text is mostly about precision. Abstracts are written to sound like their topic, so embedding search over them works well for 'find me papers about X.' It struggles with 'find the paper that used dataset Y with method Z,' because those details live in the body. One line of work shows that letting an agent grep through raw text beats dense embeddings on entity-constrained, multi-hop questions, because embeddings blur together entities that a literal string match keeps apart Can direct corpus search beat embedding-based retrieval?. The broader diagnosis is that embeddings measure *association*, not *relevance*. Adding more text doesn't fix that mismatch, and it can make it worse, since a longer document's vector averages over more unrelated content Where do retrieval systems fail and why?.

That points to the catch: chopping full text into chunks throws away what made the full text valuable. A methods section only makes sense in light of the research question, and a results table only makes sense next to its baseline. MiA-RAG addresses this by first summarizing the whole document into a 'map' and then retrieving with that map in hand, so scattered evidence can be found by its role in the paper and not only by surface similarity Can building a document map first improve retrieval over long texts?. Seen this way, an abstract is a hand-written version of that map. So the real question isn't 'abstract or full text.' It's whether your system can use the abstract-level view to steer search through the full text. Another option skips chunking: long-context models can match RAG on semantic retrieval by simply reading everything, though they break down on structured queries that need joins across tables Can long-context LLMs replace retrieval-augmented generation systems?.

There are also two hidden costs. First, PDF extraction is messy: broken columns, garbled equations, OCR errors. A system built for noisy historical newspapers succeeded only by casting a wide retrieval net and then refusing to answer anything it couldn't ground in the text Can RAG systems refuse to answer without reliable evidence?. Second, more retrievable passages means more near-misses: chunks that are on topic but don't actually match. One fix is a separate verification step after recall that checks fine-grained token-level matches instead of trusting a single similarity score Can verification separate structural near-misses from topical matches?.

The thing you may not have expected to want to know: more retrieved text can make answers *feel* better without making them better. In 24,000 search-engine comparisons, users preferred responses with more citations almost as much when the citations were irrelevant as when they were relevant Do users trust citations more when there are simply more of them?. So if you measure 'full text vs. abstract' by user satisfaction, full text may look like a win just because it produces more citable passages. Measuring whether the retrieved passage actually contains the answer is the honest test, and it's the experiment this corpus still lacks.


Sources 7 notes

Can direct corpus search beat embedding-based retrieval?

GrepSeek trains agents to retrieve via executable shell commands over raw text, achieving better multi-hop performance on entity-constrained queries than dense embeddings. The approach scaffolds unstable search mechanics with supervised trajectories, then refines task-oriented behavior through reinforcement learning.

Where do retrieval systems fail and why?

RAG systems fail at three structural levels: adaptive triggering (fixed intervals waste context), semantic-task mismatch (embeddings measure association, not relevance), and mathematical limits (embedding dimension constrains representable document sets). These require fundamentally different retrieval approaches, not tuning.

Can building a document map first improve retrieval over long texts?

MiA-RAG inverts standard RAG by summarizing documents first, then conditioning retrieval on that global view. This approach recovers discourse structure that bag-of-chunks retrieval destroys, making scattered evidence findable by their document role rather than surface similarity alone.

Can long-context LLMs replace retrieval-augmented generation systems?

The LOFT benchmark shows LCLMs match RAG on semantic retrieval without explicit training, but cannot execute relational queries requiring joins across structured tables. Context length alone cannot bridge this gap.

Can RAG systems refuse to answer without reliable evidence?

A multilingual RAG system for noisy historical newspapers succeeds by aggressively expanding retrieval while constraining generation to only grounded answers. The grounded-refusal prompt prevents hallucination when OCR errors and language drift degrade source quality, trading coverage for integrity.

Show all 7 sources
Can verification separate structural near-misses from topical matches?

A two-stage pipeline—pooled-cosine recall followed by a small Transformer verifier operating on token-token similarity maps—reliably rejects structural near-misses that MaxSim-style late interaction cannot. The verifier succeeds because it operates on full token interaction patterns rather than compressed vectors.

Do users trust citations more when there are simply more of them?

Analysis of 24,000 Search Arena interactions shows irrelevant citations boost user preference (β=0.273) nearly as much as relevant citations (β=0.285), indicating citation count functions as a decoupled trust heuristic.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.