Does learning from AI summaries produce shallower knowledge than web search?
This explores whether the convenience of LLM-generated summaries trades off against the depth of understanding people develop compared to traditional web search. The question matters because it affects how people learn and teach others.
Seven experiments across two preregistered designs, a lab replication, and a downstream-adoption test (combined n = 10,426, including a primary sample of 1,104 after exclusions) find that participants who learned about a topic via ChatGPT summaries, rather than Google web links, developed shallower knowledge and produced lower-quality advice for others. In the first experiment, ChatGPT users spent less time on the search task (MGPT = 585.41s vs. MGoogle = 742.81s, F(1,1102) = 44.61, P < 0.001) despite submitting a similar number of queries (2.06 vs. 2.15, P = 0.401) — the authors read this as "less engagement with the search results," not less interactivity. ChatGPT users also reported learning fewer new things (MGPT = 3.43 vs. MGoogle = 3.86, P < 0.001), felt less "personal ownership over the knowledge they gained" (P = 0.009), and rated the information as less comprehensive (P < 0.001). A fourth experiment showed the downstream cost: independent recipients, blind to which platform the advice-writer had used, found ChatGPT-sourced advice "sparser and more generic" and were less willing to adopt it.
The paper's mechanism is effort substitution, not raw information quality: "this lower effort in assembling knowledge from LLM syntheses...risks suppressing the depth of knowledge that users gain, which subsequently affects the nature of the advice they form." It draws on "search-as-learning" research, which holds that web search builds knowledge structures through "posing queries, gathering and interpreting information from different websites, and then assembling this knowledge into a cohesive whole" — a recursive "sensemaking" process. LLMs perform that sensemaking for the user, so "this critical ingredient in learning is often diminished." The authors frame this as a trade of efficiency for depth, and report the effect held even when LLM summaries were augmented with real-time web links and when the search engine itself was held constant (Google's own "AI Overview" vs. its standard results).
This supplies the experimental, causal counterpart to what Do large language models narrow human expression and thought? argues from training statistics without new measurement: where that paper infers narrowing from how models are trained and shared, this one randomly assigns people to a platform and measures a narrowing outcome — self-reported depth, then independently judged advice quality — directly. It also gives a concrete, measured instance of what Why do people trust AI outputs they shouldn't? describes abstractly as Trap 2 — users substituting the fluency of a synthesized answer for the grounded effort of building their own understanding; here that substitution is operationalized as time-on-task and shows up downstream in advice quality.
The depth-of-learning measure is self-report (how much participants felt they learned, felt ownership of, judged comprehensive), not an independent knowledge test; only the downstream advice-sparseness judgment was made by blind, independent raters. The flagship experiment used a single, low-stakes topic (how to plant a vegetable garden) and two platforms, though the paper states the pattern replicated across other topics and tools in supplementary experiments not excerpted here. The design establishes a causal effect of summary-format learning on self-reported depth and on the objective quality of a downstream writing task, within this learn-then-advise structure — it does not establish that LLM-mediated learning produces shallower knowledge in higher-stakes settings, with repeated use, or absent an advice-writing task.
Inquiring lines that read this note 14
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
Are AI-generated articles systematically disadvantaged in search ranking and user engagement?- Can AI search summaries influence voters without active seeking?
- How do publishers distinguish between search crawling and AI training requests?
- Can AI summaries recover lost clicks through citations or credit to sources?
- Which search queries trigger AI summaries most often on Google?
- Do users click links within AI summaries or end sessions instead?
- Do searchers prefer clarity about AI involvement when viewing search overviews?
- How much do AI Overviews currently appear in Google search results?
- Are zero-click searches rising because of AI answer summaries?
- Does effort reduction during search affect how deeply people understand topics?
Related concepts in this collection 3
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Do large language models narrow human expression and thought?
Explores whether LLMs homogenize how people write, think, and reason by reflecting narrow training distributions and subtly shifting user preferences toward model outputs.
this excerpt's experimental causal measurement where that note only argues from training statistics
-
Why do people trust AI outputs they shouldn't?
When do human cognitive shortcuts fail in AI interaction? Three compounding traps—treating statistical patterns as facts, mistaking fluency for understanding, and avoiding disagreement—may explain systematic overreliance across languages and contexts.
measures the fluency-for-effort substitution Rose-Frame names as Trap 2, as time-on-task and downstream advice quality
-
Does ChatGPT harm informal learning compared to Google Search?
An 8-day experiment tested whether using ChatGPT for self-directed learning produces different knowledge gains than Google Search, and explored what mechanisms might explain any differences.
extends A: an 8-day field experiment finds ChatGPT learners tested worse than Google searchers, attributing the gap to diminished agency and information distortions
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Experimental evidence of the effects of large language models versus web search on depth of learning
- AI Meets the Classroom: When Does ChatGPT Harm Learning?
- What and Whose Knowledge? Measuring Epistemic Diversity in Large Language Models
- Generalization Bias in Large Language Model Summarization of Scientific Research
- Explain-Query-Test: Self-Evaluating LLMs Via Explanation and Comprehension Discrepancy
- LLM Augmentations to support Analytical Reasoning over Multiple Documents
- Teaching a Large Language Model Tutor to Withhold the Answer: A Supervisor Architecture and an Evidence-Driven Method for Tuning Socratic Behavior
- LLM or Human? Perceptions of Trust and Information Quality in Research Summaries
Original note title
learning from LLM summaries yields shallower reported knowledge and sparser advice than web search across seven experiments