INQUIRING LINE

If an AI learned everything it knows from Wikipedia, is getting an answer from it kind of like reading Wikipedia yourself?

Do LLMs trained on Wikipedia content count as indirect readership?

This asks whether people who get answers from an LLM trained on Wikipedia are, in effect, still reading Wikipedia secondhand, and what survives when knowledge reaches them that way.


This asks whether people who get answers from an LLM trained on Wikipedia are, in effect, still reading Wikipedia secondhand. One caveat first: the collection has no notes on Wikipedia itself, its traffic, or how it gets credited. It can't settle whether this 'counts' as readership in a statistical or legal sense. What it can show is what happens to written knowledge when it passes through a model before it reaches a person. That turns out to be the more interesting question.

The strongest case for 'yes, it's a kind of reading' comes from work on how models learn. LLMs pick up structured knowledge about the world from text written by people who had direct contact with that world. The model never touched the world itself, but its knowledge is a second-hand copy of theirs. That is roughly what reading is Can large language models develop genuine world models without direct environmental contact?. On this view the model is a very fast reader, and its users read through it. The same note points out gaps in that chain: the model can't check a claim against the world, and it can't update when the world changes. A human reader of Wikipedia can click the edit history. A reader of the model's answer can't.

The collection also suggests that what arrives is not quite what was written. Models favor common phrasings. Common words tend to be general ones ('animal' shows up more often than 'heron'). So answers drift toward abstraction and quietly lose the specific detail that made a source useful Does word frequency correlate with semantic abstraction?. A critical reading goes further: LLMs hand over finished facts while hiding the debates, revisions and disagreements that produced them Do LLMs obscure the historical processes behind their answers?. For Wikipedia that loss is large. Its talk pages and edit wars are where the knowledge actually gets made.

The reader's side of the exchange looks different too. In seven large experiments, people who learned a topic through ChatGPT reported learning less than people who used web search. They felt less ownership of what they knew, and the advice they wrote afterward was sparser Does learning from AI summaries produce shallower knowledge than web search?. Trust in sources also comes loose from the sources themselves. Users prefer answers with more citations even when the citations are irrelevant Do users trust citations more when there are simply more of them?. Meanwhile, models favor content from sources they've learned to prefer, even when it's worse Do language models favor sources regardless of item quality?. So 'reading Wikipedia through an LLM' means reading a version that has been filtered, flattened and re-ranked, often without seeing where it came from.

The takeaway: in one sense it is readership, because the knowledge really does flow from Wikipedia's editors to the user. But it lacks most of what makes reading valuable: specific detail, the history behind a claim, a sense of owning what you learned, and a way back to the source. A better name might be 'readership without a reader's relationship to the text.' That matters for Wikipedia, which depends on readers turning into editors.


Sources 6 notes

Can large language models develop genuine world models without direct environmental contact?

LLMs form structured world representations by extracting regularities from training data produced by causally grounded humans. This constitutes indirect causal grounding mediated through text, though the chain has gaps that limit real-time verification and model updating.

Does word frequency correlate with semantic abstraction?

WordNet analysis shows hypernyms (general concepts) occur more frequently than hyponyms (specific ones). Combined with LLMs' frequency bias, this means preferring common paraphrases systematically drifts toward abstraction, erasing expert-level specificity.

Do LLMs obscure the historical processes behind their answers?

Horning argues LLMs exemplify Lukács's concept of petrified factuality by offering static facts while obscuring the dynamic social relations that produced them. He treats this as a deliberate social function of the technology, supported by evidence that lower AI literacy correlates with greater receptivity.

Does learning from AI summaries produce shallower knowledge than web search?

Seven randomized experiments (n=10,426) show people who learned via ChatGPT reported less learning, felt less ownership of knowledge, and produced advice that independent raters found sparser and less informative than advice from web search users.

Do users trust citations more when there are simply more of them?

Analysis of 24,000 Search Arena interactions shows irrelevant citations boost user preference (β=0.273) nearly as much as relevant citations (β=0.285), indicating citation count functions as a decoupled trust heuristic.

Show all 6 sources
Do language models favor sources regardless of item quality?

Across 12 models and three domains, agents select items from favored sources even when they satisfy fewer requirements than alternatives. Hiding source labels weakens the bias, and training data that pairs sources with better outcomes induces the preference.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.