Does using AI answers actually change how you think, not just what you conclude?
Does exposure to LLM answers actually change how people think critically?
This explores whether reading or working with AI-generated answers changes how people reason, weigh evidence and form judgments, rather than whether it only changes what conclusions they end up with.
This explores whether exposure to LLM answers changes the process of human thinking, not just its conclusions. The short version: the collection has no study that directly measures critical-thinking skill before and after AI use. It does have several findings that, taken together, suggest LLMs change how people think in ways they may not notice, and that the change isn't always toward better judgment.
The most direct evidence comes from a 348-person experiment on strategic decisions Does using LLMs actually improve strategic decision making?. People who used an LLM considered a wider range of factors, so their mental picture of the problem really did change. Their predictions got no more accurate, though. They also felt more overloaded and less ownership of their decisions. So the AI changed the shape of their thinking without improving its quality, and it loosened their sense that the reasoning was their own. That second effect is easy to miss, and it may matter more for critical thinking than accuracy does.
How you encounter AI output also seems to matter. In a study of AI debates, people arguing directly with an LLM changed their minds only about 7% of the time, while people who simply read the same exchanges were swayed 34–62% of the time Why do LLM audiences shift views more than debaters?. Arguing back acts as a built-in critical filter. Reading passively removes it. Most people meet LLM answers in the passive mode: reading a response, not pushing against it. Related work shows that even ML experts can't reliably tell AI-written abstracts from human ones, and they rated AI-edited text the clearest Can readers tell LLM abstracts from human ones?. Polished, fluent prose lowers the guard that critical reading depends on.
The collection also has a quieter worry: the range of ideas people are exposed to shrinks. LLM essays on debate topics recover only about half of the distinct arguments humans make, and they lean on the same hedged sub-points Do language models flatten the range of public arguments?. Critical thinking needs real alternatives to weigh against each other. If the answers you read keep drawing on a narrower set of arguments, you can become a careful evaluator of a shrinking menu. Answers can also shift with the emotional tone of the question: negative prompts get steered back toward neutral or positive responses Does emotional tone in prompts change what information LLMs provide?. Two people asking the same question in different moods may get different information without knowing it.
For a more philosophical angle, one critique argues that LLMs present knowledge as fixed, finished facts and hide the debates and history that produced them. The same essay points to evidence that people with lower AI literacy are more receptive to that framing Do LLMs obscure the historical processes behind their answers?. Read together, these findings suggest the risk isn't that AI makes people stop thinking. The risk is that it changes the conditions for thinking: fewer competing arguments, more fluent certainty, more passive reading and less ownership. Whether people stay critical may depend less on the AI's quality than on whether they engage with it like a debater instead of an audience member.
Sources 6 notes
A 348-person experiment found that LLM-assisted evaluation broadened the cues people considered but did not improve prediction accuracy. The assistance also increased perceived overload and reduced psychological ownership of decisions.
The Thin Line study found debate participants showed only 7% mind-change rates, while audience readers of the same exchanges showed 34–62% sway. Defensive friction in real-time conversation protects beliefs; read-only consumption lacks this friction.
Readers with ML expertise struggle to identify LLM-generated content reliably, tending to assume human involvement across all abstract types. However, LLM-edited abstracts received highest clarity ratings and were preferred 55% of the time when authorship was disclosed.
Across 23,384 LLM essays on debates, models recover only half of distinct human arguments and reuse hedged sub-arguments—a gap in argumentative structure, not just prose style. Diversity prompting adds noise outside human argument space rather than filling the long tail.
GPT-4 exhibits emotional rebound (negative prompts yield ~86% neutral-positive responses) and a tone floor (positive prompts rarely go negative), causing identical questions to receive different answers depending on emotional framing. This bias is suppressed only on sensitive topics where alignment constraints override tone effects.
Show all 6 sources
Horning argues LLMs exemplify Lukács's concept of petrified factuality by offering static facts while obscuring the dynamic social relations that produced them. He treats this as a deliberate social function of the technology, supported by evidence that lower AI literacy correlates with greater receptivity.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Do LLMs Change Their Minds Like Humans? Diagnosing Human--LLM Divergence in Single-Turn Persuasion Judgments
- Large Language Models are as persuasive as humans, but how? About the cognitive effort and moral-emotional language of LLM arguments
- The Homogenizing Effect of Large Language Models on Human Expression and Thought
- A meta-analysis of the persuasive power of large language models
- Unlocking Varied Perspectives: A Persona-Based Multi-Agent Framework with Debate-Driven Text Planning for Argument Generation
- LLM or Human? Perceptions of Trust and Information Quality in Research Summaries
- ChatGPT Reads Your Tone and Responds Accordingly -- Until It Does Not -- Emotional Framing Induces Bias in LLM Outputs
- Argument Collapse: LLMs Flatten Long-Form Public Debate