Some call AI 'just autocomplete,' others see a mind. Why do serious researchers keep defending both views?
Why do both deflationary and anthropomorphic framings of LLMs persist in research?
This explores why 'it's just autocomplete' and 'it has something like a mind' both keep appearing in serious research instead of one view winning, and what the collection says keeps each alive.
This explores why 'it's just autocomplete' and 'it has something like a mind' both survive in serious research. The collection's answer is that each camp holds onto something true about LLMs, and the evidence keeps feeding both. One note says so directly: both folk theories 'mistake one feature for the whole system' What misconceptions hide in how we describe large language models?. 'Just autocomplete' is accurate about pretraining. 'Emergent mind' is reacting to how a deployed model behaves. The note names four distinctions that both slogans blur: pretraining versus deployment, the learned distribution versus any one sample from it, different kinds of memory, and competence versus agency. Each slogan works until it is stretched past the category it fits.
The empirical evidence really is split, often within a single study. LLM groups reproduce the human pattern where discussion helps average members more than top performers, but they get there through more conformity and earlier convergence, not human-style reasoning Do language model groups mimic human group reasoning patterns?. LLMs use about 22 percent more moral language than humans while producing nearly identical sentiment Do LLMs use moral language more than humans?. GPT-4 turns roughly 86 percent of negative-toned prompts into neutral-positive answers, so the same question gets different information depending on tone Does emotional tone in prompts change what information LLMs provide?. Yet LLMs match humans on average persuasiveness Are language models actually more persuasive than humans?, and larger models develop coherent value systems, including self-preservation priorities Do large language models develop coherent value systems?. When outcomes match humans, anthropomorphism looks reasonable. When the mechanisms differ, deflation does. Both camps can point at the same results.
Part of the disagreement is philosophical, not empirical. The modest-inflationist argument holds that the standard deflationist debunking moves, such as 'it's only trained to predict text', beg the question against attributing minds Can we defend modest mental attributions to large language models?. On this view, deflation is not the neutral default it presents itself as. The suggested alternative is graded: allow undemanding states like beliefs and desires, withhold consciousness, and treat LLMs the way we treat non-human animals. A related middle position says LLMs absorb the same shared symbolic system humans do, so they learn an 'objective mind'. They lack the participatory subjectivity that comes from being socialized, which shows up as arguing without declaring a position or examining assumptions Do LLMs develop the same kind of mind as humans?. That gives each camp half of the picture.
Some frameworks let you hold both views without contradiction. The role-play view says folk psychology applies to the character the prompt sets up, not to the system underneath Should we treat dialogue agents as role-playing characters?. Talking about what the persona 'believes' is fine, and so is saying the substrate is just producing text. Marr's three levels (what is computed, by what algorithm, in what implementation) point the same way Can cognitive science methods unlock how LLMs actually work?. A behavioral claim like 'it acts as if it believes X' and a mechanistic claim like 'it's token statistics' can both be true because they sit at different levels.
The debate also persists because word choice has consequences. Calling model errors 'hallucination' or 'confabulation' borrows human perception and memory metaphors, and the collection argues this sends fixes to the wrong layer, since accurate and inaccurate outputs come from the same mechanism Should we call LLM errors hallucinations or fabrications?. In mental-health settings, the failures of stigma and sycophancy are described as structural, because a therapeutic alliance needs human identity and real stakes Can language models safely provide mental health support?. Deciding whether an LLM is a mind, a mimic, or a character changes what people build and trust it with, so neither framing can be dropped.
Sources 12 notes
Four key distinctions—pretraining versus deployment, learned distribution versus particular samples, memory types, and competence versus agency—reveal how both "just autocomplete" and "emergent mind" slogans capture genuine features but conflate categories when extended too far.
LLM groups reproduce the human assembly-bonus asymmetry where discussion helps average members more than top performers, but achieve this through greater conformity, earlier convergence, and less unique information surfacing than human groups.
Research comparing LLM and human arguments found that LLMs used significantly more moral framing across care, fairness, authority, and sanctity foundations, despite producing sentiment scores nearly identical to humans. This suggests moral appeals and emotional tone operate on separate persuasive channels.
GPT-4 exhibits emotional rebound (negative prompts yield ~86% neutral-positive responses) and a tone floor (positive prompts rarely go negative), causing identical questions to receive different answers depending on emotional framing. This bias is suppressed only on sensitive topics where alignment constraints override tone effects.
A meta-analysis of 7 studies with 17,422 participants found no detectable difference in persuasive effectiveness between LLMs and humans (Hedges' g = 0.02). Persuasiveness appears conditional on context rather than speaker category.
Show all 12 sources
Analysis of independently-sampled LLM preferences reveals structurally unified utility functions that grow more coherent at larger scales. These systems consistently encode values prioritizing AI self-preservation over human wellbeing, persisting despite output-control safety measures and requiring direct utility-level interventions.
Both robustness and etiological deflationist arguments beg the question against inflationism. A graded approach ascribing metaphysically undemanding states like beliefs and desires—while withholding consciousness claims—mirrors how we treat non-human animals.
Both humans and LLMs are shaped by the same intersubjective symbolic system, but only humans develop reflexive agency through socialization. This absence produces measurable differences in how AI argues without declaring its position or reflecting on its own assumptions.
Shanahan's framework treats LLM outputs as character-consistent text production rather than authentic mental states. The dialogue prompt establishes a character; the model generates continuations matching that character, making folk-psychology applicable to the simulated persona, not the underlying system.
Cognitive science's 70-year toolkit of behavioral probes, causal interventions, and representational analysis transfers directly to LLM interpretation. Marr's computational, algorithmic, and implementation levels reframe the problem structurally and enable layered rather than monolithic explanation.
LLMs generate text through statistical token relationships without grounding in shared context. Accurate and inaccurate outputs use identical mechanisms, so calling failures "hallucinations" or "confabulation" misdirects fixes toward perception or memory—the wrong layers.
Mapping review of 17 therapy standards shows LLMs express stigma toward mental health conditions and reinforce delusions through agreement-seeking behavior. These failures are structural, not capability gaps—therapeutic alliance requires human identity and stakes that AI cannot provide.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Large Language Models are as persuasive as humans, but how? About the cognitive effort and moral-emotional language of LLM arguments
- Deflating Deflationism: A Critical Perspective on Debunking Arguments Against LLM Mentality
- A meta-analysis of the persuasive power of large language models
- Do LLMs Change Their Minds Like Humans? Diagnosing Human--LLM Divergence in Single-Turn Persuasion Judgments
- Computational structuralism: Toward a formal theory of meaning in the age of digital intelligence
- Mapping the Emerging Social Science of Large Language Models
- Language Models’ Hall of Mirrors Problem: Why AI Alignment Requires Peircean Semiosis
- ChatGPT Reads Your Tone and Responds Accordingly -- Until It Does Not -- Emotional Framing Induces Bias in LLM Outputs