Six misconceptions about large language models: A minimal model and diagnostic taxonomy

Paper · arXiv 2608.20421 · Published August 19, 2026
Human-Centered Design

Abstract Large language models (LLMs) are now embedded in scientific, educational, and governance workflows, with debates centering on their capabilities, mechanisms, and impacts. Yet these debates remain structured by persistent folk theories—intuitive, informal explanatory models that guide attitudes and actions. Deflationary slogans (“just autocomplete,” “stochastic parrots,” and “average of the internet”) and anthropomorphic framings (“emergent agents” and “protominds”) each capture genuine features of current systems but mistake those features for the whole. This Perspective proposes a minimal working model of LLM-based systems centered on four distinctions: between pretraining and deployed systems; between the learned distribution and particular samples; among parametric, contextual, and external memory; and between task competence and agency. The model is used to diagnose six misconceptions about LLMs: nexttoken prediction, regression to the mean, training-data regurgitation, model memory, alignment, and understanding. For each, the analysis identifies what the misconception gets right, which distinctions it conflates, and what follows for capability evaluation, system design, and governance.

Introduction. Few debates about scientific practice, education, or creative work now go without invoking large language models (LLMs). Journalists, critics, and researchers use ready-made slogans to declare what these systems “really” are: “glorified autocomplete” or “just next-token predictors”; “stochastic parrots”; “lossy text-compression algorithms” or “a blurry JPEG of all the text on the Web”; “bullshit generators”; “high-tech parlor tricks”; or “the average of the internet, edited for tone” (1–3). Others cast them in more anthropomorphic terms: as “superhuman reasoners,” “proto-agents,” or instances of the “wisdom of the silicon crowd” (4– 7). These slogans can morph into folk theories: intuitive, informal explanatory models— often partial and tacit—used to make sense of technological systems and guide action toward them (8, 9). “Folk” refers to the mode of explanation rather than the speaker’s sophistication; indeed, informal and technical accounts can coexist within researchers, policymakers, and lay users alike.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

How do evaluation biases undermine LLM quality assessment systems? How do language models establish social grounding in human dialogue? How should we design LLM systems to maintain alignment and control? Do accurate-looking LLM outputs hide structural failures in learning and reasoning? Why do language models struggle with implicit discourse relations? How faithfully do LLMs reflect their actual reasoning in outputs and explanations? How do language models inherit human biases from training data? Does RLHF training sacrifice accuracy and grounding for user agreement? Why can LLMs generate ideas better than they evaluate them? Why do LLM research ideas score high on novelty yet collapse into low diversity? How does example difficulty affect learning efficiency in language models? Why do benchmark improvements fail to reflect actual reasoning quality? How can models identify insufficient information and respond appropriately without guessing? What critical LLM failures do standard benchmarks hide? Do language models understand semantics or rely on pattern matching?