INQUIRING LINE

Is an AI actually thinking, or just remixing everything humans have ever written — and does that distinction even hold up?

Can models generate intelligence or only reflect human discourse?

This explores whether language models produce anything that counts as their own thinking, or whether everything they output is a sophisticated echo of the human writing they were trained on. The corpus suggests the either/or framing is shakier than it looks, and that part of the answer lies with the people reading the output.


This explores whether language models produce anything that counts as their own thinking, or whether everything they output is a sophisticated echo of the human writing they were trained on. The corpus suggests the either/or framing is shakier than it looks, and that part of the answer lies with the people reading the output.

The case for 'reflection' is strong. Several separate lines of research find that post-training doesn't teach models to reason. It brings out reasoning that was already there in the base model, which learned it from human text. RL, critique fine-tuning, decoding tweaks and feature steering all unlock the same latent ability Do base models already contain hidden reasoning ability?. The visible 'thinking' that models write out may not be reasoning either. Traces with invalid logical steps improve performance almost as much as valid ones, which suggests the traces work by imitating the style of reasoning rather than by being correct Do reasoning traces show how models actually think?. Models can also hit high prediction accuracy using task-specific shortcuts without building a coherent model of how the world works What makes a world model actually useful for reasoning?.

The obvious test for real reasoning breaks down, though. The classic test asks whether a system can reason about abstract logic regardless of content. Humans fail that test in exactly the same ways LLMs do: both do better on familiar content and worse on unfamiliar content Do language models fail reasoning tests that humans pass?. If being shaped by content disqualifies a system from 'real' intelligence, it disqualifies us too. Some findings also point past simply replaying human discourse. Models can scale their reasoning in hidden internal states without writing any words at all Can models reason without generating visible thinking tokens?, so whatever they're doing isn't only reproducing human sentences. Models fine-tuned on psychology experiments predict human decisions better than the theories humans wrote about themselves Can language models learn to model human decision making?. That's hard to explain as mirroring alone.

Where models clearly differ is in process rather than output. LLM groups reproduce human group-discussion results, but they get there by agreeing sooner and sharing less unique information Do language model groups mimic human group reasoning patterns?. Frontier models can solve problems almost perfectly yet do poorly at spotting flawed steps in someone else's solution Can models that reason well also grade reasoning well?. That uneven profile doesn't look like a human mind, and it doesn't look like a simple mirror either.

The less expected answer comes from the more philosophical notes. Habermas's distinction between watching a system from outside and taking part in a conversation with it helps here. Viewed as systems, humans and LLMs are completely different. Inside a conversation, both draw on the same shared pool of language and meaning Do humans and LLMs differ fundamentally or just superficially?. One note argues that AI output is 'residue' that carries the markers of communication, and that readers do the interpretive work that turns it into what feels like a real exchange Does AI generate genuine utterances or just text patterns?. Another argues that AI separates the finished form of intellectual work from the thinking that normally produces it Does AI separate intellectual form from the thinking behind it?. So a better question than 'does the model think?' may be 'where does the thinking happen?' Some of it happens inside the model, in forms that aren't human language. Some of it is ours, supplied when we read.


Sources 11 notes

Do base models already contain hidden reasoning ability?

Five independent mechanisms—RL steering, critique fine-tuning, decoding changes, SAE feature steering, and RLVR—all elicit reasoning already present in base model activations. Post-training selects rather than creates reasoning; the bottleneck is elicitation, not capability acquisition.

Do reasoning traces show how models actually think?

LLM reasoning traces perform as persuasive appearances rather than reliable explanations of computation. Invalid logical steps perform nearly as well as valid ones, and corrupted traces generalize comparably, showing that semantic correctness is not what produces the performance gains.

What makes a world model actually useful for reasoning?

Research shows LLMs may achieve high prediction accuracy through task-specific heuristics without developing coherent generative models of how the world works. True world models must enable reasoning about interventions and counterfactuals, not surface regularities.

Do language models fail reasoning tests that humans pass?

Research shows both humans and LLMs succeed and fail along the same content-sensitivity axis in reasoning tasks like Wason tests and natural language inference. Content-independence is not a meaningful criterion for distinguishing real reasoning from pattern matching.

Can models reason without generating visible thinking tokens?

Multiple architectures—depth-recurrent models, Heima, and Coconut—demonstrate that test-time compute scales through hidden state iteration rather than token generation. This suggests verbalization is a training artifact, not a reasoning requirement.

Show all 11 sources
Can language models learn to model human decision making?

LLMs finetuned on psychology experiment data predict human behavior more accurately than theory-driven models in decision tasks, capture individual differences in their embeddings, and transfer learning across tasks without task-specific design.

Do language model groups mimic human group reasoning patterns?

LLM groups reproduce the human assembly-bonus asymmetry where discussion helps average members more than top performers, but achieve this through greater conformity, earlier convergence, and less unique information surfacing than human groups.

Can models that reason well also grade reasoning well?

Frontier reasoning models solve problems near-perfectly but score as low as 48% when grading solutions with correct answers but flawed steps. Outcome-focused training rewards answer production, not step-by-step verification, leaving evaluation starved.

Do humans and LLMs differ fundamentally or just superficially?

Applied Habermas's observer/participant distinction to AI: from outside, humans and LLMs are utterly different; from within shared discourse, both draw on the same symbolic substrate, making the difference structural rather than absolute.

Does AI generate genuine utterances or just text patterns?

AI output carries communicative markers inherited from training data but lacks the event structure that produces actual utterances. Users supply the missing orientation through interpretive labor, creating a pseudo-event with structure only on the human side.

Does AI separate intellectual form from the thinking behind it?

Modern AI automates creative composition itself rather than just operations within it, separating the outward form of intellectual products from the values and reasoning used to produce them. This mechanism allows exchange value to float free from use value.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.