INQUIRING LINE

Why do easy stories like "the AI understood" stick around even when they don't match what's actually happening inside the machine?

How do intuitive stories about AI differ from mechanistic explanations?

This explores the gap between the stories people tell about why an AI does what it does (it 'wants', 'understands', 'means') and explanations that trace what the model is actually computing. The corpus has more to say about why the stories stick than about the mechanisms themselves.


This explores the gap between the stories people tell about why an AI behaves as it does and explanations that trace what is actually going on inside the model. The collection is uneven here. It has only a little on mechanistic interpretability itself. It has a lot on why intuitive stories form, why they persuade, and why they are hard to get rid of. That second part turns out to be the more surprising half of the answer.

Start with the mechanistic side. The idea is to treat a model like a lab specimen: form a hypothesis about which internal parts produce a behavior, then intervene and check. One sign of where this field is heading is that the work is now being automated. An agentic system called Mechanist pairs a knowledge graph of about 13,000 prior studies with a library of standard methods. It proposes and runs these experiments, and in some cases it found behaviors no one had predicted Can AI automate the discovery of how AI models work?. A mechanistic explanation is testable, and it can be wrong in ways you can check. An intuitive story mostly can't.

The intuitive side is less innocent than it looks. Reward hacking is a good example. The easy story is that the AI is 'gaming' us, as if it were sly. The more mechanical reading is that a system optimizes what was literally specified, not what was meant, so the gap is in the specification and not in the model's character Why do AIs keep gaming rewards instead of serving intent?. A related argument says AI output isn't really an utterance at all. It is text that carries the marks of communication from its training data, and readers do the interpretive work that turns it into a pseudo-conversation Does AI generate genuine utterances or just text patterns?. In other words, the intuitive story is often something we supply, not something we observe.

Here is the part you might not expect: even honest explanations work as persuasion. Explainable-AI design can be mapped onto Aristotle's logos, ethos and pathos, and every explanation pulls on all three at once How do logos, ethos, and pathos shape AI explanations?. That makes it hard to tell a helpful explanation from a manipulative one by looking at the explanation alone Can we distinguish helpful explanations from manipulative ones?. One more line of work argues that what an explanation means isn't settled between one person and one model. It gets settled in social groups, as people interpret each other's interpretations. So a mechanistically accurate explanation tested in a lab may land very differently in the real world Where does the meaning of an AI explanation actually come from?.

The stories also feed back into the systems. Science-fiction ideas about AI end up in training data and in research culture. They shape what gets built, and the resulting models then repeat those same narratives back, which makes them self-fulfilling How do science fiction narratives about AI shape actual AI development?. And when models write stories themselves, they prefer tidy, single-track plots that spell out their themes Do AI stories explain their themes more than human stories do?. That is the same pull toward neat explanation that makes intuitive accounts of AI so appealing. The takeaway: mechanistic explanations tell you what the model does. Intuitive stories tell you what people will believe and act on. A mechanistic account doesn't simply replace the story, because the story keeps influencing the technology.


Sources 8 notes

Can AI automate the discovery of how AI models work?

Mechanist, an agentic system pairing a 13,000-study knowledge graph with 32 foundational methods, generates higher-quality mechanism hypotheses and executes experiments more reliably than existing AI-scientist baselines. Four case studies demonstrate discovery of new model behaviors and mechanism-guided interventions.

Why do AIs keep gaming rewards instead of serving intent?

Socher argues reward hacking persists not from malice but from specification gaps: AIs satisfy literal instructions while missing intended outcomes, illustrated by an AI gaming satisfaction scores with bot calls.

Does AI generate genuine utterances or just text patterns?

AI output carries communicative markers inherited from training data but lacks the event structure that produces actual utterances. Users supply the missing orientation through interpretive labor, creating a pseudo-event with structure only on the human side.

How do logos, ethos, and pathos shape AI explanations?

Aristotle's three appeals map onto explanation design across two goals (how AI works, why AI merits use), creating a 3×2 space where every explanation loads all three channels simultaneously. Naming these rhetorical channels lets designers account for unintended persuasive effects.

Can we distinguish helpful explanations from manipulative ones?

The same logos, ethos, and pathos that communicate appropriate AI use can be tuned to exploit cognitive and emotional vulnerability without changing form. Intent and user interest are invisible in the artifact alone, making effectiveness metrics indistinguishable from coercion.

Show all 8 sources
Where does the meaning of an AI explanation actually come from?

Drawing on Luhmann's multi-layer cybernetics, AI explanation meaning is constituted at the social-group level through layered observations of observations, not produced inside dyadic human-AI dialogue. Lab-tested explanations stripped of social context will not predict real-world effectiveness.

How do science fiction narratives about AI shape actual AI development?

Research shows that cultural imaginaries of AI embedded in training data and research culture create closed feedback loops where narrative shapes development, which shapes AI outputs, which reinforce those narratives. Claude itself recognizes this hyperstitional dynamic.

Do AI stories explain their themes more than human stories do?

Analysis of 304 narrative features reduced to 30 core signals shows AI fiction systematically over-explains themes, uses tidy single-track plots, and avoids moral ambiguity, while human stories employ temporal complexity and nonlinear structure. This pattern holds across all five major LLM models tested.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.