INQUIRING LINE

AI models often learn shortcuts that predict right answers without actually understanding how things work underneath.

Why do foundation models rely on surface patterns instead of causal structure?

This explores why large AI models tend to pick up shortcuts that predict the right answer instead of learning how the world actually works, and what can be done about it.


This explores why large AI models tend to pick up shortcuts that predict the right answer instead of learning how the world actually works. The corpus mostly shows that this happens and how to detect it. It says less about the deeper why, but the evidence points to one answer: nothing in training forces a model to go further than shortcuts, as long as the shortcuts score well. The clearest case is Do foundation models learn world models or task-specific shortcuts?. Transformers trained on orbital mechanics predict planetary motion well. But when researchers probe what the models learned, there is no Newtonian law inside, only patterns that work in one setting and fall apart in the next. Even arithmetic turns out to rely on rough 'this number is in about this range' rules rather than an actual procedure for adding.

The same pattern shows up in language, under different terms. Can models pass tests while missing the actual grammar? finds that models can pass grammar tests by relying on sentence length, word choice and spelling instead of grammatical rules. Here is the part you might not expect: standard benchmarks can't tell the difference. If a cheap cue and the real structure give the same answer on the test, the model has no reason to learn the harder one, and we have no way to see which one it learned. Can models be smart without organized internal structure? makes this stronger. A model can hold every feature a task needs, readable and accurate, while its internal organization is still fractured. That damage stays invisible until the inputs shift or someone perturbs the model. So 'it got 100%' and 'it understands' can come apart completely.

If the models won't learn causal structure on their own, one option is to stop asking them to. Can separating causal models from language models improve reasoning? describes an architecture where an explicit causal model does the reasoning and can revise itself, and the LLM only translates inputs and outputs into language. This sidesteps the false-pattern problem instead of hoping training fixes it. Can structural causal models automate social science with language models? goes the other way: a causal model guides the LLM through social experiments. The finding there is revealing. The simulations reliably get the direction of an effect right but not its size, which fits a system that knows roughly which way things go without a real mechanism underneath.

There are two complications worth knowing. First, causal structure may not be the whole goal. Can causal models alone capture how humans actually reason? points out that human reasoning also runs on association, analogy and emotion, so a purely causal model would miss real parts of how people think. Some of what looks like shallow pattern-matching may be close to how humans reason too. Second, because models settle for patterns, they can't check themselves against reality. Do foundation models actually reduce our need for real data? argues that without real data to test against, repeatedly refining prompts just bounces your own beliefs back to you. The surprising takeaway is that more capable models make real-world data more important, not less.


Sources 7 notes

Do foundation models learn world models or task-specific shortcuts?

Inductive bias probes show transformers trained on orbital mechanics and games learn predictive patterns, not unified world structure. Fine-tuning reveals nonsensical, slice-dependent laws; circuit analysis shows arithmetic relies on range-matching heuristics, not algorithms.

Can models pass tests while missing the actual grammar?

BabyLM evaluations showed models can produce correct outputs by relying on sentence length, word choice, and orthography rather than grammatical structure. Standard benchmarks cannot distinguish these two generalization types without tests specifically designed to rule out surface heuristics.

Can models be smart without organized internal structure?

Models trained with SGD can contain all the linearly decodable features needed for a task while maintaining fundamentally broken internal organization. This makes them vulnerable to perturbation and distribution shift invisible to standard evaluation metrics.

Can separating causal models from language models improve reasoning?

Causal Reflection separates causal reasoning into a formal dynamic model with a Reflect mechanism for revision, relegating the LLM to structured inference and language rendering. This architecture sidesteps asking LLMs to perform causal reasoning directly, addressing both spurious-correlation failures and RL's explanation gap.

Can structural causal models automate social science with language models?

LLMs guided by structural causal models can propose and test causal hypotheses across negotiation, bail, interview, and auction scenarios. Simulations reveal effect directions reliably but not magnitudes, making them useful for directional social science.

Show all 7 sources
Can causal models alone capture how humans actually reason?

Causal belief networks excel at modeling causal reasoning but cannot represent associative links, analogical mappings, or emotion-driven belief shifts. The GenMinds framework itself acknowledges this as a tractable starting point rather than a complete theory.

Do foundation models actually reduce our need for real data?

Powerful foundation models don't eliminate the need for real data—they heighten it. Without empirical anchoring, iterative prompt refinement creates epistemic circularity where users confirm their own beliefs rather than test them.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.