When a chess-playing AI learns concepts inside itself, do they match the ideas people use to explain chess?
How do concept vectors in neural networks relate to human-language chess explanations?
This explores whether the concepts a game-playing neural network learns internally (stored as directions, or 'vectors', in its activations) line up with the ideas humans use when they explain chess, and whether a model's explanations actually reflect what it computes.
This explores whether the concepts a neural network learns internally match the ideas people use to explain chess, and whether a model's explanation of a move tells you anything about how it chose the move. A note up front: this collection has no papers on chess engines or on probing chess concepts directly. What it does have is a cluster of work on what concept vectors are, how they can be read out of a network, and why an explanation in words can come apart from the computation underneath. That gives you a useful way to think about the chess question.
Start with what a 'concept vector' is. Mechanistic interpretability treats the most basic kind of understanding as a feature that sits along a direction in activation space Do language models understand in fundamentally different ways?. That same work describes two higher tiers: knowing facts about the state of the world, and compact circuits that apply principles. The key finding is that the higher tiers don't replace the lower ones. They sit alongside heuristic shortcuts, so the result is a patchwork. A chess network could hold a clean 'king safety' direction and still pick many moves through shallow pattern-matching. These directions can also carry real structure. The Polar Probe found that language models encode grammar using both distance and angle between embeddings, which is the kind of symbol-compatible geometry that would let a learned vector match a nameable human concept How do language models encode syntactic relations geometrically?. Networks also tend to split compositional tasks into separate subnetworks Do neural networks naturally learn modular compositional structure?, which is one reason you might expect separate chess ideas to end up in separate, findable places.
The twist is that finding a concept and explaining with it are separate abilities. Research on 'Potemkin understanding' shows models that explain a concept correctly, fail to apply it, and then recognize that they failed. That pattern points to explanation and execution running on largely disconnected pathways Can LLMs understand concepts they cannot apply?. Applied to chess, a fluent paragraph about 'controlling the center' may have little to do with the internal vector that actually drives the move. The 'imposter intelligence' work goes further. Networks with identical outputs on every input can have very different internal representations, so playing strength alone can't tell you whether the inside is organized around human-like concepts or around a tangle that happens to work Can AI pass every test while understanding nothing?.
The game-reasoning notes come at this from the other side. When language models reason about strategic games, they move away from optimal play as the game gets more complex. Structured step-by-step workflows bring them back close to optimal Do language models make rational strategic decisions in games?. Different models also show distinct reasoning styles, such as minimax-style calculation versus anticipating the other player's beliefs Do large language models use one reasoning style or many?. So a model's verbal reasoning in a game can be informative, but mostly when something outside the model holds it in place, not because the words faithfully report internal state. A related finding fits here: when a model's trained-in associations are strong, prompting it in text often can't change its behavior, and researchers had to intervene directly on its internal representations Why do language models ignore information in their context?. That is roughly how you would test whether a 'chess concept vector' really matters: change the vector and see whether the move changes.
The takeaway you might not have expected: finding a vector that matches a human chess concept is the easy half. The harder question is whether that vector causes the model's play, and whether a language explanation is connected to it at all. The collection suggests three separate layers: what the network represents, what drives its choices, and what it says. None of them is guaranteed to line up with the others, so 'human-language chess explanations' are best treated as claims to check against the vectors, not as reports of what the vectors do.
Sources 8 notes
Mechanistic interpretability reveals conceptual understanding (features as directions), state-of-world understanding (factual connections), and principled understanding (compact circuits). Crucially, higher tiers coexist with lower-tier heuristics rather than replacing them, creating a patchwork of capabilities.
The Polar Probe shows LLMs represent syntactic type and direction through both distance and angular position between embeddings, nearly doubling accuracy over distance-only methods. This demonstrates neural networks spontaneously learn structured, symbolic-compatible geometry.
Pruning experiments reveal that neural networks implement compositional subroutines in isolated subnetworks, with ablations affecting only their corresponding function. Pretraining substantially increases the consistency and reliability of this modular structure across architectures and domains.
Models can explain concepts accurately, fail to apply them, and recognize the failure—a triple pattern incompatible with human cognition. This indicates functionally disconnected explanation and execution pathways rather than simple knowledge gaps.
The Fractured Entangled Representation hypothesis shows that SGD-trained networks can produce identical outputs across all inputs while maintaining radically different internal representations. Standard benchmarks cannot detect this structural difference.
Show all 8 sources
LLMs frequently fail to compute Nash equilibria, with worse performance as game complexity increases. Structured game-theoretic workflows guide reasoning toward optimal strategies, reducing exploitability and enabling near-optimal negotiation outcomes.
Analysis of 22 LLMs across behavioral game theory reveals three dominant profiles: GPT-o1 uses minimax reasoning, DeepSeek-R1 uses trust-based reasoning, and GPT-o3-mini uses belief-anticipation. Performance correlates with game structure, not raw reasoning depth.
Research demonstrates that LMs generate outputs inconsistent with their context because parametric knowledge from training dominates over in-context information. Textual prompting alone cannot override strong priors; causal intervention in representations is required.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Break It Down: Evidence for Structural Compositionality in Neural Networks
- Game-theoretic LLM: Agent Workflow for Negotiation Games
- LLM Strategic Reasoning: Agentic Study through Behavioral Game Theory
- Strategic Reasoning with Language Models
- Word Meanings in Transformer Language Models
- Can Large Language Models Develop Strategic Reasoning? Post-training Insights from Learning Chess
- Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey
- Computational structuralism: Toward a formal theory of meaning in the age of digital intelligence