Can models be smart without organized internal structure?
Explores whether linear feature decodability proves genuine compositional reasoning or merely indicates that the right features are present but poorly organized. Critical for understanding what performance metrics actually certify.
Two findings from mechanistic interpretability appear contradictory but operate at different levels of representational analysis:
Fractured Entangled Representations (FER): Since Can identical outputs hide broken internal representations?, SGD-trained models fail catastrophically under perturbation or distribution shift in ways that well-organized representations would not. The pathology is invisible to standard evaluation.
Compositional generalization at scale: Scaling data and model size produces representations where compositional features are linearly decodable — separable task constituents can be independently identified and manipulated. This has been taken as evidence for genuine compositional understanding.
The resolution: Linear decodability tests for the presence of features, not their organization. A fractured representation could contain every linearly decodable feature while being fractured in how those features relate to each other. The compositional parts are present but their composition is broken.
This connects directly to the "imposter intelligence" post angle: Can LLMs understand concepts they cannot apply?, Does supervised fine-tuning actually improve reasoning quality?, and Do foundation models learn world models or task-specific shortcuts?. All describe the same meta-pattern: surface metrics certify capability that internal structure analysis would disqualify.
The practical implication for model evaluation: passing compositional generalization tests does not guarantee robust compositional reasoning. Evaluation under distribution shift, perturbation, and novel recombination is required to distinguish genuine compositionality from fractured representations that happen to contain the right features.
Inquiring lines that read this note 163
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
How can identical external performance mask different internal representations?- Why do only two of fourteen models improve when problem constraints are removed?
- What distinguishes minimal-pair asymmetry from standard accuracy evaluation?
- How do unstated constraints become invisible to training data distributions?
- How do surface statistical regularities enable correct outputs while degrading robustness?
- How do weight perturbations reveal what performance benchmarks cannot measure?
- Why do single function-calling benchmarks mask model weakness in specific areas?
- Can identical model performance mask fundamentally broken internal representations?
- Why do feature-based approaches struggle when privacy or latent factors are involved?
- What makes attractor-based probing better for third-party model auditing than alternatives?
- Do generic kernel-decay assumptions alone explain coarse-to-fine spectral ordering?
- How do coverage and identifiability set separate performance ceilings?
- What benefits do open foundation models create that closed systems cannot?
- How do spectral-norm constraints prevent divergence in world model rollouts?
- When does the right constraint beat additional model capacity?
- What structural constraints matter more than model depth for CF?
- What production constraints should determine paradigm selection?
- How do embedding dimension limits constrain what concept models can represent?
- Why do power-law distributions make standard ML infrastructure assumptions fail?
- Why does the right structural prior matter more than raw model capacity?
- How can expensive models efficiently support cheap models in production?
- How do unstated feasibility constraints affect model decision-making?
- What distinctive properties make open foundation models different from closed ones?
- What design changes could make constraint inference more reliable without explicit cuing?
- What is the mechanistic signature when models chain facts never presented together?
- What architectural properties of deterministic models block multi-solution reasoning?
- Can mechanistic interpretability reveal how ideologies decompose into simpler features?
- Can Kolmogorov complexity alone capture what makes intelligence general?
- How do functional features differ from representational abstract features?
- What makes linear decodability a reliable signal of compositionality?
- What happens when you remove core political features from a deep model?
- Why do models with less steerability have more abstract ideological features?
- Can mechanistic interpretability explain explanation-execution disconnection?
- Can fractured entangled representations hide undetected by standard analysis methods?
- Does the linear representation hypothesis reflect networks or reflect our analysis tools?
- Can representation engineering cleanly isolate single features in entangled semantic space?
- What are fractured entangled representations in neural networks?
- How do sparse circuits compare to the modular subnetworks that emerge naturally?
- Can sparse approximations reveal interpretable structure hidden in existing dense models?
- Can geometric structure in representations exist without supporting functional mechanisms?
- How does mechanistic interpretability complement learning mechanics in explaining deep learning?
- What distinguishes a representational feature from a causally inert correlation?
- Can interventions on model components prove mechanism without explaining encoding?
- Can representation analysis methods detect complex features models compute with?
- How can neural networks be interpretable by design rather than post-hoc?
- How do mechanistic features compare to natural language for interpretability?
- What physical structure does a Gaussian-regularized latent space actually encode?
- What makes regularization an implicit factor in embedding geometry?
- Do feature extraction methods systematically miss computationally important complex features?
- What makes a feature abstract versus concrete in neural network activations?
- What prevents representation collapse in latent-prediction world models like JEPA?
- How should benchmarks test whether models fit algorithms or patterns?
- How does optimizing model performance decouple from optimizing user interpretability?
- How much do metric choices inflate claims about model capabilities?
- How should researchers evaluate whether correct model outputs reflect real structural learning?
- Why do text-only benchmarks underestimate deployed model capability?
- What makes well-formatted outputs misleading as evidence of model capability?
- How can benchmark accuracy scores mask the absence of interpretable reasoning structure?
- How does requential coding measure true simplicity without parameter count inflation?
- Can likelihood choice matter more than architectural depth for CF?
- How does uniform code distribution make items more distinguishable?
- What sparse high-rank patterns does the deep tower fail to capture?
- Why do structural signals across edges resist noise better than single-edge counts?
- How does nesting optimization levels improve on traditional network depth?
- What compression explains why syntax fits in low-dimensional subspaces?
- Can steering vectors be combined with other compression techniques?
- Can entropy signatures alone detect whether context was model-generated or externally prefilled?
- Why do parameter-based compressors fail to measure true model simplicity?
- How should product specifications measure alignment without naming the dimension?
- Can alignment methods like DPO exploit or correct these surface feature biases?
- What makes AI-discovered architectures reveal design principles invisible to humans?
- Does architectural discovery follow an empirical scaling law like neural networks?
- Can a single SAE feature control reasoning behavior across model families?
- Can steering vectors prove that representations are genuinely organized?
- How does LatentQA differ from predefined concept steering like representation engineering?
- What other behavioral properties exist as linear directions in activation space?
- Can bilevel autoresearch succeed when the inner and outer loops use different models?
- How much does domain shift limit the mechanisms a bilevel system can autonomously discover?
- Can a world model have rich representations without adequate data coverage?
- Does model collapse occur across different architectures or only in specific conditions?
- Can seedless generation maintain explainability while scaling control?
- Do larger models develop more abstract features than smaller ones?
- Why do text-to-image models fail at composing multiple concepts together?
- Does scaling data automatically produce compositional reasoning or just better feature encoding?
- What test distinguishes genuine compositionality from fractured feature presence?
- Can granular function calling tasks learn composition from graph-sampled data?
- What makes structured stochasticity more effective than unstructured randomness in reasoning?
- Does sparsity enforce compositional structure or merely amplify existing modularity?
- Why does gradient descent discover compositional structure without explicit pressure?
- What architectural alternatives can capture compositional structure beyond pooled cosine?
- How does scaling and training data enable compositional behavior without symbolic mechanisms?
- What task structures benefit most from geometric parameter merging?
- When should model isolation be preferred over weight-averaging approaches?
- What performance trade-offs emerge when composing multiple independently trained model capabilities?
- Can we predict which tasks will decompose into modular subnetworks?
- How does discretization make item representations more distinguishable?
- Do multi-vector or cross-encoder models escape these dimensional constraints?
- Why is a combinatorial framework better than family resemblance classification?
- Can spectral eigenvector ordering serve as a model-agnostic interpretability probe?
- Can generative reconstruction preserve latent manifold structure better than geometric compression?
- How does representation-level reranking address residual gaps after decomposition?
- Can structural perturbations harm model accuracy more than semantic ones?
- Does parameter composition work when adapter alignment is imperfect?
- Why does capturing domain structure reduce data requirements more than raw volume?
- Why do models fail on logically equivalent tasks with different data distributions?
- Why do energy-based models generalize better on out-of-distribution data than standard transformers?
- How does adjacent layer sharing differ from non-adjacent weight reuse?
- Why do standard transformers fail to encode recursive structure in their hidden states?
- Does Gemma's transformer explicitly exploit the inherited hierarchical geometry?
- How do pre-norm layers enable reliable fixed-point halting signals?
- What makes some model capabilities reliable while others remain brittle?
- Can end-to-end models maintain debuggability without modular components?
- Why does the gap between theoretical expressiveness and learned capability matter?
- Can ensemble predictions be distilled back into a single deployable model?
- Why do linear research pipelines lose global context across planning and generation steps?
- When does backward decomposition fail on open-ended or unstructured tasks?
- Why do metric choices constrain which model capabilities get developed?
- What features does a sample reinforce when it moves bands?
- Why does weight sparsity reduce superposition and force disentangled representations?
- Can sparsity patterns reliably indicate how well a model knows its input?
- What distinct structural signatures do model repetition and topic volatility create?
- Can we balance interpretability with the efficiency gains of compressed inter-model communication?
- Can you steer reasoning by directly manipulating SAE features?
- What limits external scaling when a model lacks reasoning foundation?
Related concepts in this collection 1
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Can we track and steer personality shifts during model finetuning?
This research explores whether personality traits in language models occupy specific linear directions in activation space, and whether we can detect and control unwanted personality changes during training using these geometric directions.
persona vectors demonstrate a case where linear decodability corresponds to genuine functional organization (steering works), providing a positive counterexample to FER's warning that decodability alone is insufficient
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Towards Principled Evaluations of Sparse Autoencoders for Interpretability and Control
- Comprehension Without Competence: Architectural Limits of LLMs in Symbolic Computation and Reasoning
- Titans: Learning to Memorize at Test Time
- Break It Down: Evidence for Structural Compositionality in Neural Networks
- Towards Monosemanticity: Decomposing Language Models With Dictionary Learning
- Computational structuralism: Toward a formal theory of meaning in the age of digital intelligence
- Farther the Shift, Sparser the Representation: Analyzing OOD Mechanisms in LLMs
- Large Language Model Reasoning Failures
Original note title
identical performance metrics can mask fundamentally different internal representations — feature linear decodability does not guarantee representational organization