Line of inquiry
Inquiring lines›How do training choices shape mode…›How do training dynamics and archi…›this line of inquiry
What organizational structures emerge in learned representations?
A broader line of inquiry — a family of 48 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 48
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Can geometric structure in representations exist without supporting functional mechanisms?
- Can a world model have rich representations without adequate data coverage?
- How do semantic features in representations become steerable task-specific directions?
- What role does a model's representational structure play in learning?
- Can steering vectors prove that representations are genuinely organized?
- Can spectral eigenvector ordering serve as a model-agnostic interpretability probe?
- How deeply are ideological structures represented in large language models?
- Can sparse approximations reveal interpretable structure hidden in existing dense models?
- Can attractor dynamics compete with input-based probing for characterizing model knowledge?
- How do surface statistical regularities enable correct outputs while degrading robustness?
- What role does embedding space geometry play in multi-hop reasoning?
- Does base model geometry predict which associations persist through intervention?
- Why must world models be nested rather than flat and uniform?
- Can representation analysis methods detect complex features models compute with?
- Can representation engineering cleanly isolate single features in entangled semantic space?
- How does trajectory geometry relate to the need for chain-of-thought reasoning?
- Does the same spectral signature appear across different embedding models?
- Do language models and multimodal models show similar attractor-based interpretability?
- Can graph cyclicity and topology predict when reasoning systems achieve breakthrough insights?
- Do reading vectors from activation space causally control model behavior?
- Does Gemma's transformer explicitly exploit the inherited hierarchical geometry?
- Does sequence prediction accuracy prove an underlying world model exists?
- What geometric structure do language models actually use during inference?
- Why do models with less steerability have more abstract ideological features?
- How do world models decompose between representation of facts versus generative mechanisms?
- What prevents representation collapse in latent-prediction world models like JEPA?
- Can curvature measurements predict task difficulty without behavioral labels?
- Can generative reconstruction preserve latent manifold structure better than geometric compression?
- What makes regularization an implicit factor in embedding geometry?
- How do weights, selection, and prompts create different geometric landscapes of accessible behaviors?
- How do encode-decode contractive biases create stable attractors in latent space?
- What physical structure does a Gaussian-regularized latent space actually encode?
- What makes multimodal conditioning effective when features are decomposed to the right granularity?
- Can modular expert decomposition extend beyond time into other causal dimensions?
- How do weight visualizations reveal temporal structure in cyclic training?
- What makes internal embeddings useful as multimodal input for language model training?
- What test distinguishes genuine compositionality from fractured feature presence?
- How do embedding dimension limits constrain what concept models can represent?
- What spectral signatures distinguish hierarchy-driven geometry from corpus-driven geometry?
- What scaffolding tools help users specify implicit contextual boundaries to models?
- How do low-dimensional representation structures entangle multiple cultures together?
- Why does integrating world models with decision-making systems matter?
- How do repetition and inefficiency register as measurable trajectory features?
- How do you measure the depth of political representation inside a language model?
- Do generic kernel-decay assumptions alone explain coarse-to-fine spectral ordering?
- What other behavioral properties exist as linear directions in activation space?
- What happens when you remove core political features from a deep model?
- Which hyperparameter theories best explain universal behaviors across neural networks?