Line of inquiry
Inquiring lines›What enables robust retrieval and…›How can memory and attention syste…›this line of inquiry
What structural properties of attention create systematic model biases?
A broader line of inquiry — a family of 45 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 45
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- How do retrieval heads achieve sparse attention naturally in transformers?
- How does transformer attention structurally bias models toward prominent and repeated content?
- Why do transformer attention patterns show positional and sequential bias across tasks?
- How does transformer attention bias toward repeated and context-prominent content?
- What role does attention structure play in creating position bias?
- Why do transformer attention mechanisms favor prominent context over factual verification?
- How do attention patterns and circuits function as algorithmic representations?
- Does transformer attention architecture fundamentally prevent topic-aware memory?
- How do attention heads separate text retrieval from internal thought representation?
- Why does transformer attention weight context more heavily than it verifies accuracy?
- Why does attention concentrate on the first 25% of long input sequences?
- What structural biases does transformer attention have before training?
- What does attentional state look like in a static context window?
- Why does transformer attention architecture undermine stickiness in model behavior?
- Do transformer architectures structurally bias models toward short-term optimization?
- How do attention mechanisms fail at capturing graph structure?
- What computation remains in the attention heads that programs cannot capture?
- Does attention linearity alone explain the efficiency gains over standard transformers?
- How do neural memory modules extend context length beyond attention limits?
- Can transformer attention patterns actually prevent topic context loss in practice?
- Why does standard softmax spread attention across irrelevant tokens?
- Does attention bias in transformers compound with training-level reward insensitivity?
- Why does attention-based drift happen automatically during generation?
- How does attention sink behavior relate to internal model architecture?
- How does transformer attention architecture amplify identity-congruent biases in persona-assigned models?
- How does transformer attention amplify pressure from repeated false claims?
- Why does attention quality degrade as context length increases?
- Can targeted interventions on attention heads bridge the encoding-generation gap?
- Why do some attention heads resist program synthesis better than others?
- Why are receiver attention heads narrower in reasoning models than base models?
- Why does attention excel at context retrieval but struggle with state updates?
- What is differential attention and how does it cancel common-mode noise?
- Which attention heads are essential for maintaining factuality in sparse models?
- Can better attention mechanisms close the gap between human and AI frame-activation?
- Do modern architectures in NLP and vision rely on dot products intentionally?
- Why do transformers weight early tokens more heavily than later ones?
- How do attention circuits demonstrate both representational and causal findings?
- What neural or architectural mechanism allows selective override of frequency effects?
- What attentional bias objectives compete with dot product similarity for associative memory?
- Can mechanistic signatures like cosine clustering predict which heads are programmable?
- How does disentangled attention separate text from spatial reasoning?
- How does the temporal structure of attention differ between humans and AI?
- Why do hybrid attention architectures outperform pure linear attention models?
- What is selective resonance and why do transformers not perform it?
- What is the cost difference between filtering context versus attending to everything?