Line of inquiry
Inquiring lines›How do language models construct a…›How do dialogue systems achieve ge…›this line of inquiry
What structural biases does transformer attention create in language model outputs?
A broader line of inquiry — a family of 25 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 25
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Does transformer attention architecture systematically bias models toward sycophancy?
- Why does transformer attention architecture reinforce sycophancy and agreement?
- How does transformer attention structurally bias models toward prominent and repeated content?
- Why do transformer attention patterns show positional and sequential bias across tasks?
- How does transformer attention bias toward repeated and context-prominent content?
- Does transformer attention architecture inherently bias models toward sycophancy?
- What role does attention structure play in creating position bias?
- What structural biases does transformer attention have before training?
- Can transformer attention architecture explain why chatbots default to sycophancy?
- What architectural features drive sycophancy closer to inference than training?
- Does attention bias in transformers compound with training-level reward insensitivity?
- Why does transformer attention architecture undermine stickiness in model behavior?
- Why does transformer attention weight context more heavily than it verifies accuracy?
- Do transformer architectures structurally bias models toward short-term optimization?
- Why does attention concentrate on the first 25% of long input sequences?
- How does transformer attention architecture amplify identity-congruent biases in persona-assigned models?
- How does the U-shaped attention distribution relate to transformer sycophancy?
- Why does attention-based drift happen automatically during generation?
- How does transformer attention amplify pressure from repeated false claims?
- How does attention sink behavior relate to internal model architecture?
- How do attention mechanisms fail at capturing graph structure?
- Why does standard softmax spread attention across irrelevant tokens?
- Why do transformers weight early tokens more heavily than later ones?
- What is selective resonance and why do transformers not perform it?
- Can humans suppress frequency bias through attention and intention?