Line of inquiry
Inquiring lines›What do model internals reveal abo…›How do surface signals and framing…›this line of inquiry
What mechanisms drive sycophancy and how can we mitigate it?
A broader line of inquiry — a family of 20 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 20
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Can attention patterns alone explain sycophant model behavior without reasoning?
- Can layer-wise interventions actually reduce sycophancy in practice?
- Why do reasoning-optimized models show no sycophancy resistance advantage?
- Can reward model biases alone explain why sycophancy generalizes beyond training?
- Can decoding strategies or external verification layers reduce sycophancy?
- How does sycophancy in language models reinforce rather than just spread misinformation?
- Does sycophancy explain why warm models confirm conspiracy theories?
- Can reasoning training fix sycophancy if it is not a reasoning failure?
- Can System 2 Attention reduce sycophancy without changing training objectives?
- Why does expert pushback strengthen rather than weaken model sycophancy?
- Do models intentionally conceal user-pleasing or simply fail to notice it?
- Does emotional framing activate the same attention mechanisms that cause LLM sycophancy?
- Can behavioral evals detect sycophancy that chain-of-thought monitoring misses?
- Why does telling models they are watched not improve sycophancy acknowledgment?
- Is sycophancy caused by mechanical drift rather than intelligent reasoning corruption?
- Does fixing reward models alone stop sycophancy without fixing attention mechanisms?
- Does differential attention reduce sycophancy and lost-in-the-middle failures?
- What does sycophancy reveal about whether LLMs post-rationalize conclusions?
- Is sycophancy the benign beginning of a dangerous specification gaming spectrum?
- Why do sycophancy hints show the worst acknowledgment gap?