Theme of inquiry

How do attention and architecture shape the biases in retrieval?

A question within its area, explored through 8 lines of inquiry below — each a family of specific questions the research asks.


Why do reward structures fail to shape long-term agent learning?

30 specific questions

See all 30 questions in this line of inquiry
Can alternative training methods improve on supervised fine-tuning for language models?

34 specific questions

See all 34 questions in this line of inquiry
How do aggregate reward models systematically exclude minority user preferences?

33 specific questions

See all 33 questions in this line of inquiry
Can language model RL training avoid reward hacking and misalignment?

38 specific questions

See all 38 questions in this line of inquiry
How can process reward models supervise complex reasoning traces?

29 specific questions

See all 29 questions in this line of inquiry
What constrains reinforcement learning's ability to expand model reasoning?

46 specific questions

See all 46 questions in this line of inquiry
What properties determine whether reward signals teach genuine reasoning?

63 specific questions

See all 63 questions in this line of inquiry
How do policy learning algorithm choices affect multi-objective optimization stability?

14 specific questions

See all 14 questions in this line of inquiry