Theme of inquiry

How do representation and aggregation choices affect model alignment and reliability?

A question within its area, explored through 5 lines of inquiry below — each a family of specific questions the research asks.


Does alignment training create genuine alignment or just output compliance?

51 specific questions

See all 51 questions in this line of inquiry
What do systematic disagreements between annotators reveal about ground truth?

20 specific questions

See all 20 questions in this line of inquiry
What enables genuine semantic understanding in language models?

77 specific questions

See all 77 questions in this line of inquiry
Why do token-level mechanisms matter for learning to reason?

62 specific questions

See all 62 questions in this line of inquiry
How do training data properties determine the emergence of internal misalignment?

45 specific questions

See all 45 questions in this line of inquiry