Theme of inquiry

Can visible reasoning improve model evaluation and safety?

A question within its area, explored through 8 lines of inquiry below — each a family of specific questions the research asks.


Do language models reason through causal mechanisms or semantic associations?

35 specific questions

See all 35 questions in this line of inquiry
Does chain-of-thought reasoning reveal genuine computation or imitate patterns?

80 specific questions

See all 80 questions in this line of inquiry
Why doesn't reasoning volume improve theory of mind performance?

29 specific questions

See all 29 questions in this line of inquiry
Can reasoning traces and behavior monitoring reliably detect hidden AI scheming?

24 specific questions

See all 24 questions in this line of inquiry
How does reasoning length affect model performance across different tasks?

77 specific questions

See all 77 questions in this line of inquiry
Why do stronger reasoning capabilities create tradeoffs with instruction following?

69 specific questions

See all 69 questions in this line of inquiry
What is the relationship between thinking tokens and reasoning accuracy?

47 specific questions

See all 47 questions in this line of inquiry
How do soft reasoning mechanisms explore multiple paths without explicit training?

7 specific questions