Theme of inquiry

How does chain-of-thought reasoning affect model capability and monitorability?

A question within its area, explored through 8 lines of inquiry below — each a family of specific questions the research asks.


How reliably can language models perform causal versus temporal reasoning?

41 specific questions

See all 41 questions in this line of inquiry
How do thinking tokens exhibit diminishing returns in reasoning?

48 specific questions

See all 48 questions in this line of inquiry
Does chain-of-thought reasoning reveal how models actually think or merely imitate reasoning?

106 specific questions

See all 106 questions in this line of inquiry
How does scaling reasoning capabilities affect models' appropriate abstention behavior?

47 specific questions

See all 47 questions in this line of inquiry
Does scaling reasoning capability create fundamental tradeoffs in control and reliability?

85 specific questions

See all 85 questions in this line of inquiry
Can reasoning models use reflection to correct their initial outputs?

34 specific questions

See all 34 questions in this line of inquiry
Can reasoning traces reveal actual model reasoning versus plausible output?

86 specific questions

See all 86 questions in this line of inquiry
Why don't better reasoning capabilities improve theory of mind performance?

30 specific questions

See all 30 questions in this line of inquiry