Theme of inquiry

How do reasoning capabilities emerge and evolve across model scales?

A question within its area, explored through 10 lines of inquiry below — each a family of specific questions the research asks.


Why do correct reasoning traces tend to be shorter than incorrect ones?

39 specific questions

See all 39 questions in this line of inquiry
Does chain-of-thought text faithfully represent the model's actual reasoning?

31 specific questions

See all 31 questions in this line of inquiry
Do LLMs understand causality or merely semantic associations?

32 specific questions

See all 32 questions in this line of inquiry
When does chain-of-thought reasoning improve performance and when does it fail?

31 specific questions

See all 31 questions in this line of inquiry
Why is chain-of-thought effective despite invalid reasoning?

37 specific questions

See all 37 questions in this line of inquiry
Does self-reflection help reasoning models identify and fix errors?

26 specific questions

See all 26 questions in this line of inquiry
Do thinking tokens beyond critical thresholds improve or degrade reasoning?

47 specific questions

See all 47 questions in this line of inquiry
What mechanistic structures enable genuine reasoning in language models?

40 specific questions

See all 40 questions in this line of inquiry
What causes language models to reason excessively or insufficiently relative to problem difficulty?

28 specific questions

See all 28 questions in this line of inquiry
Why doesn't reasoning effort improve theory of mind performance?

21 specific questions

See all 21 questions in this line of inquiry