Line of inquiry
Inquiring lines›How can we optimize language model…›How do reasoning capabilities emer…›this line of inquiry
When does chain-of-thought reasoning improve performance and when does it fail?
A broader line of inquiry — a family of 31 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 31
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- What three factors actually drive chain of thought performance improvements?
- What makes diffusion chain-of-thought reasoning qualitatively different from sequential chain-of-thought?
- Does chain-of-thought reasoning amplify bullshit or just make it more visible?
- What happens to chain-of-thought performance across distribution shifts?
- Why does chain-of-thought fail when problems lack matching training schemata?
- Does chain-of-thought reasoning specifically improve performance on metalinguistic tasks?
- How much of chain-of-thought reasoning is actually redundant?
- What structural properties define effective long chain-of-thought reasoning?
- How does chain-of-thought training change higher layer computations?
- Does chain-of-thought monitoring fundamentally degrade under optimization pressure?
- Why does chain-of-thought fail to improve multimodal model perception performance?
- Why does extended chain-of-thought reasoning fail to improve numerical optimization performance?
- Why does long CoT training optimize for structural coherence over content correctness?
- What makes o1's chain-of-thought processing specifically effective for exploration tasks?
- Does chain-of-thought reasoning improve mental state tracking in dialogue?
- How does difficulty level change whether extended thinking provides genuine reasoning signal?
- What makes some bottlenecks invisible to chain-of-thought training?
- What distinguishes metacognitive regulation from standard chain-of-thought reasoning?
- Does deep-thinking ratio measure computational effort better than chain-of-thought length?
- What makes token-level reasoning during pretraining different from test-time chain-of-thought?
- Can extended reasoning training capture individual strategic thinking styles?
- How do thought anchors differ from individual forking tokens mechanistically?
- What distinguishes graph-of-thought reasoning from other structured reasoning topologies?
- What are the six types of reasoning steps that appear in chain-of-thought?
- What is the relationship between reasoning depth and verbalization requirements?
- How does interaction horizon differ from chain-of-thought depth?
- How do progressive abstraction chains differ from branching reasoning topologies?
- How does graph of thoughts enable divide-and-conquer reasoning patterns?
- How much does chain-of-thought reasoning narrow the decompression gap?
- How does chain of thought amplify specific forms of rhetorical bullshit?
- How does the three-component definition apply to test-time scaling laws?