Theme of inquiry
What causes systematic gaps in LLM generalization across domains?
A question within its area, explored through 1 line of inquiry below — each a family of specific questions the research asks.
137 specific questions
- Why does LLM knowledge fail to influence their actual outputs?
- Should LLM reasoning be studied as latent state trajectories rather than surface text?
- Does LLM reasoning always match the outputs it generates?
- Why does chain-of-thought reasoning alone not fix LLM performance with users?
- Do LLMs rely on surface statistical patterns instead of causal structure?
- Can evidence density alone shift an LLM from generation to reasoning?
- Can LLMs improve at simple deduction through different training approaches?