Theme of inquiry
Where and why do LLM capabilities unexpectedly break down?
A question within its area, explored through 3 lines of inquiry below — each a family of specific questions the research asks.
79 specific questions
- Does structured decomposition improve LLM reasoning in other compound tasks?
- Can symbolic solvers reliably replace LLM reasoning for logical tasks?
- Can reasoning chains work without logical validity?
- Why does augmenting symbolic reasoning outperform replacing it entirely?
- Can LLMs successfully translate natural language into formal solver specifications?
- What makes structural logic correlate so strongly with contextual consistency?
- Can language models perform purely symbolic reasoning when semantics are removed?
59 specific questions
- How does implicit meaning processing limit LLM pragmatic reasoning?
- Can LLM semantic representations exist without causally influencing their generation output?
- Can explicit connectives compensate for missing intentional tracking in LLMs?
- Why do LLMs perform better on explicit discourse connectives than implicit relations?
- Do language models need words to think or just latent structure?
- Why do language models fail at implicit discourse relations while handling explicit connectives?
- Why do explicit discourse connectives help LLMs but implicit relations cause failures?
99 specific questions
- Why do LLM outputs match researcher priors without solving tasks correctly?
- Why does LLM knowledge fail to influence their actual outputs?
- How faithful are natural language explanations from LLMs really?
- Do standard language benchmarks underestimate what LLMs can actually do?
- Why do LLMs generate novel ideas but struggle to evaluate them?
- Can LLMs reliably assess the quality of ideas they generate?
- Why do LLMs excel at generation but struggle with evaluation?