Theme of inquiry
Why do linguistic mismatches cause LLMs to fail despite apparent understanding?
A question within its area, explored through 3 lines of inquiry below — each a family of specific questions the research asks.
80 specific questions
- How faithful are natural language explanations from LLMs really?
- Does LLM reasoning always match the outputs it generates?
- Why does LLM knowledge fail to influence their actual outputs?
- Why do LLM outputs match researcher priors without solving tasks correctly?
- Can evidence density alone shift an LLM from generation to reasoning?
- When should an LLM engage extended reasoning versus responding directly?
- Can forcing warrant checking through structured prompts improve LLM reasoning?
79 specific questions
- What prevents monolithic LLMs from coordinating decomposition with execution?
- Can LLMs reliably generate novel working architectures without structured representations?
- Why do LLMs fail at directly solving stochastic control problems?
- What planning tasks benefit most from combining LLM generation with external verification?
- Can LLMs successfully translate natural language into formal solver specifications?
- What distinguishes LLM Programs from chain-of-thought and agentic frameworks?
- Do LLMs fail exploration because of context integration or computational limitations?
61 specific questions
- Can LLMs translate between natural language and formal logic faithfully?
- Why do LLMs fail at semantic generalization despite grammatical accuracy?
- Why do LLMs struggle to translate natural language into logical formalizations?
- Can language models translate theorems faithfully without semantic loss?
- Why do LLMs fail at faithful autoformalisation of reasoning problems?
- Do LLMs struggle more with semantic accuracy than syntactic correctness across domains?
- Can language models perform genuine symbolic reasoning without semantic grounding?