Theme of inquiry
How do modularity, routing, and search affect architecture efficiency?
A question within its area, explored through 5 lines of inquiry below — each a family of specific questions the research asks.
26 specific questions
- Can users experience the LLM Fallacy even when AI outputs are completely accurate?
- How does LLM hallucination risk manifest in knowledge graph construction?
- Does framing LLM output as fabrication rather than hallucination matter philosophically?
- Can surface-level correctness hide failures in structural learning by LLMs?
- Why is hallucination the wrong term for all LLM false outputs?
- What should we call errors in LLM outputs when hallucination does not apply?
- What happens when we treat LLM outputs as sampled rather than stored?
34 specific questions
- Why do standard NLP benchmarks hide the most critical language limitations?
- Do standard language benchmarks underestimate what LLMs can actually do?
- Why do language models fail at iterative numerical optimization despite scale?
- Why do NLP benchmarks hide LLM failures in ambiguity handling?
- Why do language models plateau at constraint satisfaction regardless of scale?
- Why do large language models still have systematic blind spots with complex structures?
- What prevents monolithic LLMs from coordinating decomposition with execution?
40 specific questions
- Can language models perform purely symbolic reasoning when semantics are removed?
- Can language models perform genuine symbolic reasoning without semantic grounding?
- Can LLMs translate between natural language and formal logic faithfully?
- Why does augmenting symbolic reasoning outperform replacing it entirely?
- Can language models reason without relying on learned semantic patterns?
- Why do LLMs struggle to translate natural language into logical formalizations?
- Why do LLMs fail at faithful autoformalisation of reasoning problems?
17 specific questions
- How do different LLM integration paradigms affect inheritance of pretraining biases?
- How can human-centered objectives be embedded earlier in the LLM pipeline?
- What deployment feedback loops amplify LLM pretraining popularity in live systems?
- What interaction controls matter most for effective human-LLM collaboration?
- How does content-only knowledge in LLMs enable pretraining popularity to leak through?
- What unique perspective do designers bring to LLM adaptation that engineers might miss?
- How does this differ from using LLMs as the policy itself?
47 specific questions
- Do language models show the same truth bias as humans?
- Do language models exhibit the same causal biases that humans show?
- Why do LLMs show gender bias but humans evaluators do not?
- Why do LLMs inherit causal biases from their training data?
- Why do language models approximate collective human judgment better than individuals?
- Do external perspectives fix the self-evaluation bias in language models?
- Do language models inherit gender bias from training data in grading tasks?