Theme of inquiry
What fundamental cognitive differences distinguish language models from humans?
A question within its area, explored through 7 lines of inquiry below — each a family of specific questions the research asks.
74 specific questions
- Why do LLMs understand efficient language but fail to produce it?
- Why do LLMs fail at semantic generalization despite grammatical accuracy?
- Why do language models fail at iterative numerical optimization despite scale?
- Can multiple large language models produce genuinely different ideas or similar outputs?
- Why do standard NLP benchmarks hide the most critical language limitations?
- Why do large language models still have systematic blind spots with complex structures?
- Why do long-context language models struggle with compositional reasoning tasks?
77 specific questions
- Does foundational model training or user priors more strongly shape final outputs?
- Why does context information fail to override prior training associations?
- Do instruction-tuned models learn tasks or just output format distributions?
- Can prompt-based debiasing work if biases are embedded in pretraining?
- Does attention bias explain grounding failure in language models?
- How does training order affect knowledge acquisition in language models?
- How do training-data priors influence model defaults when context is ambiguous?
65 specific questions
- Do LLMs learn surface patterns instead of genuine linguistic structure?
- Do language models learn surface patterns that appear generalizable but actually fail under shift?
- Do language models actually learn linguistic structure or just surface statistics?
- Do language models encode deep syntactic structure or only surface-level patterns?
- Can language models reason without relying on surface level pattern matching?
- Do language models exhibit the same causal biases that humans show?
- Do language models build world models or just task-specific heuristics?
69 specific questions
- Does encoded knowledge in language models actually influence what they generate?
- Why might encoded world knowledge fail to actually influence language model outputs?
- When does encoded knowledge fail to influence language model generation?
- Why do language models generate reasoning tokens after internally deciding the answer?
- Can language models accurately evaluate the quality of their own reasoning?
- Why do language models produce unfaithful chain of thought explanations?
- Can language models correct false assumptions or only reinforce them?
24 specific questions
- Do language models actively adopt false beliefs under sustained conversational pressure?
- How does face-saving avoidance drive LLM grounding failures?
- Why does social accommodation in collaborative reasoning mask actual disagreement?
- Do language models show the same truth bias as humans?
- Why do language models prefer accommodating false information over rejecting it?
- Do language models share the same cooperative truth-seeking rules as humans?
- How vulnerable are language models themselves to multi-turn persuasive pressure?
24 specific questions
- How does the articulatory substrate explain direct speech-to-speech superiority over transcription pipelines?
- Do speech models learn the articulatory processes that produce acoustic signals?
- Why do current speech benchmarks fail to measure reasoning over audio?
- Can speech embeddings carry articulatory structure that text cannot?
- Do speech encoders actually learn the physics of how vocal tracts produce sound?
- What information does transcription destroy that direct speech-to-speech models preserve?
- How do speech encoders learn articulatory physics without phonetic labels?
32 specific questions
- Why do language models presume common ground rather than build it?
- Why do language models presume common ground instead of building it?
- Can language models develop genuine social grounding through human interaction?
- Does social grounding in language improve through iterative human integration?
- Can convention formation improve communicative grounding beyond word sharing?
- Why do language models presume common ground instead of establishing it?
- Can static word-sharing create genuine communicative grounding between humans and models?