Theme of inquiry
Why are LLM outputs so inconsistent across minor contextual variations?
A question within its area, explored through 4 lines of inquiry below — each a family of specific questions the research asks.
79 specific questions
- Do LLMs genuinely internalize human psychological structure or match surface patterns?
- How do language models infer their own mental states like humans do?
- Do LLMs rely on surface statistical patterns instead of causal structure?
- Why do conventional mental models fail when applied to AI interaction?
- Does internal anomaly detection in LLMs indicate genuine self-awareness beyond role-play?
- Do realistic LLM behaviors require simulating human thought or just behavior?
- Can LLMs participate meaningfully in discourse without consciousness or understanding?
45 specific questions
- Do LLMs actually reason differently than humans about moral dilemmas?
- How do moral language patterns differ between LLM and human arguments?
- Do LLMs reason about politics differently than other domains?
- How do minimal wording changes affect LLM moral reasoning consistency?
- Why do LLMs persuade through logical appeals but humans through emotion?
- Do LLMs achieve similar persuasive outcomes through different rhetorical mechanisms than humans?
- How do knowing and doing diverge in LLM decision-making?
37 specific questions
- What makes novelty assessment harder to automate than idea generation?
- Can LLMs generate more novel research ideas than human experts?
- Why do research ideation systems suffer from diversity collapse despite high novelty metrics?
- Why do LLMs generate novel ideas but struggle to evaluate them?
- Why do LLM-generated ideas score higher novelty yet lower feasibility than expert ideas?
- Can LLMs reliably assess the quality of ideas they generate?
- Why do LLM research ideas lack diversity despite high average novelty?
43 specific questions
- What other evaluation biases exist in LLM judge systems?
- Can an LLM judge's bias be reduced through prompting or other interventions?
- How do LLM judges' built-in biases influence the policies they help align?
- Do LLM judges with diverse personas resist individual biases better than single evaluators?
- What calibration corrections can reduce LLM judge bias in automated evaluation pipelines?
- Does ensembling smaller judges reduce bias more effectively than single large judges?
- Why do LLM judges systematically favor outputs from their own model family?