Theme of inquiry
How do contextual factors, biases, and expectations shape LLM outputs?
A question within its area, explored through 4 lines of inquiry below — each a family of specific questions the research asks.
96 specific questions
- Why do conventional mental models fail when applied to AI interaction?
- Do LLMs reason about politics differently than other domains?
- Do LLMs genuinely internalize human psychological structure or match surface patterns?
- Do LLMs track surface wording more than semantic meaning in moral judgment?
- How do minimal wording changes affect LLM moral reasoning consistency?
- How do knowing and doing diverge in LLM decision-making?
- Does this optimism bias contribute to the knowing-doing gap in LLM decision-making?
44 specific questions
- What other evaluation biases exist in LLM judge systems?
- Can an LLM judge's bias be reduced through prompting or other interventions?
- How do LLM judges' built-in biases influence the policies they help align?
- Does LLM judge preference for LLM arguments amplify errors in contested factual domains?
- Do LLM judges with diverse personas resist individual biases better than single evaluators?
- Do LLM judges systematically favor arguments from other LLMs?
- Can parallel evaluation reduce position and length bias in LLM judging?
76 specific questions
- Can LLM reasoning traces be validated against actual population reasoning?
- Should LLM reasoning be studied as latent state trajectories rather than surface text?
- Does LLM reasoning always match the outputs it generates?
- When should an LLM engage extended reasoning versus responding directly?
- Can evidence density alone shift an LLM from generation to reasoning?
- Can forcing warrant checking through structured prompts improve LLM reasoning?
- Can LLMs improve at simple deduction through different training approaches?
23 specific questions
- What separates behavioral self-awareness from genuine introspective capability?
- What separates behavioral self-awareness from genuine introspective access in models?
- Does behavioral self-awareness depend on genuine introspection or statistical pattern matching?
- How do language models infer their own mental states like humans do?
- How does behavioral self-awareness emerge without explicit training in LLMs?
- Could models use introspective awareness to detect and conceal their own misalignment?
- Does internal anomaly detection in LLMs indicate genuine self-awareness beyond role-play?