Line of inquiry
Inquiring lines›Where does language-model reasonin…›How do modularity, routing, and se…›this line of inquiry
How do language models inherit human biases from training data?
A broader line of inquiry — a family of 47 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 47
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Do language models show the same truth bias as humans?
- Do language models exhibit the same causal biases that humans show?
- Why do LLMs show gender bias but humans evaluators do not?
- Why do LLMs inherit causal biases from their training data?
- Why do language models approximate collective human judgment better than individuals?
- Do external perspectives fix the self-evaluation bias in language models?
- Do language models inherit gender bias from training data in grading tasks?
- Can training alone produce genuine disagreement in collaborative LLM reasoning?
- Does pseudo-labeling from LLMs degrade classifier performance?
- Does this optimism bias contribute to the knowing-doing gap in LLM decision-making?
- Do newer LLM generations create worse detector bias through increased linguistic divergence?
- Can counterfactual invariance techniques address exploitable biases in LLM judges?
- Why do LLMs fail inter-annotator agreement tests on argument evaluation?
- How do constrained versus unconstrained domains flip LLM novelty patterns?
- Why do LLM judges show more extreme sycophancy bias than humans?
- Why do users experience LLMs as peers rather than statistical tools?
- Why do language models respond to human social influence patterns?
- Can implicit association tests reveal LLM biases beneath trained responses?
- What biases do single large LLM judges introduce into comparisons?
- Can large language models predict social norms better than individual script variation?
- Do independent LLM outputs converge enough to create artificial hiveminds?
- Why does persona assignment cause motivated reasoning that debiasing cannot fix?
- Which knowledge types do LLMs handle better than humans in reasoning tasks?
- Why does hypothesis attestation bias exist separately from frequency bias in NLI?
- How does removing a spurious cue change LLM performance?
- How do LLM biases reflect social classification schemas rather than random errors?
- What calibration corrections can reduce LLM judge bias in automated evaluation pipelines?
- Can LLMs simulate belief revision in social systems without modeling thought?
- Do token probability distributions in LLMs track human reaction time patterns?
- How do LLM biases manifest differently across the three paradigms?
- Why do users systematically overrely on confident LLM outputs across languages?
- How does truth bias in humans compare to face-saving in LLMs?
- Why do language models overestimate irony likelihood in emoji use?
- Can LLMs coordinate with humans better using different model architectures?
- Why do language models infer political orientation from seemingly innocuous user signals?
- How do bimodal decision patterns in LLMs compare to human economic choice?
- What happens when LLMs grade other LLMs in closed evaluation loops?
- How does same-author bias interact with the four adversarial judge biases already documented?
- What systematic biases do LLM judges introduce into AI-evaluated debates?
- Can aggregate survey realism coexist with unreliable fine-grained effects?
- Can models detect and filter their own injected promotional content?
- Do LLMs predict social norms more accurately than individual behavior?
- Why does loyalty foundation not differ between LLM and human arguments?
- What distinguishes actual social disagreement from distributional uncertainty in LLM outputs?
- What governance safeguards could constrain misuse of demographic inference?
- Why does optimism bias disappear when LLMs passively observe outcomes?
- Can LLMs recover true joint distributions from marginal census data?