Line of inquiry
Inquiring lines›What enables humans to maintain au…›What makes preference signals vali…›this line of inquiry
Does debiasing during pretraining prevent biases in downstream model behavior?
A broader line of inquiry — a family of 19 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 19
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Can data filtering during pretraining prevent cognitive biases in language models?
- Can prompt-based debiasing work if biases are embedded in pretraining?
- Why does eliminating proxy-model filtering improve reasoning emergence in pretraining?
- Why does diversity in training data enable denoising rather than reinforce shared biases?
- Does removing cognitive bias from training signals accidentally break what makes alignment work?
- Do pretraining biases and traditional selection bias compound in production recommenders?
- Can dataset-level debiasing methods fix popularity bias inherited from pretraining?
- Does debiasing training data actually solve the bias problem in machine learning?
- Can self-training drift be prevented by applying student compatibility filtering?
- How much can mitigation techniques like augmentation reduce priming without harming learning?
- Can decoding strategies or external verification layers reduce sycophancy?
- Does filtering passages before generation improve large model answer quality?
- Can a rejected-edit buffer work like hard negatives in contrastive learning?
- How does Western-dominance bias propagate through multimodal training data?
- Why can data filtering fail to remove transmitted behavioral traits?
- Can selective history filtering address topic drift that generation-time topic following cannot prevent?
- Can humans suppress frequency bias through attention and intention?
- Why are documents read but not cited harder distractors than random samples?
- Can rejected edits serve as negative feedback like hard negatives in contrastive learning?