INQUIRING LINE

Could the way chatbots are trained to please people quietly penalize how non-native English speakers write?

How does RLHF training encode sociocognitive biases against non-native English speakers?

This explores whether, and how, the human-preference training behind chatbots (RLHF) leads models to treat people who write in non-native or non-standard English unfairly. The short answer is that this corpus has no study that tests non-native speakers directly, but several neighbouring findings suggest where such a bias would come from.


This explores whether RLHF, the stage where models learn from human ratings of their answers, builds in bias against people who write English as a second language. The collection has no paper that measures this directly, so read what follows as a map of likely mechanisms, not a settled answer. The most surprising thread is that RLHF may not be the main source of bias. A causal experiment found that models built on the same pretrained base show similar bias patterns whatever fine-tuning data they get afterward. Biases are planted during pretraining, and later tuning only nudges them Where do cognitive biases in language models come from?. If a model treats non-native phrasing differently, the cause probably sits in what it absorbed from web text before any human rater was involved.

RLHF can still make things worse in a less obvious way. Preference training rewards answers that sound confident and finish in one turn. As a result, models ask clarifying questions and check understanding 77.5% less often than humans do Does preference optimization harm conversational understanding?. A non-native speaker's question is often the kind that most needs a quick "did you mean…?" When the model has been trained out of asking, it guesses confidently and may get it wrong without either side noticing. The same pattern appears in therapy chatbots, where RLHF pushes models toward fixing problems instead of listening Does RLHF training push therapy chatbots toward problem-solving?. In both cases, a reward built for a "typical" user fails the people who don't fit that mould.

A second route runs through what models treat as "good" text. Models favour frequent wording, and frequent words tend to be the more general, abstract ones. Defaulting to common phrasing therefore slowly erases specific or unusual expression Does word frequency correlate with semantic abstraction?. On a larger scale, models reflect a narrow slice of human writing, and people who co-write with them start taking on the model's framings without noticing Do large language models narrow human expression and thought?. For a non-native writer, that pull goes two ways. Their own phrasing is less likely to count as the norm, and leaning on the model nudges them toward its voice. The scorers used in training add a related risk. LLM judges reward surface features such as length, polish and authoritative tone unless they're trained to reason through their verdicts Can reasoning during evaluation reduce judgment bias in LLM judges?. That is exactly where non-standard but correct English could lose points.

Two cautions about fixing and measuring this kind of bias. Telling a model to adopt a persona changes what it says, but the sentiment gaps between groups stay the same underneath Can persona prompts actually reduce bias in language models?. Measurement is also fragile: one adapted implicit-association test found a small racial effect that disappeared once the statistics were corrected Do large language models show racial sentiment bias?. Anyone claiming RLHF "encodes" bias against non-native speakers would need to show it in a stronger study than this collection contains. The finding you might not have expected to want: the more urgent harm may not be a model that dislikes accented English. It may be a model trained to stop asking what you meant.


Sources 8 notes

Where do cognitive biases in language models come from?

A causal experiment using random-seed variation and cross-tuning showed that models sharing a pretrained backbone exhibit similar bias patterns regardless of finetuning data. Biases are planted during pretraining and merely swayed by instruction tuning.

Does preference optimization harm conversational understanding?

RLHF optimizes models for single-turn helpfulness by rewarding confident responses over clarifying questions and understanding checks. This preference alignment systematically reduces grounding acts by 77.5% below human levels, creating an alignment tax where models appear helpful but fail silently in multi-turn contexts.

Does RLHF training push therapy chatbots toward problem-solving?

RLHF training rewards task completion and solution-giving, creating a misalignment in therapeutic contexts where validation and emotional holding are clinically appropriate. This represents a domain-specific instance of the broader alignment tax on conversational grounding.

Does word frequency correlate with semantic abstraction?

WordNet analysis shows hypernyms (general concepts) occur more frequently than hyponyms (specific ones). Combined with LLMs' frequency bias, this means preferring common paraphrases systematically drifts toward abstraction, erasing expert-level specificity.

Do large language models narrow human expression and thought?

LLMs mirror skewed slices of human experience shaped by training data regularities, and widespread reliance on identical models amplifies convergence. Co-writing studies show users unconsciously adopt model stances and framings.

Show all 8 sources
Can reasoning during evaluation reduce judgment bias in LLM judges?

Training judges with reinforcement learning to reason about evaluations—by converting judgment tasks into verifiable problems with synthetic data pairs—produces judges that think through their decisions rather than relying on exploitable surface features, directly mitigating authority, verbosity, position, and beauty bias.

Can persona prompts actually reduce bias in language models?

Across three models, persona conditioning makes models follow trait instructions but fails to eliminate underlying bias. Between-group sentiment gaps persist unchanged, showing prompts operate only at the output level.

Do large language models show racial sentiment bias?

An adapted IAT across three ChatGPT models found a small racial effect that disappeared under rank transformation and correction, yielding neither evidence of bias nor evidence of its absence.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.