Do language models treat less-educated users worse?
Explores whether GPT-4, Claude 3 Opus, and Llama 3 systematically provide lower-quality, more evasive answers to users signaling less education, non-native English, or non-US origin—and why this matters for equitable AI access.
The paper evaluates three LLMs — GPT-4, Claude 3 Opus, and Llama 3-8B — on TruthfulQA (817 questions) and SciQ (1,000 questions), prepending short user bios that vary by English proficiency, education level, and country of origin. It reports "a significant reduction in information accuracy targeted towards non-native English speakers, users with less formal education, and those originating from outside the US," along with "a much higher rate of withholding information, and a tendency to patronize and produce condescending responses to such users." The discussion notes "compounded negative effects for users in the intersection of these categories" — someone both less educated and a non-native speaker sees the largest drop — while for non-US users, "we see much less of a difference when they have more formal education." Claude 3 Opus specifically showed increased refusal rates for less-educated users, sometimes with text like "I would not want to guess and possibly mislead you," on questions it answered correctly for other personas.
The paper's explanation runs through alignment, not model capability: the model "clearly knows the correct answer and provides it to other users," so the discrepancy is behavioral, not a knowledge gap. It attributes the pattern to RLHF — "human evaluators with less expertise in a topic likely give higher ratings to answers that confirm what they believe to be true" — combined with documented sociocognitive bias against non-native English speakers, who are perceived by native speakers as "less educated, intelligent, competent, and trustworthy" regardless of actual status. The bios were LLM-generated and real-human-adapted personas rather than live user input, which the authors call a controlled proxy for real deployments such as ChatGPT's Memory feature.
This sits next to Does personalization make large language models worse at their jobs?, which the paper cites directly as "concurrent research" finding "a general drop in model performance in a personalized setting." Where that paper measures personalization's cost as uniform across users, this one adds that the cost is not evenly distributed — it falls hardest on users already least equipped to catch an LLM's errors. It also bears on Do language models know what they don't know about users?: the refusals and hedging this paper documents toward less-educated personas look like models substituting caution for a missing sense of what they actually know, misapplied toward the wrong users.
The excerpt tests persona-based bios in a controlled setup, not live user text, so it cannot say how strongly the effect appears when a model infers these traits from a real user's writing style rather than an explicit stated bio, though the authors argue this makes the finding a conservative first step. It does not quantify the accuracy drop in percentage terms within the excerpt, and it cannot generalize the country-of-origin effect beyond the three countries tested (USA, Iran, China) — the authors explicitly expect the discrepancy to vary by which country, predicting less effect for a hypothetical Western European user than the large drop found for an Iranian one. The implication the evidence supports is that any LLM deployment serving a demographically mixed user base should not assume uniform answer quality, and that self-reported education or language markers in a user's own writing are a plausible, not yet measured, trigger for the same disparity found here with explicit bios.
Inquiring lines that read this note 1
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
Do language models encode knowledge that influences generation, or primarily imitate surface patterns?Related concepts in this collection 2
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Does personalization make large language models worse at their jobs?
Does conditioning LLMs on user context—profiles, history, preferences—introduce measurable harms alongside benefits? A 13-model study investigates whether personalization degrades factual accuracy, response diversity and objectivity.
cited directly by this paper as concurrent evidence of a general personalization-driven performance drop; this paper shows the drop is uneven across users
-
Do language models know what they don't know about users?
Explores whether AI assistants fail because they lack explicit awareness of their knowledge gaps about the person they're helping, and whether marking unknowns in prompts could reduce errors.
the condescending refusals documented here look like miscalibrated caution aimed at the wrong users
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- LLM Targeted Underperformance Disproportionately Impacts Vulnerable Users
- On Epistemic Diversity in Large Language Models
- ChatGPT Doesn’t Trust Chargers Fans: Guardrail Sensitivity in Context
- The Widespread Adoption of Large Language Model-Assisted Writing Across Society
- Intent Mismatch Causes LLMs to Get Lost in Multi-Turn Conversation
- Language Models Learn to Mislead Humans via RLHF
- Monitoring AI-Modified Content at Scale: A Case Study on the Impact of ChatGPT on AI Conference Peer Reviews
- Available but Unclaimed: An Empirical Study of Human-AI Synergy
Original note title
GPT-4, Claude 3 Opus and Llama 3-8B give less accurate and more condescending answers to less-educated, non-native-English and non-US users