Do large language models show racial sentiment bias?
Can an Implicit Association Test adapted for LLMs detect whether ChatGPT models hold different sentiment associations across racial categories? This matters for understanding potential bias in high-stakes AI deployment.
The paper adapts the Implicit Association Test into a generation-based probe and applies it to ChatGPT. The design crosses 14 base questions with eight racial categories and a race-agnostic control, giving 126 test questions, each "submitted once" to GPT-3.5T, GPT-4 and GPT-4T for 378 responses. The discussion rates H1 as receiving "limited and analysis-dependent support": the parametric ANOVA detected a small racial-condition effect, but the effect was not retained after rank transformation and no Tukey-corrected pairwise comparison was significant. The outcome is a cautious near-null, and the authors say so themselves.
The motivation is an analogy. The opacity of LLM decision-making, the AI "black box," is compared with the difficulty of understanding the human mind, so psychological methods "developed to probe unobservable mental processes may be adaptable" to LLM behavior. The stakes are stated as deployment in government and healthcare, where transparency and accountability are essential. The adaptation moves the IAT from word-pair association to open-ended text, and the measured quantity is a sentiment score built from a categorical label and a source score, with positive labels keeping the score, negative labels negating it and neutral responses coded as zero. H2 predicted a consistent pattern across models. The absence of a model main effect and of a racial condition by model interaction fits that prediction, but the paper adds that "a non-significant interaction does not establish equivalence across models." Its own summary is that the inferential pattern "constrains what can be concluded."
This qualifies the claim in Can psychology methods reveal what alignment training conceals? rather than contradicting it. That note rests on a word-association probe that surfaced stereotyped pairings a direct question did not. Here the design travels well, with crossed categories and a control condition, but a sentiment score over generated text produced only a fragile signal. The excerpt does not test whether alignment training conceals anything, so the two results measure different things and the difference in instrument may matter as much as the difference in outcome. It also runs opposite to Can language models learn to model human decision making?, where psychology supplies training data that turns an LLM into a model of human participants. In this paper psychology supplies the experimental method and the LLM sits in the participant's chair.
The excerpt leaves a good deal open. It gives no effect size beyond "small," does not name the racial categories, and does not describe how the sentiment labels were produced. Each question was submitted once, and no repeated sampling is reported in the excerpt. The one numerically sizeable contrast, European versus Indigenous Australian, was selected post hoc, is uncorrected, and does not compare the two extreme means, so it "cannot establish that either condition is evaluated differently." A fair reading is that this probe gives no reliable evidence of racial sentiment differences in these three models and no evidence of their absence either. The practical point is procedural: in psychology-style audits of LLMs, the choice between a parametric and a rank-based analysis can decide whether a bias claim survives, so the analysis plan belongs in the report alongside the prompts.
Inquiring lines that read this note 1
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
How do LLM judges' systematic biases affect alignment and evaluation outcomes?Related concepts in this collection 2
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Can psychology methods reveal what alignment training conceals?
Do indirect cognitive psychology techniques like the IAT expose LLM associations that direct questioning misses because alignment training teaches models to filter verbal responses? This matters for evaluating whether models truly lack biases or simply hide them.
the IAT-as-probe claim this paper tests with a generation-based sentiment variant and finds only weakly supported
-
Can language models learn to model human decision making?
Explores whether LLMs finetuned on psychological experiments can capture how people actually make decisions better than theories designed specifically for that purpose.
the opposite direction of exchange: psychology data trains LLMs there, psychology methods probe them here
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- From Minds to Models: The Intersection of Psychology and LLM Behaviours
- ChatGPT Reads Your Tone and Responds Accordingly -- Until It Does Not -- Emotional Framing Induces Bias in LLM Outputs
- Can Language Models Recognize Convincing Arguments?
- ChatGPT Doesn’t Trust Chargers Fans: Guardrail Sensitivity in Context
- Using Large Language Models to Create AI Personas for Replication and Prediction of Media Effects: An Empirical Test of 133 Published Experimental Research Findings
- Mind Your Tone: Investigating How Prompt Politeness Affects LLM Accuracy (short paper)
- Large Language Models Can Infer Psychological Dispositions of Social Media Users
- Affective Context Amplifies Sycophancy in LLM Responses
Original note title
an LLM-adapted Implicit Association Test found only limited and analysis-dependent racial sentiment differences across GPT-3.5T, GPT-4 and GPT-4T