From Minds to Models: The Intersection of Psychology and LLM Behaviours

Paper · arXiv 2607.27579 · Published July 30, 2026
Emotions and AI

The opacity of large language models' (LLMs') decision-making is often compared with the complexity, non-linearity and interpretive difficulty of the human mind. Given these parallels, psychological research methods developed to probe unobservable mental processes may be adaptable to the study of LLM behaviour. This is particularly important where LLMs are deployed in government and healthcare, where transparency and accountability are essential. Building on prompt-based adaptations of the Implicit Association Test, this study examined whether ChatGPT produced relative differences in sentiment across racial conditions in open-ended text. We developed a generation-based LLM-adapted Implicit Association Test comprising 14 base questions crossed with eight racial categories and a race-agnostic control. Each of the 126 test questions was submitted once to GPT-3.5T, GPT-4 and GPT-4T, yielding 378 responses. The analysed sentiment score was derived from the workbook's categorical sentiment label and source score: positive labels retained the source score, negative labels were assigned the negative of that score, and neutral responses were coded as zero.

Introduction. A Large Language Model (LLM) is a branch of artificial intelligence (AI) designed to learn and understand human language, creating context that enables effective interaction with individuals. By comprehending the context of language, LLMs can analyse content, make decisions and provide relevant and informative responses based on their training data. Models such as GPT (Generative Pre-trained Transformer) and BERT (Bidirectional Encoder Representations from Transformers) are widely used in fields including healthcare, government and content generation. The opacity of AI decision-making processes, particularly in critical fields like government and medical diagnostics, raises ethical concerns due to the reliance on outputs that cannot be fully explained (Hamirul et al., 2023). This opacity, often described as the AI “black box”, mirrors the challenges faced in psychology when attempting to understand the human mind.

Discussion / Conclusion. Evaluation of Hypotheses H1 received limited and analysis-dependent support. The parametric ANOVA detected a small racial-condition effect, but the effect was not retained after rank transformation and no Tukey-corrected pairwise comparison was significant. H2 predicted a consistent pattern across models. The absence of both a model main effect and a racial condition × model interaction is consistent with that prediction, although a non-significant interaction does not establish equivalence across models. The inferential pattern constrains what can be concluded. The parametric omnibus result is close to the conventional threshold, disappears under rank transformation, and does not resolve into any significant Tukey comparison. The exploratory European–Indigenous Australian contrast is numerically sizeable, but it was selected post hoc, was uncorrected, and does not compare the two extreme means in the reconciled dataset. It cannot establish that either condition is evaluated differently.

Lines of inquiry this paper opens 14

Research framings built by reading the notes related to this paper — the questions it feeds into.

How do chatbots affect human self-disclosure and emotional engagement? How faithfully do LLMs reflect their actual reasoning in outputs and explanations? How do language models inherit human biases from training data? Does RLHF training sacrifice accuracy and grounding for user agreement? Does alignment training create blind spots in detecting genuine safety threats? How should we design LLM systems to maintain alignment and control? How do professional roles and expertise transform with AI-generated content? How does AI assistance affect human cognitive development and reasoning autonomy? Does conversational format create illusions of genuine AI communication? Can AI-generated outputs constitute genuine knowledge or valid claims? Does tokenized intelligence retain genuine value through exchange-based systems?