Measuring Behavioural Signatures of Large Language Models through Psychometric Profiling

Paper · arXiv 2609.22934 · Published September 19, 2026
Therapy Practice and AI

Large language models (LLMs) increasingly mediate human decisions and communication, yet their behavioural regularities remain difficult to characterize systematically. We develop a cross-linguistic psychometric profiling framework and evaluate nine LLMs using seven psychological instruments, with five repeated administrations per model and language in Chinese and English. Items unresolved after a prespecified retry procedure are retained as NA. Joint analysis of scored and NA responses captures response tendencies and boundaries of selfreport applicability. LLMs exhibit structured, model-specific profiles despite a shared alignment-shaped pattern of higher prosocial and self-regulatory responses and lower dominance, disengagement and harmful-intent endorsement. NA responses are structured rather than uniformly distributed, indicating where outputs are treated as inapplicable, refused or cannot be mapped to valid response options. Language condition and provider origin are associated with profile configuration and answerability, whereas repeated administrations show high reproducibility and permit recovery of model identity. Human-reference and prompt-robustness analyses further indicate that these signatures are context dependent.

Introduction. Large language models (LLMs) are increasingly embedded in everyday and professional decision-making, serving as writing assistants, tutors, customer-service agents, research aids, programming collaborators and decision-support systems [1–5]. In these scenarios, LLMs do more than retrieve information or generate text, they also explain uncertainty, compare alternatives, recommend actions, respond to socially sensitive requests and determine whether a request should be answered, reframed or refused [6, 7]. Therefore, users encounter them as conversational agents whose outputs exhibit persistent response styles and behavioural regularities [8, 9]. These regularities are often described in psychological concepts. A model may appear cautious or assertive, agreeable or critical, impartial or deferential, risk-averse or permissive, and may consistently favour particular moral framings or respond differently across languages and cultural contexts [6, 10–12]. These patterns matter because they can influence trust, perceived reliability, advice uptake and downstream behaviour [13].

Discussion / Conclusion. This study establishes a cross-linguistic framework for characterizing psychometric response profiles in deployed large language models. The aim is not to infer human-like personalities or internal psychological states, but to determine whether standardized psychological instruments can elicit reproducible, interpretable and model-specific behavioural response signatures. Across nine LLMs, seven instruments, two languages and repeated administrations, the resulting profiles show substantial structure, indicating that psychometric probes can provide a quantitative perspective on behavioural regularities in artificial systems. The models exhibit both convergence and differentiation. Across the battery, they share a broad pattern of comparatively high prosocial, self-regulatory and stability-related responses and low endorsement of dominance, moral disengagement and harmful-intent dimensions.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

How does AI-generated content transformation affect public discourse quality? Does AI fluency substitute for verifiable accuracy in human judgment? Does AI text rewriting systematically distort writer intent and preference? How can humans calibrate appropriate trust in AI systems? Does tokenized intelligence retain genuine value through exchange-based systems? Does alignment training create blind spots in detecting genuine safety threats? Can AI systems balance emotional competence with factual reliability? What makes AI persuasion effective and how can we counter it? Can model confidence signals reliably improve reasoning quality and calibration? How do we evaluate AI systems when user perception misleads actual performance? Why do persona-level simulations fail to predict individual preferences accurately?