Natural Language Processing Psychometrics

Paper · arXiv 2608.07316 · Published August 7, 2026
Therapy Practice and AI

Natural Language Processing (NLP) models predicting mental health outcomes rarely specify what they measure: contextual knowledge, emotional content, or syntactic structure. NLP Psychometrics treats psychological prediction from text as a psychometric problem, linking scores to interpretable linguistic evidence and testing beyond the training text format. Nine LLMs, conditioned on controlled personas (cognitive digital shadows), completed psychometric questionnaires with textual explanations per item. We extracted emotional profiles and syntactic-semantic structure via textual forma mentis networks, combined with personality and sociodemographic variables in ablated random forest (RF) regressors, using SHAP to identify which features drove performance and in which direction. Full RF models explained up to 70.8% of variance in life satisfaction (SWLS), 55.7% in depression (PHQ-9), and, for DASS-21, 68.5% depression, 76.0% anxiety, 72.4% stress. Sociodemographics alone explained no meaningful variance in depression, anxiety, or stress, but did so for life satisfaction, where emotion features and income were the strongest predictors; neuroticism and network topology instead dominated depression and anxiety, reversing direction between them.

Introduction. Measuring psychological phenomena depends on inner experience leaving quantifiable traces behind [1–3]. Some of those are explicit, as in questionnaire ratings: Individuals consciously rate experiences and leave correlated numerical sequences, or item scores, as traces of their inner world [1, 3]. Other traces are linguistic [4–6], distributed across the words people choose, the emotions they express, and the concepts they connect, either explicitly or implicitly. In this view, language is not only a vehicle for communication [7, 8] but a key trace for psychological measurement, whose structured knowledge can open the way to psychological measurements [9, 10]. This premise is deeply connected to research on the mental lexicon [8, 11, 12]. In the modelling metaphor of the mental lexicon, language is reflected within cognition as a complex system of interconnected concepts or words [13]. The latter are not independent labels attached to experience, but structured cognitive representations embedded in networks of meaning [8].

Discussion / Conclusion. From our pioneering work with NLP Psychometrics, three findings stand out. First, language carried most of the recoverable psychometric signal: emotion and network features formed a predictive core of machine learning features across all tested scales (SWLS, PHQ-9 and DASS-21), whereas sociodemographics alone rarely explained meaningful variance. Second, the markers identified by SHAP scores were interpretable and construct-specific, ranging from family income and affect for life satisfaction to neuroticism and discourse topology for depression. Third, the machine learning "feature to psychometric score" mapping transferred, with reduced yet significant accuracy, to both out-of-genre LLM-generated diaries and to human clinically labelled data [75, 77], i.e., speech transcripts of clinically depressed patients and controls.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

What role does compression play in language model capability and generalization? What prevents language models from reliably adopting diverse personas? Why do persona-level simulations fail to predict individual preferences accurately? How can persona representations reduce language model variance and improve task accuracy? Why can't humans reliably detect AI-generated text despite measurable linguistic signatures? What factors beyond surface content determine how readers extract meaning differently? How does latent reasoning compare to verbalized chain-of-thought? How can emotions function as reliable information in reasoning and cognitive systems? Can model confidence signals reliably improve reasoning quality and calibration? Why do language models struggle with implicit discourse relations? How can real-time alliance measurement improve therapy outcomes? Does conversational format create illusions of genuine AI communication?