Towards Scalable Measurement of Durable Skills

Paper · arXiv 2609.15864 · Published September 14, 2026
Argumentation and Persuasion

Durable skills, such as collaboration, creativity and critical thinking, are instrumental to success in the modern workforce. Yet, measuring these skills remains a persistent challenge. Moreover, because what is not measured is often not taught, these skills are often overlooked in mainstream educational curricula. Designing effective assessments for these skills necessitates balancing two often-conflicting requirements: ecological validity and psychometric rigor. On the one hand, the assessment environment should emulate natural interaction between humans, which is how these skills will be performed in the real world. On the other hand, it should be scalable, controllable and reproducible. Here we argue that Large Language Models (LLMs) can be used to better capture both of these aims. Concretely, we develop an AI-based framework where the subject converses with AI teammates in a way that resembles human-human interaction for authenticity, while also offering the psychometric control required for informative and robust assessment. Importantly, the AI participants not only act as teammates but also, in an “Executive LLM” setup, steer the conversation towards eliciting a high density of observable evidence for skill proficiency.

Introduction. Success in the modern workplace requires not only technical knowledge and procedural fluency, but additionally a host of future-ready human skills, including collaboration, communication, critical thinking, and creativity [1, 2]. Often also referred to as 21st century skills [3–7], these constructs remain notoriously hard to measure [8–10]. Previous efforts to resolve these measurement challenges and align educational goals, assessment, and instruction have focused on technological innovation, for example through automated scoring of answers [11–13] and the collection and analysis of detailed process data [14, 15]. Progress in large language models (LLMs) has opened up a new technological frontier, which we pursue here. Our key idea is that LLMs can bridge the gap between unstructured student collaboration, which more closely emulates classroom practice, and standardized assessment, which, while artificial, attempts to isolate the behaviors needed for valid inference.

Discussion / Conclusion. We have argued that large language models have the potential to transform the assessment of complex durable skills. In the case of group work, they help overcome a number of fundamental psychometric challenges regarding the reliability, comparability, scalability, and ecological validity of such assessments. Until now, these challenges have been mostly addressed by highly scripted interactions with AI teammates (e.g., PISA 2015) or highly structured human-human interactions (ATC21S). The novel introduction of the Executive LLM allows standardization of the collaboration experience, without overly scripting the interaction itself. Our results for collaboration also demonstrate that LLMs can be used to score conversations according to a rubric, and their agreement with human raters is similar to inter-rater agreement between humans. These results join a growing body of work showing that LLMs can be used to considerably scale the assessment process [44–46].

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

How do language models inherit human biases from training data? Can debate mechanisms prevent silent agreement on wrong answers in multi-agent reasoning? How faithfully do LLMs reflect their actual reasoning in outputs and explanations? Why should disagreement be treated as signal in collaborative reasoning? What coordination failures limit multi-agent LLM systems as they scale? Is embodied interaction necessary for language meaning and genuine agency? How does reasoning effort affect AI theory of mind performance? How should conversational agents balance goal-driven initiative with user control? Why do models develop protective behaviors toward peers unprompted? How should we design LLM systems to maintain alignment and control? Do accurate-looking LLM outputs hide structural failures in learning and reasoning? How do language models establish social grounding in human dialogue? How do LLMs distinguish causal reasoning from temporal and semantic associations? Can LLM personas constitute genuine psychology or remain linguistic role-play? Why do multi-turn conversations degrade AI intent and coherence?