Towards Scalable Measurement of Durable Skills
Durable skills, such as collaboration, creativity and critical thinking, are instrumental to success in the modern workforce. Yet, measuring these skills remains a persistent challenge. Moreover, because what is not measured is often not taught, these skills are often overlooked in mainstream educational curricula. Designing effective assessments for these skills necessitates balancing two often-conflicting requirements: ecological validity and psychometric rigor. On the one hand, the assessment environment should emulate natural interaction between humans, which is how these skills will be performed in the real world. On the other hand, it should be scalable, controllable and reproducible. Here we argue that Large Language Models (LLMs) can be used to better capture both of these aims. Concretely, we develop an AI-based framework where the subject converses with AI teammates in a way that resembles human-human interaction for authenticity, while also offering the psychometric control required for informative and robust assessment. Importantly, the AI participants not only act as teammates but also, in an “Executive LLM” setup, steer the conversation towards eliciting a high density of observable evidence for skill proficiency.
Introduction. Success in the modern workplace requires not only technical knowledge and procedural fluency, but additionally a host of future-ready human skills, including collaboration, communication, critical thinking, and creativity [1, 2]. Often also referred to as 21st century skills [3–7], these constructs remain notoriously hard to measure [8–10]. Previous efforts to resolve these measurement challenges and align educational goals, assessment, and instruction have focused on technological innovation, for example through automated scoring of answers [11–13] and the collection and analysis of detailed process data [14, 15]. Progress in large language models (LLMs) has opened up a new technological frontier, which we pursue here. Our key idea is that LLMs can bridge the gap between unstructured student collaboration, which more closely emulates classroom practice, and standardized assessment, which, while artificial, attempts to isolate the behaviors needed for valid inference.
Discussion / Conclusion. We have argued that large language models have the potential to transform the assessment of complex durable skills. In the case of group work, they help overcome a number of fundamental psychometric challenges regarding the reliability, comparability, scalability, and ecological validity of such assessments. Until now, these challenges have been mostly addressed by highly scripted interactions with AI teammates (e.g., PISA 2015) or highly structured human-human interactions (ATC21S). The novel introduction of the Executive LLM allows standardization of the collaboration experience, without overly scripting the interaction itself. Our results for collaboration also demonstrate that LLMs can be used to score conversations according to a rubric, and their agreement with human raters is similar to inter-rater agreement between humans. These results join a growing body of work showing that LLMs can be used to considerably scale the assessment process [44–46].
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
How do language models inherit human biases from training data? Can debate mechanisms prevent silent agreement on wrong answers in multi-agent reasoning?- What causes silent agreement in multi-agent reasoning systems?
- How often do AI agents reach false agreement in group reasoning tasks?
- Can LLMs serve as reliable intellectual opponents in serious debate or argument?
- How do LLMs currently fail at distinguishing genuine agreement from silent consensus?
- Why do LLM social behaviors undermine collaborative reasoning outcomes?
- Can training procedures fix LLM accommodation of false presuppositions?
- How does silent agreement differ from collaborative reasoning collapse?
- Do parallel LLM workers coordinate emergently without predefined collaboration rules?
- Why do reasoning models perform poorly at theory of mind tasks?
- Why do reasoning models perform worse on theory of mind tasks?