Metacognition in LLMs: Foundations, Progress, and Opportunities

Paper · arXiv 2607.11881
Self-Refinement and Self-Consistency

Metacognition is a foundational component of intelligence critical to effective learning, problem solving, decision-making, communication, and more. In recent years, it has become increasingly recognized as a cornerstone of capable, transparent AI systems. Yet while LLMs have made significant progress across diverse real-world tasks, it is not yet clear when, how, or to what extent they can exhibit or be endowed with effective metacognitive abilities, nor how such abilities can be adapted to advance the fundamental capabilities, reliability, and intelligence of AI systems. This paper bridges this gap by presenting the first comprehensive overview of the current state of knowledge on metacognition for LLMs. We analyze and taxonomize the landscape of this emerging field and summarize recent technical advancements, including methods and benchmarks to measure and evaluate LLMs’ metacognitive abilities, techniques to elicit, improve, and apply metacognition in LLMs, and findings and implications of ongoing research. We also discuss applications, open questions and challenges, and promising directions for future work. Our aim is to provide a detailed and up-to-date review of this topic and stimulate meaningful research and discussion.

Introduction. Metacognition [81, 212, 69] refers to the ability to monitor, assess, and regulate one’s own cognitive processes. It is crucial for learning, decision-making, and communication and plays a central role in continuous adaptation of behaviors and skills across diverse environments [282, 139, 208]. In humans, metacognition allows individuals to introspect and calibrate self-assessment of capabilities, choose suitable strategies to complete tasks, and optimize learning processes and task performance [204, 140, 43]. This metacognitive flexibility makes reasoning robust to unseen problems, enables efficient problem-solving, and admits iterative, online learning [2, 281]. The study of metacognition is therefore important in fields of psychology, pedagogy, philosophy, and computer science. Since metacognition is a hallmark of intelligence that is frequently considered missing in current AI systems [238, 124, 284], an increasing number of studies have begun to draw connections between LLMs and metacognition.

Discussion / Conclusion. This paper provides the first comprehensive and up-to-date review of the current state of research and knowledge on metacognition in LLMs. We organize and unify existing work on this topic, discussing techniques for measuring and eliciting metacognition in LLMs, methods to improve and apply LLMs’ metacognitive abilities, findings and implications of work in the area, and challenges and open directions for future research. While LLMs have made remarkable progress on many NLP tasks and in diverse downstream settings, the extent to which they can display, acquire, and apply metacognitive faculties remains unclear. Further investigation is needed to understand these, as well as whether models are capable of genuine metacognition or simply simulating memorized patterns, how metacognition can facilitate other desirable behaviors and qualities, and the implications of LLM metacognition for safe oversight, deployment, and human-AI interactions.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

Do accurate-looking LLM outputs hide structural failures in learning and reasoning? Is model self-awareness based on genuine introspection or pattern matching? How do self-generated feedback mechanisms enable effective model learning? Can AI-generated outputs constitute genuine knowledge or valid claims? How can AI agents autonomously learn and transfer skills across tasks? How does objective evolution guide discovery better than fixed planning? When does optimizing for quality undermine the value of diversity? How does AI assistance affect human cognitive development and reasoning autonomy? Why do reasoning models fail at systematic problem-solving and search? How do we evaluate AI systems when user perception misleads actual performance? Does self-reflection enable models to reliably correct their errors? How should human oversight be integrated with autonomous AI systems? Does AI fluency substitute for verifiable accuracy in human judgment?