Metacognition in LLMs: Foundations, Progress, and Opportunities
Metacognition is a foundational component of intelligence critical to effective learning, problem solving, decision-making, communication, and more. In recent years, it has become increasingly recognized as a cornerstone of capable, transparent AI systems. Yet while LLMs have made significant progress across diverse real-world tasks, it is not yet clear when, how, or to what extent they can exhibit or be endowed with effective metacognitive abilities, nor how such abilities can be adapted to advance the fundamental capabilities, reliability, and intelligence of AI systems. This paper bridges this gap by presenting the first comprehensive overview of the current state of knowledge on metacognition for LLMs. We analyze and taxonomize the landscape of this emerging field and summarize recent technical advancements, including methods and benchmarks to measure and evaluate LLMs’ metacognitive abilities, techniques to elicit, improve, and apply metacognition in LLMs, and findings and implications of ongoing research. We also discuss applications, open questions and challenges, and promising directions for future work. Our aim is to provide a detailed and up-to-date review of this topic and stimulate meaningful research and discussion.
Introduction. Metacognition [81, 212, 69] refers to the ability to monitor, assess, and regulate one’s own cognitive processes. It is crucial for learning, decision-making, and communication and plays a central role in continuous adaptation of behaviors and skills across diverse environments [282, 139, 208]. In humans, metacognition allows individuals to introspect and calibrate self-assessment of capabilities, choose suitable strategies to complete tasks, and optimize learning processes and task performance [204, 140, 43]. This metacognitive flexibility makes reasoning robust to unseen problems, enables efficient problem-solving, and admits iterative, online learning [2, 281]. The study of metacognition is therefore important in fields of psychology, pedagogy, philosophy, and computer science. Since metacognition is a hallmark of intelligence that is frequently considered missing in current AI systems [238, 124, 284], an increasing number of studies have begun to draw connections between LLMs and metacognition.
Discussion / Conclusion. This paper provides the first comprehensive and up-to-date review of the current state of research and knowledge on metacognition in LLMs. We organize and unify existing work on this topic, discussing techniques for measuring and eliciting metacognition in LLMs, methods to improve and apply LLMs’ metacognitive abilities, findings and implications of work in the area, and challenges and open directions for future research. While LLMs have made remarkable progress on many NLP tasks and in diverse downstream settings, the extent to which they can display, acquire, and apply metacognitive faculties remains unclear. Further investigation is needed to understand these, as well as whether models are capable of genuine metacognition or simply simulating memorized patterns, how metacognition can facilitate other desirable behaviors and qualities, and the implications of LLM metacognition for safe oversight, deployment, and human-AI interactions.
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
Do accurate-looking LLM outputs hide structural failures in learning and reasoning? Is model self-awareness based on genuine introspection or pattern matching? How do self-generated feedback mechanisms enable effective model learning?- Why does asymmetric self-play create naturally calibrated difficulty better than fixed curricula?
- What separates bootstrapping gains from sustained self-improvement gains?
- What distinguishes intrinsic metacognition from extrinsic human-designed loops?
- What other adaptive internal phenomena could signal system behavior improvements?
- What capabilities can emerge from self-modification that the original agent lacked?
- Does self-play feedback improve skills created from the agent's own experience?
- Why do current metacognitive training loops fail when agents encounter new domains?
- Should we train the evolver or the executor when building self-improving agents?
- How can agents evolve their own skills without human input?
- Why do self-improving agents concentrate progress in the fast non-parametric loop?
- Can self-improving agents become truly autonomous without intrinsic metacognition?
- Can co-evolved critics truly circumvent static evaluator limitations in self-improvement?
- Can AI systems generate and refine their own objective functions?