Beyond Fixed Representations: The Vocabulary and Verifier Gaps in Open-Ended AI
Abstract Modern AI systems are increasingly being evaluated for their ability to reason, code, prove theorems, use tools, and long-horizon research tasks. These are powerful capabilities, but they share a structural limitation: the representational frame within which the model operates, including its conceptual vocabulary, the space of admissible solutions it can search, and the criteria by which success is evaluated, is typically fixed and supplied in advance. This paper argues that building stronger intelligent systems capable of open-ended innovation requires additional classes of operations: the creation, stabilization, and reuse of new representational primitives, which alter the space being searched rather than simply searching within it. We characterize the distance between current AI systems and genuinely open-ended intelligence through two gaps. The first is the vocabulary gap, the difficulty of inventing and stabilizing new representational primitives rather than merely recombining existing ones. The second is the verifier gap, the difficulty of judging the value of a new primitive when its full payoff may be visible only after future reuse. We interpret both gaps through a unified framework of intelligence as cognitive discrepancy reduction. By viewing intelligent behaviors as a sequence of cognitive transformations, we distinguish intra-space transformations which operate within a fixed representational frame, from generative transformations which may modify the frame itself. On this basis, we propose a ladder of innovation autonomy and outline several directions for advancing open-ended AI, including objectives that reward useful representational change, persistent memory architectures for invented primitives, and adaptive verification mechanisms capable of evolving alongside the representations they evaluate.
Introduction. There are two types of “intelligent” activities that are often conflated. The first is solving problems within a given frame: given a goal, a trained model with fixed representation space, and a way to check answers, find a good solution. The second requires changing the frame itself, for example creating new concepts, relations, measurements, abstractions, or evaluators that makes a class of problems previously unsolvable under pre-existing frames expressible and solvable. The first activity searches within a space, while the second changes the space being searched itself. Contemporary AI excels at the first type and is still primarily benchmarked on tasks in this category. Reasoning puzzles, competition mathematics, code generation, theorem proving, tool use, and agentic task completion all evaluate performance within problems whose representational frame is largely fixed in advance—a model’s weights are fixed by training data distribution; A benchmark supplies the task, the admissible form of an answer, and the standard of success; A coding environment supplies a programming language and tests; A theorem prover provides a formal language and a checker. An AI-for-science pipeline supplies a topic, a tool stack, a literature corpus, and procedures for validation. Such framed problems are useful for measuring progress, but they also impose an important limitation: our headline metrics are, by construction, weak tests of the second that require stronger capability. They tell us relatively little about whether a system can tackle genuinely open-ended innovation tasks in which the crucial step is not merely finding a solution within a given space, but recognizing that the space itself is inadequate and that a new representational primitive must be introduced. The history of scientific and social progress can be understood, in large part, as a history of expanding representational spaces through open-ended new concept creation. Civilization advancement depends on natural, mathematical, and logical languages that progressively enlarge what can be symbolized, measured, and reasoned about. The concept of “number” reorganized heterogeneous collections into a common quantitative representation. The invention of “negative numbers” made previously impermissible algebraic operations meaningful. “Entropy” introduced a vocabulary for expressing regularities that were not naturally visible within older representation frameworks. Similarly, concepts such as energy, information, feedback, gene, algorithm, and attention have become representational primitives which restructure the space of what can be described, searched, measured, and reused. In the social domain, concepts such as rights, markets, incentives, and laws have played a comparable role in organizing collective behavior and political reasoning. Progress, in this sense, is not simply the accumulation of facts within a fixed framework. It is the continual transformation of the framework itself, making previously inaccessible patterns thinkable. This is where current AI systems remain limited. Existing models are increasingly powerful at solving the type-one problems that operate within a given and frozen representational framework, where the representation primitives, search space, and evaluation criteria are largely fixed in advance. Recent analysis on LLM-driven autonomous research shows that LLM-generated research ideas are narrow and highly-concentrated on the pattern of recombining and synthesizing multiple existing ideas [4, 8, 55], which differs significantly from the human ideation patterns. To climb higher on the ladder of AGI requires systems that can also handle type-two problems: problems whose solutions require open-ended representational expansion, concept formation, and the creation of new primitives that were not supplied by the initial framework to enable autonomous innovation. In this paper, we analyze this limitation and make the following claims. First, we argue that the distance between current AI systems and genuinely open-ended intelligent systems can be characterized by two central gaps. The first is the “vocabulary gap”, the difficulty of inventing and stabilizing new representational primitives, rather than merely recombining existing ones supplied from outside. The second is the “verifier gap”, the difficulty of judging whether a newly introduced primitive is worth retaining when its value is not yet validated by any criterion the system currently possesses. Second, we interpret these innovation gaps through the lens of cognitive discrepancy reduction. Under this unified view, intelligence is driven by the pressure to reduce gaps between current states and unresolved demand states by prediction, explanation and creation. Concept invention becomes necessary when existing representations cannot sufficiently reduce such discrepancies.
Related work. This paper connects four lines of work that are often discussed separately: open-endedness and its crucial role in intelligence, cognitive foundations of innovation, self-modifying discovery systems, recursive self-improvement and automated research agents. The common thread is the question of whether an intelligent system can move beyond solving tasks inside a fixed representation frame and instead revise the primitives, search space, and evaluative criteria that make meaningful discovery possible.
6.1 Open-Endedness: Necessity and Limits Open-ended innovation is what allowed humans to create knowledge, make discoveries, and build civilization. Modeling this capacity is substantially harder than modeling closed-ended intelligence, where the task, vocabulary, objective, and verifier are fixed in advance. [6] argues that creativity is a core feature of intelligence and an unavoidable challenge for AI, grounded in cognitive capacities such as association, analogical thinking, perception, search, and reflective self-criticism—the operations that we formalize as discrepancy-reducing transformations in Section 3.1. Recent position work sharpens this necessity claim. Hughes [26] argue that open-endedness is essential for artificial superintelligence, defining it behaviorally through the novelty and learnability of a system’s output stream relative to an observer. Genewein et al. [17] argues that if intelligence is viewed as search, then a suitable open-ended search process yields progressively stronger general performance, identifying superintelligence with super-creativity. These framings are complementary to ours, differing in level of analysis: behavioral definitions ask whether a system produces a continuing stream of novel, learnable artifacts, whereas this paper asks what internal operation produces such behaviors—representation-space expansion, constrained by the verifier gap. Avestimehr et al. [3] provide a complementary limitation result: if retraining is support-preserving, an autonomous system cannot discover valid artifacts outside its initial generative support, so breaking exploration barriers, valid verification, and human amplification through guidance are each necessary for strong knowledge discovery. This is a formal complement to our informal diagnosis—their support barrier is closely related to the vocabulary gap, and their verification condition to the verifier gap.
6.2 Cognitive Foundations of Innovation Three operations underlie the production of novelty: abstraction, analogy, and composition. In our framework, these operations are treated as discrepancy-reducing transformations over representational space (Section 3.1). We situate the relevant prior work below.
Method. 2 The Distance from Open-ended Intelligence As argued above, existing AI systems have become increasingly capable within fixed representational schemes, predefined objectives, and established problem frames. However, they remain far from stronger forms of open-ended intelligence. Such intelligence does not merely solve given problems, it also autonomously expands the space of possible problems, concepts, and methods through which new solutions can be formulated. In this section, we identify two central gaps that must be narrowed for AI systems to acquire such open-ended capabilities: the vocabulary gap and the verifier gap.
2.1 The Vocabulary Gap Genuine discovery rests on concept invention. Before a problem can be searched, someone must supply the primitives in which candidate solutions are written, and the deepest advances are often not new answers but new primitives: a variable that did not previously exist, a relation no one had isolated, a measurement that makes a hidden regularity legible. Scientific and social progress, read this way, is in large part a history of vocabulary changes. The matrix is invented to abstract arrays of numbers and eigenvalues to capture a matrix’s invariant structure. The concept of economy abstracts the production, exchange, and management of resources into an object that can be measured and studied. Current AI systems are powerful searchers and recombiners within a provided vocabulary, but they rarely decide, on their own, how and when the vocabulary should grow. Concept creation is an act of abstraction, it extracts invariant attributes shared across a diverse range of objects. Grouping a wide range of observations under a single reusable unit shortens the descriptions of everything that uses it. This is the minimum-description-length (MDL) view of abstraction, where a primitive justifies its value by compressing a family of observations, reducing the representational budget required to express them. In trivial cases, attaching almost any symbol to a set of raw perceptions reduces the description length and renders something newly reachable within budget. However, the critical requirement is amortization: a primitive must justify its representational value across a family of problems, not a single local instance. “Number” applies to any collection of objects, not only those currently being observed. “Policy” applies to any rule an organization might enforce, not only the actions already taken. To make this precise, let LL(·) denote description length under a language L, and let LB L(f) denote the length of the shortest solution to task f discoverable within the search budget B under L (taken as ∞if no solution is found within budget). For a task family F, adding a primitive π to obtain L′ = L ∪{π} is generative with respect to F and B when both hold:
Current models do form concepts and build vocabularies, but these are inherited from the training distribution rather than initiated autonomously.1 The abstractions a model acquires are shaped by the data distribution and limited contexts where they appear. Open-ended intelligence requires the stronger capacity to introduce primitives without that guidance. Imagine a model that has never been trained on mathematics so it has no concept of “number”. Now show the model scenes involving objects in different quantity, like five apples, seven chairs, three people. Would it be able to autonomously create a concept of “number” and augment its own vocabulary and representation space for future reuse? No current system does this, and without such an autonomous vocabulary expansion, open-ended innovation is out of reach. 2 2.2 The Verifier Gap When the representation frame is fixed, evaluation is usually fast, cheap, and decisive. A generated program passes its tests or fails, a proof is accepted by a checker or rejected, a design either improves an objective or it does not. The verifier is already defined over a fixed representation and candidate space. The strongest successes in recent AI-assisted discovery occur largely in this regime.
Discussion. 5 What Would Have to Change If the vocabulary and verifier gaps are central, then the path to open-ended AI cannot be reduced to scaling models, extending context windows, adding tools, or increasing inference-time search. These techniques improve search within an existing frame, but they do not address how a system should recognize that the frame itself is inadequate. The deeper challenge is to build systems that can generate, stabilize, test, revise, and reuse new primitives under uncertainty about their future value, and in the hardest cases, construct the very criteria by which that value can be judged. Several directions follow.
Objectives that reward useful representation change Next-token prediction (NTP) can be interpreted as a specific instance of the discrepancy-reduction schema in Section 3.1: the model is trained to reduce the discrepancy between a context representation and the next-token distribution, typically measured by cross-entropy. In this case R is the model’s representation of the token context, G is the observed next token as a target distribution over the fixed vocabulary, D is the cross-entropy loss, and the transformation T is identity T = I since there is no representation frame change. This objective has been remarkably effective, but it rewards abstraction only indirectly. A representation is favored only insofar as it improves prediction over the training distribution. The model is not explicitly trained to notice that its current vocabulary is inadequate, introduce a new abstraction, and preserve that abstraction for future reuse.
The objective should therefore be expanded from answer prediction to representation revision. A training episode should not only ask which output is correct, but also whether the current representation leaves a recurring discrepancy unresolved. A system should receive credit for triggering an operation T that produces a more useful representation R′ = T(R) for solving a problem, including abstraction, compression, relational mapping, composition etc. For example, the current state R may contain a collection of observations that appear heterogeneous under the existing vocabulary. The target G is not a single next token, but a representation that captures their shared structure. The discrepancy D then measures how poorly the current representation explains, compresses, or transfers across the observations. A useful transformation is one that reduces this discrepancy by introducing an abstraction that can be reused in later cases. Under the general schema of Eq.1, richer objectives can therefore be constructed by varying the representation R, the transformation T, and the discrepancy measure D. The goal is not to replace prediction, but to augment it with explicit reward for useful representational change. This would make concept formation and abstraction direct targets of learning rather than accidental byproducts of NTP.
Data that elucidates invention, not only its outcome Enrichment of objectives requires corresponding change in data. Almost all existing training data exhibits concept use rather than concept invention. Mathematics already uses the established concept of “eigenvalue”, physics already uses the concept “entropy”. A distribution that only displays established frames is a weak teacher of frame change, because the crucial act of novel primitive introduction is usually absent. Data for open-ended intelligence should instead expose trajectories of representational revision. A useful trajectory would show an initial representation Rt, a set of observations or tasks that produce a persistent discrepancy under that representation, a transformation Tt that introduces a candidate primitive, and a revised representation Rt+1 in which the discrepancy is reduced.
Conclusion. Human intelligence is defined not only by task performance but by innovation. Civilization is, in large part, an unbounded recursion in what can be represented, constructed, and transformed. Every scientific theory, medical treatment, transportation system, and social institution demonstrates that human intelligence does not merely solve predefined problems but creates problem spaces that did not previously exist. This capacity, to reshape environments, generate new concepts, and expand the space of what can be known and built, is largely absent from today’s AI systems. While rapid progress on intellectually demanding tasks such as coding and mathematics has been made, these are mainly advances within fixed representational frames. Higher forms of open-ended intelligence require something further, so that a system can refine and expand its representational space on its own. This paper has argued that two gaps stand in the way.
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
Can we trust AI-generated mathematical proofs without understanding them?- What would it mean for mathematics to define itself before AI transformation?
- Does verification by inspection scale for AI mathematics discoveries?
- What verification methods can prove AI mathematical proofs are sound?
- Does AI-assisted research hollow out the understanding that producing proofs generates?
- Can self-administered surveys establish trustworthy AI capability benchmarks?
- Can self-reported AI reliability metrics hide confounding factors like task complexity?
- What counts as knowledge versus skilled performance in AI-mediated learning?
- Does the answer-versus-tutor distinction hold across subjects beyond math and programming?
- Can independent validation of AI output substitute for method disclosure?
- Can self-reported confidence measures predict actual AI task performance?
- Why do users interpret AI outputs through frameworks meant for human experts?