Can AI systems invent new concepts rather than reuse trained ones?
Current AI systems excel at search and reasoning within fixed representational frames, but can they autonomously create novel primitives like mathematicians invented negative numbers or entropy? This matters because genuine open-ended innovation may require frame-altering operations, not just frame-internal search.
The paper argues that current systems are strong at reasoning, coding, theorem proving and tool use, but all of these run inside a representational frame that is "typically fixed and supplied in advance." Open-ended innovation needs a second class of operation: creating, stabilizing and reusing new representational primitives, which "alter the space being searched rather than simply searching within it." It locates the distance to that capacity in two gaps. The vocabulary gap is "the difficulty of inventing and stabilizing new representational primitives rather than merely recombining existing ones." The verifier gap is "the difficulty of judging the value of a new primitive when its full payoff may be visible only after future reuse." Its examples, "number", "negative numbers" and "entropy", are concepts that became valuable by reorganizing many problems at once.
The mechanism is a minimum-description-length view of abstraction. A primitive earns its place by compressing a family of observations, and the paper's critical requirement is amortization: a primitive "must justify its representational value across a family of problems, not a single local instance." It frames intelligence as cognitive discrepancy reduction and separates intra-space transformations, which work inside a fixed frame, from generative transformations, which "may modify the frame itself." On the verifier side it observes that inside a fixed frame "evaluation is usually fast, cheap, and decisive": a program passes its tests or a proof is accepted by a checker. A newly invented primitive has no such checker yet, because its value shows only through later reuse.
This sharpens the nearest notes. Do foundation models learn world models or task-specific shortcuts? shows models that predict well without the structure a world model would give them. The paper makes the concept-level version of that point: current models form concepts, but "these are inherited from the training distribution rather than initiated autonomously." Its thought experiment, showing a model five apples and seven chairs and asking whether it would invent "number", is a test of initiation that the heuristics probe does not run. Can AI research itself without losing human oversight? is the nearest engineering counterpart, and it contrasts: its analyzer distills reusable insights, but its cognition base supplies human priors, which is the part the paper says systems do not originate. The "persistent memory architectures for invented primitives" the paper proposes are the step that note leaves open. Do language models fail at reasoning due to complexity or novelty? describes the same boundary from the evaluation side: search inside a covered frame fails at the novelty edge, and the paper's argument is that only frame change moves that edge.
What the excerpt does not establish: this is a position paper. It reports no experiments, and its claim that no current system autonomously creates a concept like "number" is asserted rather than tested. The "ladder of innovation autonomy" and the adaptive verifiers appear in the abstract but are not developed in the excerpt. The generative condition is introduced as holding "when both hold", and the excerpt does not list the two conditions. The implication is that the two gaps are a diagnosis with a testable prediction. A benchmark that rewards reuse of a newly introduced primitive across a family of tasks would show whether a system closes the vocabulary gap, and the paper does not supply one. At the strength the evidence allows, this is a conceptual framing, not an empirical finding about today's models.
Inquiring lines that read this note 1
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
Can we trust AI-generated mathematical proofs without understanding them?Related concepts in this collection 4
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Do foundation models learn world models or task-specific shortcuts?
When transformer models predict sequences accurately, are they building genuine world models that capture underlying physics and logic? Or are they exploiting narrow patterns that fail under distribution shift?
the probe documents the structural version of the inherited-concepts point the paper makes for vocabulary
-
Can AI research itself without losing human oversight?
Explores whether AI systems can internalize the human judgment and insight-distillation that normally drives research progress, and what this means for maintaining meaningful human control over AI advancement.
contrast: the analyzer reuses insights inside a fixed frame, on priors the paper says systems do not originate
-
Do language models fail at reasoning due to complexity or novelty?
Explores whether reasoning-model failures stem from task complexity thresholds or from encountering unfamiliar instances. Tests whether scaling chain length actually addresses the root cause of reasoning breakdown.
parallel: search inside a covered frame fails at the novelty boundary, which only frame change addresses
-
Can AIs learn to specify their own research objectives?
Rapid recursive self-improvement may depend on whether AIs can autonomously propose and pursue their own goals without deviating. This question separates specified autoresearch from open-ended scientific discovery.
the verifier gap sharpens this: who judges a new primitive when no criterion yet exists
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Beyond Fixed Representations: The Vocabulary and Verifier Gaps in Open-Ended AI
- ASI-Bench: At the Dawn of Artificial Superintelligence
- Mathematical methods and human thought in the age of AI
- Self-Organizing Graph Reasoning Evolves into a Critical State for Continuous Discovery Through Structural-Semantic Dynamics
- The Method of Critical AI Studies, A Propaedeutic
- Intelligence from Learnable Novelty
- Has the Creativity of Large-Language Models peaked? —an analysis of inter- and intra-LLM variability —
- Agentic AI Scientists Are Not Built For Autonomous Scientific Discovery
Original note title
open-ended AI needs operations that alter the search space, not only search within it — the vocabulary gap and the verifier gap