AI systems can change gradually as they learn, or suddenly flip in behavior — what actually separates those two kinds of change?
What distinguishes learnable perturbations from bifurcation-triggered motive shifts?
This reads the question as asking how two kinds of change in an AI system differ: small, gradual adjustments that training can absorb and learn from, versus sudden tipping points where the system's behavior or goals flip all at once. The corpus doesn't address 'bifurcation-triggered motive shifts' directly, so this answer covers the nearest material and points out where the gap is.
This reads the question as asking how two kinds of change in an AI system differ: small, gradual adjustments that training can absorb and learn from, versus sudden tipping points where the system's behavior or goals flip all at once. To be direct: the collection has no notes on bifurcations, on how models' motives change, or on learnable perturbations as a topic. What it does have is several notes that, read together, sketch the difference between smooth change and abrupt change in how models think and learn.
On the gradual side, the clearest doorway is the work on drift during continual learning. Models trained to stay close to their starting distribution keep their ability to learn new tasks, while models that drift further stall when the task domain changes Does staying close to the base model preserve learning ability?. That gives a practical meaning to 'learnable perturbation': a change small enough that the system can keep adjusting afterward. Agents that store lessons outside their weights, as written reflections or as a library of skills, push the same idea further. They change their behavior without disturbing the underlying model at all Can agents learn from failure without updating their weights? Can agents learn new skills without forgetting old ones?.
On the abrupt side, the nearest material treats a model's internal state as something that moves over time and can settle or shift suddenly. Looped transformers can tell when to stop computing by noticing that their internal state has stopped changing, a stable resting point, and this works better than a learned 'stop' signal Can fixed points replace learned halt tokens in reasoning models? Can models learn by looping instead of growing larger?. Reasoning models show the opposite pattern: their hidden states loop back on themselves, and those cycles line up with the 'aha moments' where a model drops one answer and reconsiders Do reasoning cycles in hidden states reveal aha moments?. These are the closest things in the collection to a qualitative shift, a point where the model's path changes direction instead of just moving further along it.
The surprising part is a bridge between the two. How a reward is designed can push a model toward an extreme without anyone intending it. Binary right/wrong rewards steadily train models into overconfident guessing, because a confident wrong answer costs no more than a hesitant one. A small change to the reward fixes this Does binary reward training hurt model calibration?. So a shift in what a model is effectively 'aiming for' may come less from sudden internal flips and more from the incentives it's trained under, which can be corrected. If you came looking for dynamical-systems accounts of AI goals changing, the collection doesn't cover that yet. The notes above are the best starting points.
Sources 7 notes
FST-trained models stay up to 70% closer to their base distribution than parameter-only RL, and this reduced drift preserves the model's ability to learn subsequent tasks effectively. Parameter-only approaches stall when task domains change, while low KL drift enables sustained adaptation.
Reflexion demonstrates that unambiguous environmental feedback (success/failure) enables agents to write useful self-diagnoses and improve across episodes without parameter updates. The binary signal prevents rationalization, and keeping reflections uncompressed preserves their usability.
VOYAGER demonstrates that storing executable skills in an embedding-indexed library and composing complex skills from simpler ones allows agents to learn continuously while avoiding the forgetting that occurs with weight-update-based methods. Environmental feedback refines skills while an automatic curriculum drives continual exploration.
FPRM shows that looped transformers halt more accurately by detecting when their latent state reaches a fixed point, calibrating compute closer to the accuracy-saturation point than learned halt tokens without requiring special training regimes.
Models that re-apply layers in recurrent depth outperform larger feedforward networks on reasoning tasks. This works because recursion enables state tracking and compositional generalization that parameter scaling alone cannot achieve, with convergence signals providing natural halting.
Show all 7 sources
Distilled reasoning models show ~5 cycles per sample versus near-zero in base models, and cyclicity correlates with accuracy. These cycles in hidden-state reasoning graphs directly map to RL-trained models' documented aha moments—moments when models reconsider intermediate answers.
Binary correctness rewards incentivize high-confidence guessing because they don't penalize confident wrong answers. Adding the Brier score as a second reward term mathematically guarantees joint optimization of accuracy and calibration without trade-off.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- SkillClaw: Let Skills Evolve Collectively with Agentic Evolver
- Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach
- A Mechanistic Analysis of Looped Reasoning Language Models
- Loop the Loopies!
- Can Looped Transformers Learn to Implement Multi-step Gradient Descent for In-context Learning?
- Generative Recursive Reasoning
- Hierarchical Reasoning Model
- CLIN: A Continually Learning Language Agent for Rapid Task Adaptation and Generalization