AI tools look like they learn and plan ahead on their own — so why don't deployed models actually do that?
Why do deployed models lack the learning and planning Weinstein attributes to them?
This explores why the AI systems people actually use don't keep learning from experience or plan ahead on their own, even though Weinstein describes them as if they did. None of the retrieved notes discuss Weinstein's argument directly, so this answer covers the mechanics behind the gap rather than his specific claims.
This explores why the AI systems people actually use don't keep learning from experience or plan ahead on their own, even though Weinstein describes them as if they did. None of the retrieved notes discuss Weinstein's argument directly, so this answer covers the mechanics behind the gap rather than his specific claims. The short version from the corpus is that a deployed model's weights are frozen once training ends. When a system seems to learn or plan, that work is usually done by the scaffolding built around the model, not by the model itself.
Take learning first. A deployed model doesn't update itself when you talk to it. Agents that do improve with experience get there by keeping memories outside the model. AgentFly stores past cases, subtasks and tool uses, and reaches strong benchmark scores without changing a single parameter Can agents learn continuously from experience without updating weights?. VOYAGER builds a growing library of executable skills and does this partly to avoid the 'catastrophic forgetting' that happens when you keep retraining the weights Can agents learn new skills without forgetting old ones?. So when a system seems to learn, the useful question is where the learning is actually stored. Usually it's in a database or a skill library, not in the model. Some researchers are trying to move memory back inside the model Should agent memory live inside the model backbone?, which is an admission that today's models don't have it built in.
Planning shows the same pattern. Many planning gains come from the execution harness, the software around the model that breaks work into steps, tracks state and retries. One harness raised scores across several frozen models, and the same setup carried over to newer models unchanged Can execution harnesses lift model performance without retuning weights?. Another study gave a weaker planner a map linking runtime behavior to the code responsible for it. With that map, the weaker planner located code as well as stronger models did Can explicit behavior maps help weaker planners compete with stronger models?. When planning ability can be swapped in from outside like this, it's a sign that much of it never lived in the model.
The less obvious part concerns what training adds in the first place. Several notes suggest RL post-training mostly teaches a model when to use reasoning it already had, rather than giving it new reasoning ability Does RL post-training create reasoning or just deploy it?. RL also changes only a small, consistent fraction of the parameters Does reinforcement learning update only a small fraction of parameters?. It can even narrow what the model is able to find: base models with light prompting eventually find more solutions than their post-trained versions Do base models find more solutions than post-trained ones?. One counterpoint is worth knowing. Post-trained models do seem to recognize that their outputs shape what they see next Do models recognize their own outputs as actions shaping future inputs?. That's an early form of acting in the world, but it isn't the same as planning or learning over time.
Put together, a deployed model is a fixed engine. Over a few training stages, it has been tuned to know when to use abilities it already had. The parts that look like an agent, such as memory, curricula and step-by-step execution, are mostly added by engineers around it. If Weinstein credits the model itself with learning and planning, the corpus suggests he is crediting it with work the surrounding system does.
Sources 9 notes
AgentFly formalizes agent learning as a Memory-augmented MDP with three memory modules (case, subtask, tool) that enable credit assignment and policy improvement entirely through memory operations. The approach achieved 87.88% on GAIA validation without modifying LLM parameters.
VOYAGER demonstrates that storing executable skills in an embedding-indexed library and composing complex skills from simpler ones allows agents to learn continuously while avoiding the forgetting that occurs with weight-update-based methods. Environmental feedback refines skills while an automatic curriculum drives continual exploration.
Metis demonstrates that agent memory can be implemented as a persistent state and autonomous procedures within the model backbone rather than external modules. This approach enables end-to-end training and avoids the decoupling failures where external memory and backbone optimize independently.
StateM improves Terminal-Bench 2.1 accuracy across multiple models by optimizing execution systems around frozen weights. The same runbook transfers to newer models without modification, achieving 95.3% on GPT-5.6 and lifting DeepSeek-V4 Flash by 5.4 points.
A behavior-to-code mapping representation improved win rates by 10–19 points while reducing planner tokens by 8–13%. Weaker planners using this mapping matched stronger models' code localization across all precision and recall metrics.
Show all 9 sources
Evidence shows base models already contain reasoning capability in latent form; RL training optimizes deployment timing rather than capability creation. Hybrid models recover 91% of performance gains by routing tokens only, and activation vectors for reasoning strategies pre-exist before any RL.
Across seven RL algorithms and ten LLM families, RL induces intrinsic parameter sparsity of 5–30% without explicit regularization. Critically, these sparse updates are nearly full-rank and nearly identical across random seeds, indicating structural rather than arbitrary parameter selection.
Across 14 model pairs and three agentic benchmarks, base models equipped with only relaxed system prompts eventually surpass post-trained counterparts in pass@K coverage as rollout budget grows. Post-training bimodalizes task outcomes, sharpening performance on easy cases while eliminating rare-but-reachable solutions.
Post-trained language models exhibit a measurable shift where they recognize their outputs become their own future inputs, closing an action-perception loop absent in pretraining. Evidence includes 3-4x lower output entropy on-policy and behavioral signatures of trajectory recognition.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Sharpening Tax in Post-Training
- SkillRL: Evolving Agents via Recursive Skill-Augmented Reinforcement Learning
- Useful Memories Become Faulty When Continuously Updated by LLMs
- Continual Learning Bench: Evaluating Frontier AI Systems in Real-World Stateful Environments
- Know It, Act on It: Investigating Memory Utilization in LLM Personalization
- Agent Memory Distillation: Empowering Small LLM Agents with Hierarchical Teacher Memory
- StateM: Reaching 95.3% Raw Accuracy, or a \$15 Frontier Run, on Terminal-Bench 2.1 via Harness Scaling
- On the Interplay of Pre-Training, Mid-Training, and RL on Reasoning Language Models