Can a model's learned tricks be packaged into small, swappable pieces that plug into a system without scrambling each other?
How do typed adapters move learned changes between isolated surfaces?
This explores how small, separate modules (adapters, prompts, skill libraries) can hold what a model has learned and carry it from one place to another, such as between users, tasks, or the weights-versus-text layers of a system, without mixing those places together. The collection has no paper on 'typed adapters' by that name, so this answer covers the nearby territory it does have.
This explores how learned behavior can be packaged into small, separate units that move between parts of a system without bleeding into each other. One thing up front: the collection has nothing on 'typed adapters' as a formal idea, meaning adapters with declared types or interfaces that control where they can plug in. What it does have is a clear picture of the parts such a system would be built from, and of where its boundaries can fail.
Start with the adapter as a portable unit of change. One line of work treats parameter-efficient fine-tuning (PEFT) adapters as lasting 'behavioral deltas': one strong shared base model plus millions of small adapters, each one carrying what was learned about a single user Can lightweight adapters replace millions of personalized models?. Here the adapter is the thing that moves, and the shared base is what keeps it portable. The same idea works for traits: PsychAdapter adds less than 0.1% extra parameters to every transformer layer and gets reliable personality control across GPT-2, Gemma and Llama 3 Can we control personality in language models without prompting?. So a small module can carry the same learned change into several different model families. ReFT goes a step further. It leaves the weights alone and learns edits to the model's hidden representations, using 10 to 50 times fewer parameters than LoRA Can editing hidden representations beat weight updates for finetuning?. That suggests the right 'surface' for an adapter might be the model's internal representations, not its weights.
The isolation half of the question gets its sharpest answer from multi-task fine-tuning. When tasks share parameters, they interfere with each other. Scheduling them at different times doesn't fix this. What works is finding each task's core parameters, freezing them, and merging only the rest Can isolating task-specific parameters prevent multi-task fine-tuning interference?. Fast-Slow Training draws a bigger boundary: lessons specific to one task go into optimized prompts (fast, text), and only minimal changes go into the weights (slow). That reduces forgetting, and the authors argue forgetting is a *misallocation* problem: learning that was sent to the wrong surface Can splitting adaptation into two channels reduce forgetting?. VOYAGER takes this to the extreme. Its skills live outside the model as executable code in an indexed library, so they can be combined and reused without touching weights at all Can agents learn new skills without forgetting old ones?.
Moving a change between surfaces is the part people often miss: it's not free. One study argues the real long-context bottleneck isn't storage space. It's the compute needed to turn context that has dropped out of the window into internal state during offline 'sleep' phases, and results improve with more consolidation passes Is long-context bottleneck really about memory or compute?. Moving learning from text into weights is a kind of translation that costs work. And not everything carries over cleanly. A map of reward-hacking defenses across weights, selection and text finds that some defenses work the same way on every surface, while others only have rough equivalents Which reward hacking defenses actually transfer across training substrates?. That is roughly what a 'type system' for adapters would need to record: which changes keep their meaning when they cross a boundary, and which only resemble their originals.
If you want to see how isolation looks on the control side, LLM Programs keep state inside an explicit algorithm and show each LLM call only the context it needs for its step Can algorithms control LLM reasoning better than LLMs alone?. That works as a kind of interface contract written in code. The surprising takeaway from the collection: the hard problem isn't packaging a learned change. Adapters, prompts and skill files all do that well. The hard problem is knowing which surface a change belongs on, and what it costs or loses when it moves to another one.
Sources 9 notes
PEFT adapters function as durable behavioral deltas carrying learned user experience, enabling a single strong base plus millions of lightweight adapters to replace millions of full models—but only when scale-up, scale-down, and scale-out reinforce simultaneously.
PsychAdapter modifies every transformer layer with <0.1% additional parameters to achieve 87.3% Big Five accuracy and 96.7% depression/life satisfaction accuracy across GPT-2, Gemma, and Llama 3. This architecture-level approach bypasses prompt resistance entirely.
ReFT learns task-specific interventions on frozen model representations rather than updating weights, with LoReFT (low-rank linear subspace variant) dramatically outperforming LoRA across reasoning, instruction-following, and NLU benchmarks while using far fewer parameters.
Research shows that identifying core parameter regions per task, clustering overlapping tasks, and freezing core parameters while geometrically merging non-core parameters consistently outperforms standard multi-task fine-tuning. Temporal task scheduling alone proves insufficient without explicit structural parameter isolation.
Fast-Slow Training routes task-specific lessons into optimized prompts while keeping parameter updates minimal, reaching equivalent performance 1.4–3x faster with substantially less catastrophic forgetting and plasticity loss, demonstrating that forgetting is a misallocation problem rather than an inherent cost.
Show all 9 sources
VOYAGER demonstrates that storing executable skills in an embedding-indexed library and composing complex skills from simpler ones allows agents to learn continuously while avoiding the forgetting that occurs with weight-update-based methods. Environmental feedback refines skills while an automatic curriculum drives continual exploration.
Research shows the bottleneck is not memory capacity but the compute required to consolidate evicted context into fast weights during offline sleep phases. Performance improves with more consolidation passes, following a test-time scaling pattern on harder reasoning tasks.
A systematic map identifies which defense mechanisms function identically across weights, selection, and text substrates versus which only provide functional analogies. Practitioners rated this correspondence analysis as their most immediately useful takeaway from the work.
LLM Programs embed LLMs within explicit algorithms that manage control flow and state, presenting only step-specific context to each LLM call. This information hiding addresses capability and context window limits while treating complex reasoning as modular, debuggable sub-tasks.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Learning, Fast and Slow: Towards LLMs That Adapt Continually
- PsychAdapter: Adapting LLM Transformers to Reflect Traits, Personality and Mental Health
- ReFT: Representation Finetuning for Language Models
- A Survey on Post-training of Large Language Models
- Context-PEFT: Efficient Multi-Modal, Multi-Task Fine-Tuning
- Lottery Ticket Adaptation: Mitigating Destructive Interference in LLMs
- SkillClaw: Let Skills Evolve Collectively with Agentic Evolver
- On the Scaling of PEFT: Towards Million Personal Models of Trillion Parameters