INQUIRING LINE

Teaching an AI new facts takes different routes, and each one trades upfront training cost against how easily you can change them later.

What knowledge injection routes trade flexibility against training cost?

This explores the different ways to get new knowledge into a language model (retrieval at query time, baking it into the weights, plug-in adapters, prompting, separate memory) and what each one costs in training effort compared with how easily you can update or swap it later.


This explores the main ways to give a language model knowledge it doesn't already have, and how each one trades training cost against flexibility. The corpus offers a useful starting map: four routes, each suited to a different deployment constraint How do knowledge injection methods trade off flexibility and cost?. Dynamic injection (RAG) needs no retraining and updates the moment you change the documents, but every query pays for the lookup in latency. Static embedding (fine-tuning the knowledge into the weights) gives the fastest answers but costs the most up front, and the knowledge is hard to revise afterward. Modular adapters sit in between: you train small add-on modules that can be swapped in and out. Prompt optimization costs nothing to train. Notably, the same note reports that combining routes beats any single one, so the real design question is how to mix them.

The cheap end has a hard limit. Prompting can only bring forward knowledge the model already absorbed in pretraining. It reorganizes what is there and cannot add anything new Can prompt optimization teach models knowledge they lack?. If the domain knowledge was never in the training data, no clever prompt will supply it. That is why prompting counts as a separate route rather than a cheaper version of the others.

The tradeoff isn't fixed, either, and two findings suggest that how you train matters as much as how much. StructTuning reaches about half of full knowledge-injection performance with just 0.3% of the training data. It does this by sorting text chunks into an automatically generated domain taxonomy, so the model learns where each fact sits in the field the way a student learns from a textbook's chapter structure Can organizing knowledge structures beat raw training data volume?. A related note argues that small amounts of explicit, structured knowledge fix problems that learning purely from data leaves behind, such as poor robustness and representations nobody can interpret Does refusing explicit knowledge harm AI system performance?. On the training-signal side, RLAG uses reinforcement learning that rewards both correct answers and sound explanations. It embeds domain knowledge more coherently than standard supervised fine-tuning Can reinforcement learning embed domain knowledge more effectively than supervised fine-tuning?. A broader overview warns that every adaptation method has a domain-specific sweet spot and hidden costs. Visible accuracy gains can come with quieter losses in reasoning faithfulness and format flexibility How do domain training techniques actually reshape model behavior?.

The less obvious part of the map is a set of routes that avoid touching the main model at all. MeMo trains a separate memory model to hold the new knowledge Can a separate memory model inject knowledge without touching the LLM?. Unlike RAG, its inference cost doesn't grow with the size of the corpus, and it works with frozen proprietary models you can't fine-tune. The price is up-front training and a fixed capacity limit. In agent research, the same idea shows up as learning without weight updates. VOYAGER stores skills as executable code in a searchable library and builds complex skills out of simpler ones, which avoids catastrophic forgetting Can agents learn new skills without forgetting old ones?. AgentFly improves an agent's behavior entirely through episodic memory modules and reaches 87.88% on GAIA with the model's parameters untouched Can agents learn continuously from experience without updating weights?.

So the choice is less a two-way split between flexible and expensive than a question of where the knowledge lives: in the prompt, in an external store, in a plug-in module, in a companion model, or in the weights. Each location fixes when you pay (at training time or on every query) and what you can still change later. The corpus has less on direct, head-to-head cost comparisons across all these routes. The four-way taxonomy note is the closest thing to one, so start there before the specialized papers.


Sources 9 notes

How do knowledge injection methods trade off flexibility and cost?

Dynamic injection (RAG) maximizes flexibility but adds latency; static embedding is fastest but costly and inflexible; modular adapters balance efficiency with swappability; prompt optimization requires no training but only activates existing knowledge. Combining all three outperforms any single approach.

Can prompt optimization teach models knowledge they lack?

Prompting works entirely within a model's pre-existing training distribution and cannot supply domain knowledge absent from training data. This creates a hard ceiling: no prompt strategy can compensate for missing foundational knowledge, only reorganize what already exists.

Can organizing knowledge structures beat raw training data volume?

StructTuning achieves 50% of full-corpus performance using only 0.3% of training data by organizing chunks into auto-generated domain taxonomies. The model learns knowledge position within conceptual structures rather than raw text patterns, matching how students learn from textbooks.

Does refusing explicit knowledge harm AI system performance?

AI systems that learn exclusively from data produce uninterpretable representations, inherit statistical biases uncorrected by normative rules, and fail to generalize beyond training distributions. Structured knowledge injection at minimal corpus cost substantially improves performance.

Can reinforcement learning embed domain knowledge more effectively than supervised fine-tuning?

RLAG rewards both answer accuracy and explanation rationality by cycling between augmented and unaugmented generation, progressively internalizing coherent knowledge structures. This outperforms SFT because it prioritizes reasoning quality over token-level correctness.

Show all 9 sources
How do domain training techniques actually reshape model behavior?

Research shows every adaptation method—from parameter-efficient tuning to knowledge graph curricula—has optimal conditions tied to specific domains. The key finding: visible benefits like performance gains often come with hidden degradation in reasoning faithfulness, capability transfer, and format flexibility.

Can a separate memory model inject knowledge without touching the LLM?

MeMo trains a dedicated memory model to encode new knowledge, eliminating inference-time search costs that scale with corpus size. It avoids fine-tuning risks and works with frozen proprietary models, but trades this for up-front training cost and capacity limits.

Can agents learn new skills without forgetting old ones?

VOYAGER demonstrates that storing executable skills in an embedding-indexed library and composing complex skills from simpler ones allows agents to learn continuously while avoiding the forgetting that occurs with weight-update-based methods. Environmental feedback refines skills while an automatic curriculum drives continual exploration.

Can agents learn continuously from experience without updating weights?

AgentFly formalizes agent learning as a Memory-augmented MDP with three memory modules (case, subtask, tool) that enable credit assignment and policy improvement entirely through memory operations. The approach achieved 87.88% on GAIA validation without modifying LLM parameters.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.