INQUIRING LINE

What kind of training material teaches an AI to explain known ideas clearly, not just know facts?

What training data would let models learn to organize established knowledge compellingly?

This explores what kind of training material would teach a model not just to know facts, but to arrange what is already known into clear, well-structured explanations. The corpus has strong material on the organizing half and very little on the 'compelling' half.


This explores what training data would help a model take knowledge that already exists and arrange it well: putting ideas in order, showing how they connect, and building from basics to harder material. The corpus's clearest answer is that structure matters more than volume. In StructTuning, training text was sorted into an automatically generated taxonomy of the domain before training, so the model learned where each piece of knowledge sits within the larger subject. That reached about half of full-corpus performance using just 0.3% of the data Can organizing knowledge structures beat raw training data volume?. The authors compare it to learning from a textbook rather than from a pile of loose pages. Medical knowledge graphs point the same way. Paths through the graph were turned into 24,000 reasoning tasks that combine simple facts into more complex chains, and a 32B model trained on them beat much larger systems across 15 medical specialties Can knowledge graphs teach models deep domain expertise?.

A less obvious point is that the useful material may be how-to text, not lists of facts. An analysis of 5 million pretraining documents found that reasoning draws on broad procedural knowledge (worked examples and step-by-step methods spread across many sources), while factual recall depends on memorizing particular documents Does procedural knowledge drive reasoning more than factual retrieval?. If organizing knowledge is a skill, then demonstrations of organizing are probably worth more than more facts to organize. A related result: models trained to predict whole concepts as well as the next word matched a comparable model's training loss using about half the tokens Can models learn faster by predicting their own concepts?. This suggests that directly rewarding larger units of meaning is efficient.

The corpus also shows that this structure can be produced in bulk. LLMs can draft and refine a set of categories for a body of text, then label the text with them Can LLMs efficiently generate taxonomies and label training data?, which is the kind of scaffolding StructTuning depends on. Models also connect knowledge on their own: they can reconstruct facts that no single document states by linking scattered hints across their training data Can LLMs reconstruct censored knowledge from scattered training hints?. So the training data may not need to spell out every link. It may only need to make the overall layout of the subject easy to find.

There are two limits worth knowing. First, organizing is not the same as storing. Memorizing facts in the weights is capped by model size, and fine-tuning in new facts can overwrite older ones Can models store unlimited facts without growing larger?. Prompting can only rearrange what the model already has; it cannot add knowledge Can prompt optimization teach models knowledge they lack?. One sensible split is to let the model learn how to structure knowledge and let retrieval or tools supply the facts. Second, domain training often has hidden costs, such as less faithful reasoning or less flexibility with formats, even when benchmark scores rise How do domain training techniques actually reshape model behavior?. One encouraging result: small models learned to stick to their source passages, quote them, and decline to answer rather than invent, because of how their training examples were designed, not because of model size Can small models learn to ground answers in context?.

The gap is the word 'compellingly.' Nothing here studies narrative, persuasion, explanation quality or teaching style as a training target. The corpus can tell you how to make a model's knowledge well organized, but not how to make its explanations engaging. That is an open question, not a settled one.


Sources 10 notes

Can organizing knowledge structures beat raw training data volume?

StructTuning achieves 50% of full-corpus performance using only 0.3% of training data by organizing chunks into auto-generated domain taxonomies. The model learns knowledge position within conceptual structures rather than raw text patterns, matching how students learn from textbooks.

Can knowledge graphs teach models deep domain expertise?

Fine-tuning a 32B model on 24,000 reasoning tasks derived from medical knowledge graph paths produces state-of-the-art performance across 15 medical domains, demonstrating that structured knowledge composition matters more than scale.

Does procedural knowledge drive reasoning more than factual retrieval?

Analysis of 5 million pretraining documents shows reasoning relies on broad, transferable procedural knowledge from diverse sources, unlike factual recall which depends on narrow, document-specific memorization of target facts.

Can models learn faster by predicting their own concepts?

An 8.9B model trained to predict both tokens and learned concepts from its hidden states matched OLMo-3-7B's final loss using only 51.3% of training tokens and outperformed it by 2.45 points downstream. This suggests explicit supervision of multi-token semantic structure improves compute efficiency.

Can LLMs efficiently generate taxonomies and label training data?

TnT-LLM automates text mining by using LLMs for open-ended reasoning to create and refine label taxonomies and generate training labels, then distilling these into lightweight classifiers for cost-effective deployment at scale.

Show all 10 sources
Can LLMs reconstruct censored knowledge from scattered training hints?

Language models perform out-of-context reasoning across the full training distribution, reconstructing information never explicitly stated in any single document. Experiments show models can infer city identities from scattered distance relationships and apply them downstream without in-context learning.

Can models store unlimited facts without growing larger?

A formal proof and experiments show in-weight memorization is bounded by model size, while tool-use enables unbounded factual recall through a simple circuit. In-weight finetuning also degrades general capability by overwriting prior knowledge.

Can prompt optimization teach models knowledge they lack?

Prompting works entirely within a model's pre-existing training distribution and cannot supply domain knowledge absent from training data. This creates a hard ceiling: no prompt strategy can compensate for missing foundational knowledge, only reorganize what already exists.

How do domain training techniques actually reshape model behavior?

Research shows every adaptation method—from parameter-efficient tuning to knowledge graph curricula—has optimal conditions tied to specific domains. The key finding: visible benefits like performance gains often come with hidden degradation in reasoning faithfulness, capability transfer, and format flexibility.

Can small models learn to ground answers in context?

Sub-2B models trained on synthetic multi-hop QA can ground answers in passages, cite literal quotes, and abstain from confabulation. The OCC-RAG work shows faithfulness emerges from training curriculum design, not parameter count.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.