Line of inquiry
Inquiring lines›How do training methods and scalin…›How do training methods and scalin…›this line of inquiry
What capability trade-offs arise from domain specialization through fine-tuning?
A broader line of inquiry — a family of 72 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 72
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Does fine-tuning actually change model capabilities or only output distribution?
- How do finetuning and pretraining improvements differ in their effects on model capabilities?
- Does specialized training in one domain create capability cliffs elsewhere?
- How do retrieval and fine-tuning trade off flexibility against training cost?
- How much can externalized skills improve models before hitting diminishing returns?
- Why do proprietary models improve with training while open-source models decline?
- What performance trade-offs emerge when composing multiple independently trained model capabilities?
- Does parameter isolation per task enable online updates without retraining?
- Does fine-tuning a small model match fine-tuning a large one?
- How does behavioral fine-tuning differ from factual knowledge encoding in models?
- Why does fine-tuning change how models process retrieved context?
- What hidden costs emerge when you fine-tune models for a single domain?
- What training cost tradeoffs exist between fine-tuning and other knowledge injection methods?
- What capabilities actually require massive scale versus specialized training regimes?
- Does pretraining data size matter less than base model scale for finetuning?
- What happens to base model capabilities when you apply finetuning?
- Can expert vectors learned offline transfer across multiple model architectures?
- Can models adapt and combine search strategies beyond their training algorithm?
- Why do fine-tuned models fail outside their specialized domains?
- What role does inductive bias play versus model capacity in practice?
- How does the functional separation of knowledge and reasoning affect adaptation methods?
- Do different domains require different types of model investment?
- What makes frozen model reasoning different from weight-based parameter updates?
- What capability risks emerge when models are optimized for single domains?
- How much performance is lost when converting pretrained checkpoints versus training from scratch?
- What causes models to develop domain capability cliffs after specialization?
- Why does over-specialization create a domain capability cliff in LLMs?
- Can extracted skills transfer effectively across different domains and model architectures?
- Which finetuning method works best across different task and data regimes?
- Why do metric choices constrain which model capabilities get developed?
- Why does the same training data produce different gains across models?
- How does pretrained knowledge constrain what adaptation strategies can achieve?
- Does the optimal model size depend on what capabilities you actually need?
- What role does pretraining play in distinguishing system capability from deployed behavior?
- Can specialized components replace single fully-trained models in deployment?
- How does over-specialization create capability cliffs outside target domains?
- What trade-offs emerge between training objectives and model reliability?
- How do ensemble methods apply within a single model?
- What hidden costs might fine-tuning retrieval models introduce on out-of-distribution queries?
- How should rapidly evolving domains choose knowledge injection methods?
- How much does workflow architecture matter versus raw model capability?
- What task structures benefit most from geometric parameter merging?
- What alternatives exist when required knowledge is absent from training?
- What access constraints allow description-based adaptation but block conventional techniques?
- How should guidance levels adapt as the model's capability boundary shifts?
- What techniques work best for injecting domain knowledge at training time?
- Can we predict out-of-distribution generalization without access to downstream tasks?
- Why does parameter-efficient tuning scaling fail to improve finetuning performance?
- How do labs actually train next-generation models from previous ones?
- How does joint backpropagation differ from training separate ensemble models?
- Can per-user adapters remain consistent without drifting or leaking?
- How can expensive models efficiently support cheap models in production?
- What counts as full capability recovery versus partial restoration?
- What benefits do open foundation models create that closed systems cannot?
- What distinguishes new associations from existing ones at the computational level?
- How much does pretraining quality affect the modularity of fine-tuned models?
- How do fast skill injection and slow gradient updates work on different timescales?
- Do open model properties like customizability create net new misuse opportunities?
- What distinctive properties make open foundation models different from closed ones?
- Why do medical and mathematical tasks require fundamentally different model capabilities?
- When should full-parameter post-training be used instead of LoRA adaptation?
- Why do production teams choose expensive frontier models over fine-tuning?
- Why do production systems optimize for three model classes instead of foundation models?
- What trade-offs emerge between graph staleness and recommendation freshness?
- How can post-training research become reproducible without releasing full interfaces?
- Why does verification sit on a different scaling axis than pre-training?
- Can different teams train specialist adapters that share one frozen base model?
- How should training distribution distance be defined when the policy evolves?
- How does mixture of experts enable flexible capacity sharing between modalities?
- How similar must a model organism be to its wild case for findings to transfer?
- How do aligned LoRA adapters compose through parameter-space arithmetic?
- Can this approach handle continuously changing product inventories in production?