Line of inquiry
Inquiring lines›How do training signals reliably a…›What training signals and data cur…›this line of inquiry
What determines how much models can improve capabilities through specialized training and composition?
A broader line of inquiry — a family of 50 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 50
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- How much can externalized skills improve models before hitting diminishing returns?
- Does specialized training in one domain create capability cliffs elsewhere?
- Does curriculum-based training keep small models perpetually at their learning edge?
- How do self-evolving curricula help RL break beyond base model capability boundaries?
- What performance trade-offs emerge when composing multiple independently trained model capabilities?
- Can smaller models produce skill updates as useful as frontier model updates?
- Do emergent abilities result from genuine new capabilities or implicit in-context learning?
- What skills can large models identify and organize about their own abilities?
- Do text-space skills transfer learning across different frontier models?
- Can granular sub-task training for function calling improve both open and proprietary models?
- Can model training address failures that really originate in harness gaps?
- Can expert vectors learned offline transfer across multiple model architectures?
- Does task superposition explain how models learn from multiple in-context trajectories?
- Can external managers optimize context better than the model itself?
- Can models optimized for solo capability support productive human collaboration?
- Can extracted skills transfer effectively across different domains and model architectures?
- Which model capabilities actually matter for sustained workflow delegation?
- Can models adapt and combine search strategies beyond their training algorithm?
- What structural differences emerge between early generic skills and later meta-strategy skills?
- Where does skill extraction fail compared to genuine model adaptation?
- Why do metric choices constrain which model capabilities get developed?
- Does importance sampling actually recover capabilities lost to hard sample training?
- Does the pretrained prior actually constrain what internalized search can discover?
- Can capability boundary collapse be reversed through external data?
- Do different function-calling subtasks have different entropy profiles during training?
- When and what should a model actually decide to delegate?
- How should multi-objective post-training balance competing behavioral goals?
- What makes skills worth externalizing into a persistent harness?
- Why does delegation training help models that work alone?
- Why does externalizing bookkeeping raise effective feedback compute?
- Can text-space optimization and audit governance coexist in a single skill lifecycle?
- What emergent behaviors do models develop when trained on underspecified pedagogical tasks?
- Can explicit goal state scaffolding at inference time transfer to autonomous tracking through training?
- How do ensemble methods apply within a single model?
- How do external invocation latencies drive technique convergence?
- Can curated demonstrations compensate for smaller or simpler training environments?
- Can models generate their own training curriculum during offline dreaming?
- Can specialized components replace single fully-trained models in deployment?
- Can contextual design decisions resist formalization into evaluation rubrics?
- What makes a model fail to activate relevant skills from its own harness?
- How does student capacity limit what it can learn from teachers?
- What makes skills suitable for retrieval and chaining in repositories?
- How do composite rewards attribute curation outcomes to specific skill library changes?
- How does trajectory burstiness compare to other structural properties that shape emergent capabilities?
- What access constraints allow description-based adaptation but block conventional techniques?
- Does extended exoskeleton use eventually produce meaningful skill transfer?
- How does scaffolding unstable mechanics improve reinforcement learning for search?
- How can post-training research become reproducible without releasing full interfaces?
- What makes exploration a verifiable and measurable training objective?
- Can backward transfer measurements reliably predict optimal multi-task training order?