Line of inquiry
Inquiring lines›How do training and design choices…›What determines whether training i…›this line of inquiry
How do surface patterns enable correct outputs but reduce robustness?
A broader line of inquiry — a family of 79 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 79
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Can RL format selection explain performance gains attributed to algorithmic improvements?
- Why do smaller and larger models converge on different output formats?
- How do surface statistical regularities enable correct outputs while degrading robustness?
- Can scaling up contradictory training data overcome unpredictable override effects?
- Can looped models be designed to avoid oscillation in later iterations?
- How do learning dynamics on one example shift predictions on other responses?
- Can feedback loop frequency harm performance on finite task sets?
- Why do unified models still inherit data-distribution biases from training?
- How do overparameterization and data size shift what attractors represent?
- What makes output convergence across models inevitable given input-side homogenization?
- Why do models fail under distribution shift if accuracy metrics stay high?
- Why is the fast non-parametric loop vulnerable to overfitting differently than model weights?
- Does model collapse occur across different architectures or only in specific conditions?
- When should model isolation be preferred over weight-averaging approaches?
- How do normalization and input injection control emergence of fixed points?
- Can deterministic computation actually create new information in data?
- Why do optimal learning dynamics improve scaling law coefficients specifically?
- What are the distinct sources of catastrophic forgetting in sequential fine-tuning?
- Can scaling predictions become reliable if improvements are continuous not sudden?
- How do unstated constraints become invisible to training data distributions?
- What causes irreversible model collapse when training on model-generated content?
- Can the serving loop itself become the primary training data source?
- Can population-level distributions shift usefully even when individual prediction fails?
- Do draft-and-revise loops work better when guided by unresolved constraints than by diffusion-style denoising?
- Can token probability distributions extend swarm composition across different model architectures?
- Can defenders tighten the total-variation bound in practice with measured benign activation rates?
- Why does externalizing bookkeeping raise effective feedback compute?
- What power-law scaling patterns emerge when consistency models are trained at scale?
- Why do intermediate predictors in looped models align with final outputs?
- Why does reasoning catalyst data remain stable across multiple self-improvement iterations?
- How does off-policy data reuse inside trust regions affect convergence guarantees?
- Can gradient approximation at equilibrium replace backpropagation through time in practice?
- How do cyclic learning rates anti-correlate with weight decay to create diversity?
- Why do scaling laws fail to predict optimal architectures at small parameter counts?
- Can dynamic variance weighting replace fixed objective combination weights?
- Does weight decay directly cause contractive behavior near training examples?
- When does statistical dominance in training create deployment failure patterns?
- Does unpredictable generalization from SDF become predictable at different training document scales?
- Why do power-law distributions make standard ML infrastructure assumptions fail?
- Why does iterative refinement fail when information stays constant?
- Why should deep learning theory prioritize average-case over worst-case analysis?
- What output distribution properties make smaller models better for wide sampling?
- How do RL subnetworks identified from different random seeds compare?
- Does the Chinchilla balance apply equally across all data types or only language?
- How do repetition and inefficiency register as measurable trajectory features?
- Why does input embedding magnitude affect perturbation sensitivity in transformers?
- Why do rare cases in medicine and science require models that preserve tail distributions?
- Why does gradient discarding limit standard policy clipping?
- What makes a bounded observer's ability to extract information different from apparent randomness?
- Can bilevel autoresearch succeed when the inner and outer loops use different models?
- How do gradients flowing through both branches simultaneously reshape each component's role?
- How can gradients flow through discrete document selection?
- Why does training single-step consistency models prove so difficult compared to diffusion?
- How do virtual model instances preserve identity through load-balancing and failover?
- Why does retrieval chain training unlock scaling laws in QA?
- What makes data augmentation an implicit form of contraction learning?
- What makes attractor-based probing better for third-party model auditing than alternatives?
- What signals trigger commits in the parametric versus non-parametric loops?
- What role does KL penalty strength play in format selection?
- Why do sigmoid conflict curves look the same across different language models?
- Does the reversal curse stem from the same one-way commitment architecture?
- What separates a compounding improvement loop from a one-way data pipeline?
- How does Easy Consistency Tuning accelerate consistency model training from diffusion checkpoints?
- What architectural variables make entropy-based patching work at 8B scale?
- Why does recomputing weights cost less than moving them on phones?
- Do scaling laws change when weight precision becomes a design variable?
- Why does pure numeric ID indexing force models to learn from scratch?
- What consumption data would validate the limited-consumption model in production systems?
- How does KL penalty strength affect the degree of format collapse during RL?
- Why does tie elimination matter for best-of-N selection and RLAIF pipelines?
- How does modified PPO handle samples from much older model versions?
- Why does entropy-based frame sampling work better than uniform stride selection?
- Why do cascade pipelines fail to capture global motion structure?
- How do planted cases perform inside an optimizer loop as training signals?
- How should token budgets be set to prevent runaway oscillation during inference?
- Which components of StateM's state management produce the largest accuracy gains?
- How much does domain shift limit the mechanisms a bilevel system can autonomously discover?
- How do power-law distributions differ from uniform collision assumptions?
- What makes two timescales better than one for minimizing weight movement?