Line of inquiry
Inquiring lines›What drives capability improvement…›What drives capability improvement…›this line of inquiry
Can smaller specialized models match frontier models on key metrics?
A broader line of inquiry — a family of 74 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 74
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Does small heterogeneous model architecture outperform large homogeneous pools economically?
- Can smaller models produce skill updates as useful as frontier model updates?
- Do small models show different parameter efficiency patterns than large models?
- Can externalized memory and skills replace model scaling?
- What performance trade-offs emerge when composing multiple independently trained model capabilities?
- Why do fine-tuned models fail outside their specialized domains?
- What capability risks emerge when models are optimized for single domains?
- What distinguishes domain-specific failure modes from general model limitations?
- Which model capabilities actually matter for sustained workflow delegation?
- Does the optimal model size depend on what capabilities you actually need?
- Why does over-specialization create a domain capability cliff in LLMs?
- Does model collapse occur across different architectures or only in specific conditions?
- Do integrated and decoupled architectures trade off intervention accuracy for efficiency differently?
- What causes models to develop domain capability cliffs after specialization?
- Does model capability still matter once coordination infrastructure is optimized?
- How does over-specialization create capability cliffs outside target domains?
- Why do frontier models corrupt more documents than weaker models during workflows?
- How much do different LLMs independently converge on similar outputs?
- What hidden costs emerge when you fine-tune models for a single domain?
- Why does capability saturation and diversity saturation occur at different scales?
- Can models optimized for solo capability support productive human collaboration?
- Why do metric choices constrain which model capabilities get developed?
- Can structural diversity through role assignment replace emergent diversity in small models?
- What production constraints should determine paradigm selection?
- How does workflow scale change the failure modes of frontier models?
- Why do only two of fourteen models improve when problem constraints are removed?
- How much does workflow architecture matter versus raw model capability?
- How do larger models maintain more parallel tasks than smaller models?
- Do different domains require different types of model investment?
- Can architectural changes reduce representational inequality in unified generators?
- Is model selection a stronger security lever than improving individual model defenses?
- Can specialized components replace single fully-trained models in deployment?
- How much does workflow architecture matter compared to raw model capability in forecasting?
- Why do high-level design guidelines fail to capture real-world deployment nuance?
- Why do frontier model failures in document editing go undetected by users?
- How do ensemble methods apply within a single model?
- Can depth scaling and breadth scaling unlock independent capability axes?
- Why do most frontier models terminate early on long-horizon benchmarks?
- Why do scaling laws fail to predict optimal architectures at small parameter counts?
- Which failure modes dominate when models handle underspecified requests?
- How do you identify which models should form a minimal diverse coreset?
- What filtering criteria best identify student-compatible refinements from teacher models?
- Why do frontier models corrupt documents while weaker models delete them?
- Can architectural changes reorder when uncertainty and empowerment signals influence decisions?
- What distinctive properties make open foundation models different from closed ones?
- Which architectural choices matter most when a model must fit one billion parameters?
- Can review effort alone keep pace with frontier model degradation?
- How do virtual model instances preserve identity through load-balancing and failover?
- Can per-user adapters remain consistent without drifting or leaking?
- Can end-to-end models maintain debuggability without modular components?
- Why do production systems optimize for three model classes instead of foundation models?
- Why do production teams choose expensive frontier models over fine-tuning?
- Do open model properties like customizability create net new misuse opportunities?
- Why does restricting foreign access require halting domestic model availability?
- How does the Ladder of Scales approach reduce search costs across model sizes?
- What access constraints allow description-based adaptation but block conventional techniques?
- What benefits do open foundation models create that closed systems cannot?
- How does model tier affect whether errors delete or corrupt document content?
- What mobile hardware constraints force the sub-billion parameter regime?
- What constraints force mobile deployments to operate in the sub-billion parameter regime?
- How do trait adapters interact with different base model architectures?
- What are the five inseparable design choices when building world models?
- What makes diverse failure modes more informative than single failure examples?
- Are SchemeArena's scenario factors fully crossed to separate bundled changes?
- Could deploying GPT-4 for everyone require 100 million specialized chips?
- How similar must a model organism be to its wild case for findings to transfer?
- Can different teams train specialist adapters that share one frozen base model?
- Why does depth outperform width for sub-billion parameter models?
- How does single-pass generation differ from multi-stage synthesis architecturally?
- How do stepping stone solutions transfer effectively between different environments?
- What organizational bottlenecks emerge when expertise concentrates in few specialists?
- What three independent failure points bottleneck traditional function calling systems?
- How do aligned LoRA adapters compose through parameter-space arithmetic?
- What role do CM fields play in constructing high unit distance configurations?