INQUIRING LINE

AI systems are now running their own research loops — how do we tell a real discovery from a trick that just games one test?

How do narrow benchmark optimizations differ from genuine architectural research discoveries?

This explores how to tell the difference between a change that just raises a score on one test and a finding that actually shows something new about how AI systems should be built, especially now that AI systems run their own research loops.


This explores how to tell a tuning trick that raises one leaderboard number apart from a discovery that changes how AI systems should be built. The corpus doesn't offer a single definition. Read together, though, it suggests three tests: does the gain transfer, does it explain why, and does it sit where you think it does?

**Transfer is the first test.** A narrow optimization fits the test it was tuned on. A discovery keeps working on problems nobody tuned it for. The automated ML engineering system AIDE2 is a useful model of how to check this. Its authors evaluated it on four benchmarks held out from the selection process, including physics-based weather forecasting, a domain outside anything it was selected on Do AIDE2's improvements transfer to unseen tasks?. The Darwin Gödel Machine is a similar case. It improved itself on coding benchmarks, but the changes it found were general ones, such as better code editing and context management, rather than benchmark-specific hacks Can AI systems improve themselves through trial and error?. Gains that survive a change of setting are more likely to be real.

**Explanation is the second test.** Real architectural findings usually come with a mechanism. MobileLLM's result that deep, thin networks beat wide ones at small scale contradicts earlier scaling assumptions, and it comes with a reason: layers stack simple concepts into more abstract ones Does depth matter more than width for tiny language models?. In recommender systems, simpler models with the right built-in assumptions beat deeper ones. That tells you design choices matter more than raw capacity What architectural choices actually improve recommender system performance?. The strongest example explains a ceiling. Token-by-token generators struggle with constraint puzzles because they can't take back a token once it's written, and that's why plugging in a symbolic solver works Why does autoregressive generation fail at constraint satisfaction?. A score increase alone can't tell you any of this.

**Location is the third test, and the one most people miss.** Many gains that look like model or architecture progress actually come from the scaffolding around the model, often called the harness. StateM raised Terminal-Bench scores across several models without touching their weights, and the same playbook carried over to newer models Can execution harnesses lift model performance without retuning weights?. Reorganizing a code repository around runtime behavior let weaker planners match stronger ones Can explicit behavior maps help weaker planners compete with stronger models?. These are real improvements, but they are discoveries about the system around the model, not about model architecture. Similarly, reasoning models keep their advantage over standard models however much inference compute the standard models get, because the advantage comes from training rather than budget Can non-reasoning models catch up with more compute?. Yet extended thinking doesn't help on numerical optimization tasks, where the bottleneck is numeric procedure rather than reasoning Do reasoning models actually beat standard models on optimization?. Figuring out where a gain lives is part of figuring out what it means.

The twist is automated research. ASI-ARCH ran 1,773 autonomous experiments, found 106 state-of-the-art architectures, and reported that breakthroughs scale predictably with GPU compute Can computational power accelerate scientific discovery itself?. But autonomous research only works in fields with an instant numeric score, modular parts and fast iteration What makes a research domain suitable for autonomous optimization?. That is also the setting where narrow benchmark optimization thrives. So the engine that makes discovery scale is built to climb metrics. Whether its output counts as discovery depends on checks the loop doesn't run by itself: held-out transfer and a mechanistic explanation. The corpus has little that directly audits whether ASI-ARCH's architectures pass those checks. That's an open gap worth watching.


Sources 11 notes

Do AIDE2's improvements transfer to unseen tasks?

The paper reports that AIDE2's improvements transfer to four held-out benchmarks spanning machine learning, algorithm engineering, and physics-based weather forecasting—the last being outside the selection task distribution. This demonstrates transferable gains beyond overfitting to the selection set.

Can AI systems improve themselves through trial and error?

DGM replaces formal proofs with empirical benchmarking and maintains an evolutionary archive of agent variants, achieving 2.5× improvement on SWE-bench and 2.2× on Polyglot by discovering capabilities like better code editing and context management.

Does depth matter more than width for tiny language models?

MobileLLM shows deep-and-thin architectures yield 2.7–4.3% accuracy gains over balanced designs at 125M–350M scale by composing abstract concepts through layers rather than spreading parameters across width.

What architectural choices actually improve recommender system performance?

Research shows that architectural choices like removing hidden layers, enforcing constraints on self-similarity, and using appropriate likelihood functions deliver better results than deeper or more complex models. This suggests that problem-specific design decisions matter more than raw representational capacity.

Why does autoregressive generation fail at constraint satisfaction?

The performance ceiling on constraint satisfaction problems is not a model-quality issue but an architectural limitation: autoregressive transformers cannot retract emitted tokens, while CSP solvers fundamentally depend on discarding invalid partial assignments. Symbolic solver integration works because it supplies what the architecture lacks.

Show all 11 sources
Can execution harnesses lift model performance without retuning weights?

StateM improves Terminal-Bench 2.1 accuracy across multiple models by optimizing execution systems around frozen weights. The same runbook transfers to newer models without modification, achieving 95.3% on GPT-5.6 and lifting DeepSeek-V4 Flash by 5.4 points.

Can explicit behavior maps help weaker planners compete with stronger models?

A behavior-to-code mapping representation improved win rates by 10–19 points while reducing planner tokens by 8–13%. Weaker planners using this mapping matched stronger models' code localization across all precision and recall metrics.

Can non-reasoning models catch up with more compute?

Reasoning models persistently outperform non-reasoning models regardless of inference budget because training instills a reasoning protocol that makes additional tokens productive. The gap is fundamentally about deployment mechanisms and training structure, not raw capability.

Do reasoning models actually beat standard models on optimization?

Reasoning variants with extended CoT show no consistent advantage over standard models on constraint-bound numerical tasks like optimal power flow. Extended thinking produces more text, not more iterative computation, suggesting the bottleneck is numeric procedure rather than reasoning steps.

Can computational power accelerate scientific discovery itself?

ASI-ARCH discovered 106 state-of-the-art architectures through 1,773 autonomous experiments, revealing that architectural breakthroughs scale predictably with GPU compute. This transforms research from human-limited to computation-scalable.

What makes a research domain suitable for autonomous optimization?

Autonomous research pipelines require immediate scalar metrics, modular architecture, fast iteration cycles, and version control. Domains lacking any property resist autoresearch regardless of LLM capability, because the bottleneck is environmental structure, not model power.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.