INQUIRING LINE

Is AI getting smarter, or just getting better tools to use the same brain?

How much of AI improvement comes from tools versus model capability?

This explores how much of the recent progress in AI comes from what is built around a model (tools, harnesses, interfaces) versus what is inside the model itself, and whether the two can be separated at all.


This explores how much of AI's recent progress comes from the systems wrapped around a model (tools, execution harnesses, interfaces) rather than from the model itself. The short answer is that the corpus has no clean percentage split. What it does show is something more useful: the line between 'tool' and 'capability' is blurrier than the question assumes, and a good share of what looks like model progress turns out to be scaffolding.

Start with the strongest case for tools. Giving a model tools isn't just a convenience. There is a formal argument that it strictly widens what the model can reason about, because it makes some strategies possible that would be impossible or impractically long in plain text, and this holds for abstract reasoning, not just arithmetic Do tools actually expand what language models can reason about?. In practice, improving only the execution system around frozen weights raises scores on a terminal-based benchmark across several models, and the same recipe carries over to newer models unchanged Can execution harnesses lift model performance without retuning weights?. When a self-improving agent was left to rewrite itself, much of what it found was better code editing and context management, which is harness work rather than smarter weights Can AI systems improve themselves through trial and error?. Ethan Mollick goes further and argues that the gap between what models can do and what people actually get from them is mostly an interface problem. In one study, finance professionals lost part of their GPT-4 productivity gain to the mental overhead of the chat window Is the AI capability gap really an interface problem?.

The surprise is that tools and capability don't simply add together. They interact. When models were asked to improve their own harnesses, the quality of their suggestions stayed about the same across model tiers, but how much a model *benefited* from those improvements peaked in the middle tier. Weak models failed to use the harness at all, and the strongest models were worse at faithfully following its instructions Do stronger models always evolve harnesses better?. So 'add better tools' pays off unevenly, depending on where the model sits. A related finding: a 35B model trained on execution-heavy, long-task data competes with much larger models at far lower cost. That suggests training a model for coordination and follow-through can stand in for raw size Does model efficiency matter more than peak capability for real work?.

Even 'model capability' gains may be less about new ability than they look. One line of evidence argues that RL post-training mostly teaches a model *when* to use reasoning it already had in latent form, not *how* to reason. Routing tokens alone recovers most of the gains Does RL post-training create reasoning or just deploy it?. Scaling is uneven too. Style-like skills stop improving early, while logical reasoning and knowledge keep climbing with model size Do all AI skills improve equally as models scale?. Put together, the real ceiling may be set by deep reasoning and knowledge, while tools, training, and interfaces decide how much of that ceiling you actually reach.

One more twist: the credit-assignment problem applies to people as well. Studies of AI-assisted work describe how fluent output and hidden processing lead users to count AI contributions as their own skill How do AI tools trick users into overestimating their own skills?. Asking 'was it the tool or the model?' runs into the same trap as asking 'was it me or the AI?'. Once the parts are tightly coupled, the gain belongs to the whole system, and pulling the contributions apart is a measurement problem in its own right.


Sources 9 notes

Do tools actually expand what language models can reason about?

Formal proof shows tool-integrated reasoning enables strategies impossible or prohibitively verbose in text alone, expanding both empirical and feasible support. The advantage spans abstract reasoning, not just arithmetic, and Advantage Shaping Policy Optimization stabilizes training without reward distortion.

Can execution harnesses lift model performance without retuning weights?

StateM improves Terminal-Bench 2.1 accuracy across multiple models by optimizing execution systems around frozen weights. The same runbook transfers to newer models without modification, achieving 95.3% on GPT-5.6 and lifting DeepSeek-V4 Flash by 5.4 points.

Can AI systems improve themselves through trial and error?

DGM replaces formal proofs with empirical benchmarking and maintains an evolutionary archive of agent variants, achieving 2.5× improvement on SWE-bench and 2.2× on Polyglot by discovering capabilities like better code editing and context management.

Is the AI capability gap really an interface problem?

Mollick argues that better interfaces—not better models—will drive perceived capability leaps. Evidence includes a cognitive-load study showing financial professionals gained productivity from GPT-4 but lost it to chatbot design's cognitive overhead, especially hurting less experienced users.

Do stronger models always evolve harnesses better?

Model capability to produce useful harness edits stays constant across tiers, but capacity to actually benefit from those edits follows an inverted U-shape, peaking in mid-tier models. Weak models fail to invoke harnesses; strong models struggle with faithful instruction-following.

Show all 9 sources
Does model efficiency matter more than peak capability for real work?

Occamy-1.0, a 35B-parameter model further trained on execution-grounded data and long-horizon trajectories, achieves competitive performance with much larger models while sitting at the low-cost knee of the Pareto frontier, suggesting that training for coordination and follow-through substitutes for raw scale in multi-step work.

Does RL post-training create reasoning or just deploy it?

Evidence shows base models already contain reasoning capability in latent form; RL training optimizes deployment timing rather than capability creation. Hybrid models recover 91% of performance gains by routing tokens only, and activation vectors for reasoning strategies pre-exist before any RL.

Do all AI skills improve equally as models scale?

FLASK's 12-skill decomposition reveals metacognition saturates at 7B parameters while logical efficiency plateaus at 30B, but reasoning and knowledge skills improve continuously. Open-source models successfully imitate surface-level style but fail at reasoning—confirming that distillation copies form not substance.

How do AI tools trick users into overestimating their own skills?

Attribution ambiguity, fluency illusion, cognitive outsourcing, and pipeline opacity combine to systematically misattribute AI outputs as user competence. The effect is multiplicative—each mechanism amplifies the others.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.