Is a model's math skill one single ability, or many separate skills that just happen to move together?
Can mathematical capability distributions be read as unified rather than separate?
This explores whether a model's mathematical ability is best understood as one underlying capability or as a bundle of separate skills that happen to show up together in a score. The corpus has no paper that tackles this head-on, but several notes approach it from different directions, and they don't all point the same way.
This explores whether a model's math ability is one thing or several things wearing one number. The corpus has no study aimed squarely at this, but it has a useful tension. Some notes find that what looks like a split is unified underneath. Others find that what looks like one ability is really several pulling apart.
The case for 'unified' comes from a surprising place. In reinforcement learning on math problems, researchers have long assumed a trade-off: a model either explores new solution paths or exploits the ones it already knows. But when you look at the model's internal states instead of its output words, the two barely correlate, and both can be improved at once, which produced large gains on a Chinese university entrance math exam Is the exploration-exploitation trade-off actually fundamental?. The lesson goes beyond RL. Splits that appear at the level of output text can disappear when you measure what's happening inside the model.
The case for 'separate' is just as strong. A study using simple cipher puzzles broke chain-of-thought performance into three independent ingredients. One is how likely the answer is in general. One is memorization of patterns seen in training. The last is real step-by-step reasoning that picks up errors as it goes. Output probability alone moved accuracy from 26% to 70% What three separate factors drive chain-of-thought performance?. So a single 'math score' can blend memory, prior expectations and genuine reasoning in proportions you can't see. Agent evaluation shows a similar pattern: capability splits into axes that rank models differently, so any single score misleads Does a single benchmark score actually predict agent readiness?.
The most unsettling doorway shows that even perfect accuracy tells you little about how a skill is organized inside. Models can contain every feature a task needs while their internal structure stays fractured. That breakage only shows up when inputs shift Can models be smart without organized internal structure?. So 'unified vs. separate' may be a question about the inside of the model, and benchmarks can't answer it. Another note makes a related point: the strong math scores of a small 3B model come from how it was post-trained, not from its size, and those results hold only where answers can be checked Can small models match frontier reasoning without massive scale?. That suggests 'math capability' is partly something the training process builds, not a fixed property of the model.
The takeaway you might not have expected: the answer depends on what you measure. At the level of output words and benchmark scores, math ability looks like separable pieces. In the model's hidden states, some of those apparent splits merge, and seemingly solid skills can turn out fractured. If you meant something narrower, such as how difficulty or accuracy is spread across math benchmarks, the corpus doesn't cover that directly.
Sources 5 notes
Hidden-state analysis using Effective Rank metrics shows near-zero correlation between exploration and exploitation, revealing the trade-off emerges only at token level. VERL demonstrates simultaneous enhancement achieving 21.4% accuracy gains on Gaokao 2024.
A shift cipher study decomposed CoT into three independent factors: output probability alone swings accuracy from 26% to 70%, memorization matches pre-training frequency patterns, and genuine reasoning exists but accumulates error with each step. This resolves the reason-or-memorize debate by showing LLMs do both simultaneously.
Capability decomposes into task success, privacy compliance, long-horizon retention, mode-shift behavior, and ecosystem readiness. Models ranked highest on one axis often rank lower on others, making single-score evaluations systematically misleading for real deployment.
Models trained with SGD can contain all the linearly decodable features needed for a task while maintaining fundamentally broken internal organization. This makes them vulnerable to perturbation and distribution shift invisible to standard evaluation metrics.
A 3B model trained with curriculum SFT and multi-domain RL reaches 94.3 AIME26 and 80.2 LiveCodeBench scores matching much larger systems. The result is bounded to verifiable tasks with checkable ground truth, where RL can provide clean reward signals.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Deciphering the Factors Influencing the Efficacy of Chain-of-Thought: Probability, Memorization, and Noisy Reasoning
- Beyond the Exploration-Exploitation Trade-off: A Hidden State Approach for LLM Reasoning in RLVR
- Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?
- Is Chain-of-Thought Reasoning of LLMs a Mirage? A Data Distribution Lens
- When More is Less: Understanding Chain-of-Thought Length in LLMs
- RL Squeezes, SFT Expands: A Comparative Study of Reasoning LLMs
- VibeThinker-3B: Exploring the Frontier of Verifiable Reasoning in Small Language Models
- Local Coherence or Global Validity? Investigating RLVR Traces in Math Domains