Does AI text that sounds substantial actually carry more ideas per word, or is fluency doing the work?
How does propositional density differ between dense and sparse generated text?
This explores how much actual content (distinct claims or ideas per word) AI-generated text carries, and what separates text that is genuinely packed with meaning from text that only looks substantial.
This explores how much actual content AI-generated text carries per word, and what separates text that is genuinely packed with ideas from text that only looks substantial. To be direct: the collection has no study that counts propositions in dense versus sparse model output. It does come at the question from several other directions, and together they suggest that a model's density and its fluency are two separate things.
The most concrete evidence comes from work that prunes reasoning chains token by token. When researchers strip out the tokens a model can most easily do without, the model clearly ranks words by how much work they do. Symbolic computation (the numbers and operations that carry the answer) is kept to the end, while grammar and meta-discourse such as "let me think about this" or "so now we see" are dropped first Which tokens in reasoning chains actually matter most?. One surprising result: student models trained on these pruned, denser chains do better than students trained on compressions written by frontier models. In practice, the sparse text is mostly scaffolding, and you can measure how much of it is removable. A related finding shows that only about 20% of tokens in a reasoning chain are real "forks" where the model is uncertain and a decision gets made. Training on just those tokens matches training on everything Do high-entropy tokens drive reasoning model improvements?. If most of the learning signal sits in a fifth of the words, most generated text is low-density by construction.
What you see on the surface can also mislead in both directions. Models trained to output filler tokens still compute the correct answer in their early layers, then actively overwrite it to produce the expected format Do transformers hide reasoning before producing filler tokens?. Text that looks empty can hide real work. Diffusion-based models point the other way: answer confidence settles early while the reasoning text keeps being refined, so half the compute can be cut without losing accuracy Can reasoning and answers be generated separately in language models?. Much of that extra reasoning text was not carrying the answer.
The more critical notes explain why sparse text can still feel full. Token prediction moves smoothly toward what is typical in the training data. The result is claims that multiply without opening new perspectives: more sentences, but not more ideas Does LLM generation explore competing claims while producing text?. Models also favor high-frequency phrasings over equivalent rare ones Do language models really understand meaning or just surface frequency?, and widespread reliance on the same models pulls everyone's expression toward the same narrow range Do large language models narrow human expression and thought?. Familiar wording reads as fluent and easy to process, which can pass for substance.
The takeaway you might not have expected: density in generated text may be easier to measure than it sounds. You can delete tokens and check whether the model's confidence or a student's performance holds up. By that test, a lot of fluent AI prose is padding that a model would cut from its own output. Padding can also cost something. Reasoning accuracy drops sharply as inputs get longer, even far below context limits Does reasoning ability actually degrade with longer inputs?, so sparse text can actively crowd out the thinking. If you want the measurement question itself, a study that counts claims per sentence across model styles, the collection doesn't have it yet.
Sources 8 notes
Greedy likelihood-preserving pruning reveals six functional token categories; symbolic computation tokens are preferentially preserved while grammar and meta-discourse are pruned first. Student models trained on these pruned chains outperform those trained on frontier-model compression.
Only ~20% of tokens exhibit high entropy as pivotal reasoning decision points; RLVR primarily adjusts these forking tokens. Training exclusively on them matches or exceeds full-gradient performance, revealing that the minority carries the learning signal.
Logit lens analysis shows models trained with hidden CoT tokens compute correct answers in layers 1-3, then actively suppress these representations in final layers to produce format-compliant filler output. The reasoning is fully recoverable from lower-ranked token predictions.
ICE shows that bidirectional attention in diffusion LLMs enables in-place prompting—embedding reasoning directly in masked positions refined alongside answers. Answer confidence converges early while reasoning continues refining, allowing early-exit mechanisms to cut compute by 50% while maintaining accuracy.
Token prediction trains models to continue toward the training distribution, not to explore logically related counterpositions. This smoothness in process produces smooth claims that multiply without generating new perspectives.
Show all 8 sources
LLMs show consistent preference for higher-frequency surface forms over semantically equivalent rare paraphrases across math, machine translation, commonsense reasoning, and tool calling. This suggests models track statistical mass from pretraining rather than meaning-recognition as their primary mechanism.
LLMs mirror skewed slices of human experience shaped by training data regularities, and widespread reliance on identical models amplifies convergence. Co-writing studies show users unconsciously adopt model stances and framings.
FLenQA shows reasoning accuracy drops from 92% to 68% at just 3000 tokens of padding, far below context window capacity. The degradation is task-agnostic, uncorrelated with language modeling performance, and persists even with chain-of-thought prompting.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Farther the Shift, Sparser the Representation: Analyzing OOD Mechanisms in LLMs
- Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning
- Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey
- Argument Collapse: LLMs Flatten Long-Form Public Debate
- Six misconceptions about large language models: A minimal model and diagnostic taxonomy
- On the Reasoning Capacity of AI Models and How to Quantify It
- Beyond Semantics: The Unreasonable Effectiveness of Reasonless Intermediate Tokens
- The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem Complexity