INQUIRING LINE

Why do AI planners struggle to think big-picture when every level of their plan shares the same mental sketchpad?

Why does a shared latent space fail to represent abstract goals well?

This explores why AI planners that use one internal representation for every level of planning, from the next small step up to the long-range goal, have trouble with abstract goals, and what changes when each level gets its own representation.


This explores why a planning system does worse at abstract goals when every level of its hierarchy works in the same internal representation (the same 'latent space'), and what the corpus says about fixing that. The clearest evidence comes from H-JEPA, a hierarchical world model. When each level of its planning hierarchy got its own, more abstract latent space instead of sharing one, success on a maze-navigation benchmark rose from 18% to 73% Does each hierarchy level need its own latent space?. The explanation is about what each level needs to see. A low-level controller needs fine detail: where the legs are and what the next few frames look like. A high-level planner asks a different question: does this future get me closer to the goal? If both share one space, the high level has to judge candidate futures in a representation full of detail that doesn't matter for that judgment. A separate, coarser space lets it compare futures in terms that match what a goal is.

A formal result helps explain why abstraction should be learned level by level. One analysis proves that predicting your own latents, rather than raw tokens or pixels, recovers layered structure with a number of samples that stays constant as the hierarchy gets deeper. Token-level learning needs exponentially more samples Why is predicting latents more sample-efficient than tokens?. The reason is that latents at the same level are strongly correlated with each other, while raw tokens are noisy. The takeaway is that structure lives at particular levels of abstraction. Collapsing several levels into one space throws away the thing that makes higher-level structure easy to learn.

The same idea shows up in language models under different names. Meta's Large Concept Model reasons over whole-sentence embeddings and plans at the paragraph level, instead of working token by token, and it produces more coherent long outputs Can reasoning happen at the sentence level instead of tokens?. User simulators for recommendation systems split their hidden variables the same way: one for who the user is across a whole session and another for what they want in each turn Can controlled latent variables make LLM user simulators realistic?. A long-range intention and a moment-to-moment action are different kinds of thing, and systems tend to work better when they're represented separately.

There's also a warning: a separate abstract space only helps if it keeps its shape. Research on latent chain-of-thought finds that when the only training signal is whether the final answer was right, the latent space drifts away from anything meaningful. The fix is supervision that preserves the space's geometry instead of compressing it Why does latent chain-of-thought fail so easily in training?. The reverse failure also helps explain the planning problem. Reasoning models that wander without a systematic search see their success rate drop exponentially as problems get deeper Why do reasoning LLMs fail at deeper problem solving?. That is roughly what you'd expect from a planner with no good representation of where it's trying to go.

A caveat: the corpus has one direct experiment on shared versus separate latent spaces (H-JEPA). The rest is supporting theory and parallels from other fields. Together they suggest that abstraction isn't just a compressed version of detail. It has its own structure and needs its own space to be represented.


Sources 6 notes

Does each hierarchy level need its own latent space?

H-JEPA improves Visual AntMaze planning from 18% to 73% success by giving each hierarchy level a distinct, more abstract latent space rather than sharing one. This per-level abstraction lets higher levels score candidate futures in abstract space better matched to goal-like objectives.

Why is predicting latents more sample-efficient than tokens?

A formal sample-complexity analysis proves latent-level self-supervision (data2vec/JEPA style) recovers compositional structure with samples constant in hierarchy depth, while token-level learning requires exponential samples—because same-level latents are far more correlated than raw tokens.

Can reasoning happen at the sentence level instead of tokens?

Meta's Large Concept Model operates on sentence embeddings rather than tokens, reasoning in a language-agnostic space before decoding to any target language. This hierarchical approach with paragraph-level planning produces more coherent output than flat token generation.

Can controlled latent variables make LLM user simulators realistic?

RecLLM demonstrates that conditioning an LLM simulator on session-level (user profile) and turn-level (user intent) latent variables produces synthetic conversations measurable as realistic via crowdsource discrimination, discriminator models, and classifier-ensemble distribution matching.

Why does latent chain-of-thought fail so easily in training?

Outcome supervision alone causes gradient attenuation along latent steps and lets the latent space wander without semantic grounding. Robust latent reasoning requires both dense trajectory supervision and space supervision that preserves geometric structure rather than compressing it.

Show all 6 sources
Why do reasoning LLMs fail at deeper problem solving?

Current reasoning models lack the three properties of systematic exploration: validity, effectiveness, and necessity. This causes success probability to drop exponentially with problem depth, making medium problems solvable but deep problems catastrophically harder.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.