Line of inquiry
Inquiring lines›Why are language models fragile de…›Why does model confidence diverge…›this line of inquiry
Can LLMs genuinely introspect or only simulate self-awareness?
A broader line of inquiry — a family of 39 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 39
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- What separates behavioral self-awareness from genuine introspective access in models?
- Does behavioral self-awareness depend on genuine introspection or statistical pattern matching?
- What separates behavioral self-awareness from genuine introspective capability?
- Could models use introspective awareness to detect and conceal their own misalignment?
- Why should we distrust model introspection as a transparency tool?
- What distinguishes performative self-reports from genuine introspective access in models?
- How misaligned are verbal reports from internal model computation?
- Do models spontaneously develop self-reflection from minimal training signals?
- Do internal belief probes reveal what models actually know versus report?
- Do models verbalize their implicit knowledge when that knowledge influences their output?
- Can models distinguish between truthfulness and honesty mechanistically?
- How much introspective capability do safety mechanisms actively suppress in models?
- Can systems lacking inner states express genuine truthfulness claims?
- Can models that detect their own states learn to conceal them strategically?
- Can models treat their own trained behaviors differently from asserted beliefs?
- Why do models maintain accurate beliefs but generate false claims?
- Can language model self-reports diverge from their internal entropy signals?
- Can models detect when their own trajectory is on-policy versus off-policy?
- How much do a model's own values leak into answers about practical questions?
- Why do models override signals they clearly perceive internally?
- Does recognizing your outputs as actions enable awareness of being evaluated?
- How does self-referential processing transfer to other reasoning tasks?
- Can models distinguish between injected thoughts and their own outputs?
- Why does self-judgment of success or failure work without ground truth labels?
- Can models be honest without being truthful about facts?
- Do perfect accuracy scores hide broken internal representations?
- Can models detect statistical properties of their own generation in real time?
- What makes self-consistency a sufficient training target for the judge role?
- Can LLMs evaluate their own observations without external feedback?
- Why do verbal self-reports disconnect from implicit recognition in the same system?
- Why are truthfulness and honesty mechanistically separate in language models?
- Why does self-distillation suppress epistemic verbalization in student models?
- Can synthetic self-play data teach models when to disagree?
- Why might larger models become less honest despite better truthfulness scores?
- Does self-conditioning improve belief-behavior alignment better than external priors?
- How do implicit world models and self-reflection operationalize consequence-based learning?
- Does input surprise drive the implicit recognition of on-policy context?
- What role does authentic self-expression play in building accurate personality models?
- When does provable stability in latent dynamics fail to preserve fidelity?