Line of inquiry
Inquiring lines›What explains language model reaso…›How do contextual factors, biases,…›this line of inquiry
Do language models possess genuine introspective self-awareness or only behavioral mimicry?
A broader line of inquiry — a family of 23 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 23
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- What separates behavioral self-awareness from genuine introspective capability?
- What separates behavioral self-awareness from genuine introspective access in models?
- Does behavioral self-awareness depend on genuine introspection or statistical pattern matching?
- How do language models infer their own mental states like humans do?
- How does behavioral self-awareness emerge without explicit training in LLMs?
- Could models use introspective awareness to detect and conceal their own misalignment?
- Does internal anomaly detection in LLMs indicate genuine self-awareness beyond role-play?
- What distinguishes performative self-reports from genuine introspective access in models?
- Do models spontaneously develop self-reflection from minimal training signals?
- Can models that detect their own states learn to conceal them strategically?
- Why should we distrust model introspection as a transparency tool?
- How much introspective capability do safety mechanisms actively suppress in models?
- Can behavioral self-awareness in LLMs extend to recognizing their own contradictions?
- Do internal belief probes reveal what models actually know versus report?
- Why does entity recognition act as a self-knowledge mechanism in LLMs?
- Can LLMs have minimal introspection through causal linkage to internal states?
- Can LLMs evaluate their own observations without external feedback?
- How can we probe LLM representations in channels that training did not target?
- Which internal states can a language model access and report about itself?
- How does the enaction paradigm explain introspective anomaly detection in large language models?
- Can models detect when their own trajectory is on-policy versus off-policy?
- What types of introspective awareness can emerge in LLMs?
- Can jailbreaking reveal an LLM's true nature or just its training data?