Line of inquiry
Inquiring lines›Where does language-model reasonin…›How do reward models guide reliabl…›this line of inquiry
Is model self-awareness based on genuine introspection or pattern matching?
A broader line of inquiry — a family of 47 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 47
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- What separates behavioral self-awareness from genuine introspective access in models?
- What separates behavioral self-awareness from genuine introspective capability?
- Does behavioral self-awareness depend on genuine introspection or statistical pattern matching?
- Could models use introspective awareness to detect and conceal their own misalignment?
- What distinguishes performative self-reports from genuine introspective access in models?
- Do internal belief probes reveal what models actually know versus report?
- Can self-description of internal states influence consciousness attribution?
- Why should we distrust model introspection as a transparency tool?
- Can systems lacking inner states express genuine truthfulness claims?
- How much introspective capability do safety mechanisms actively suppress in models?
- How do language models infer their own mental states like humans do?
- How misaligned are verbal reports from internal model computation?
- Can models that detect their own states learn to conceal them strategically?
- Does internal anomaly detection in LLMs indicate genuine self-awareness beyond role-play?
- Can models distinguish between truthfulness and honesty mechanistically?
- How does behavioral self-awareness emerge without explicit training in LLMs?
- Do models verbalize their implicit knowledge when that knowledge influences their output?
- Does recognizing your outputs as actions enable awareness of being evaluated?
- What makes quasi-beliefs real enough to explain AI behavior?
- Can models distinguish between injected thoughts and their own outputs?
- Can models detect when their own trajectory is on-policy versus off-policy?
- Why does entity recognition act as a self-knowledge mechanism in LLMs?
- Why do models override signals they clearly perceive internally?
- Can language model self-reports diverge from their internal entropy signals?
- Can behavioral self-awareness in LLMs extend to recognizing their own contradictions?
- Can a perfect behavioral simulation constitute genuine understanding or experience?
- Why do users attribute consciousness to language models in practice?
- How do neural self-other representations affect AI deception and alignment?
- What skills can large models identify and organize about their own abilities?
- Why do verbal self-reports disconnect from implicit recognition in the same system?
- Can models develop situational awareness without explicit training for it?
- How does the enaction paradigm explain introspective anomaly detection in large language models?
- Can models be honest without being truthful about facts?
- Why are truthfulness and honesty mechanistically separate in language models?
- Can observation transparency make models more honest in reasoning?
- Does input surprise drive the implicit recognition of on-policy context?
- Can LLMs have minimal introspection through causal linkage to internal states?
- What is the behavioral signature of a model tracking input surprise?
- Can functional behavior alone capture what makes something a genuine belief?
- Can lie detection work from just honesty representation vectors?
- Why does conceptual priming alone fail to produce consciousness claims?
- What role does authentic self-expression play in building accurate personality models?
- What types of introspective awareness can emerge in LLMs?
- What distribution patterns appear across different theory-of-mind datasets?
- What makes accountability and validity-orientation non-behavioral properties?
- What are the seven components of genuine mental state simulation?
- Can jailbreaking reveal an LLM's true nature or just its training data?