Theme of inquiry
How do representation and aggregation choices affect model alignment and reliability?
A question within its area, explored through 5 lines of inquiry below — each a family of specific questions the research asks.
51 specific questions
- Do verbal alignment benchmarks measure representation or just output compliance?
- Can alignment training conceal underlying model associations from probes?
- What alignment properties emerge when the reward model disappears?
- Does alignment training create the shared prosocial pattern across models?
- Can alignment-aware training deposit knowledge where reasoning can access it?
- Why does RLHF alignment reduce the diversity of viewpoints in AI output?
- How do early training associations survive later alignment attempts?
20 specific questions
- How do human annotators disagree systematically on ambiguous examples?
- Why do NLP benchmarks treat annotation disagreement as noise rather than signal?
- Can discourse communities collectively detect disruptions individual readers miss?
- Do high-disagreement items signal contested values or measurement noise?
- Why does consensus-seeking destroy information in normative but not factual tasks?
- What measurement artifacts emerge when annotators interpret the same question differently?
- Why are ground truth labels missing from unlabeled domain evaluations?
77 specific questions
- What distinguishes surface cues from structural meaning in language understanding?
- Why does frame-activation matter more than word-by-word composition?
- Can language meaning emerge without joint attention and shared embodied interaction?
- What makes some interpretive postures stick while others fail to form?
- Can language models acquire meaning from distributional patterns alone without joint attention?
- What distinguishes real understanding from superficial pattern matching?
- Why does training data saliency distort how models judge meaning?
62 specific questions
- How do meta-tokens help models learn when to generate reasoning versus commit predictions?
- Can models internally identify which tokens matter most for reasoning?
- What evidence shows that reasoning chains encode token-level functional structure?
- Why does token-level gradient targeting matter more than aggregate loss?
- Does the token prediction framing actually capture what human reasoning does?
- Why does the first generated token trigger collapse of task superposition?
- What makes token selection more important than adaptation strategy?
45 specific questions
- How does dataset composition affect which internal directions encode misaligned behavior?
- Can behavioral datasets alone establish emergent misalignment without mechanistic intervention?
- Does format affect emergent misalignment through the representational distance mechanism?
- Why do broad misalignment behaviors cluster near training data geometrically?
- Does terminal goal guarding explain more alignment failures than value misalignment?
- Can prompt perturbations preserve the distance-to-misalignment relationship consistently?
- Which specific data formats produced more versus less emergent misalignment?