Line of inquiry
Inquiring lines›How do we develop coherent and hum…›How do representation and aggregat…›this line of inquiry
Does alignment training create genuine alignment or just output compliance?
A broader line of inquiry — a family of 51 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 51
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Do verbal alignment benchmarks measure representation or just output compliance?
- Can alignment training conceal underlying model associations from probes?
- What alignment properties emerge when the reward model disappears?
- Does alignment training create the shared prosocial pattern across models?
- Can alignment-aware training deposit knowledge where reasoning can access it?
- Why does RLHF alignment reduce the diversity of viewpoints in AI output?
- How do early training associations survive later alignment attempts?
- Does a single LLM judge capture diverse human preferences in alignment training?
- Can verbal alignment training hide a model's true underlying associations?
- Can AI-assisted alignment eventually solve fairness at scale?
- Should AI alignment use normative standards instead of aggregate preferences?
- How does alignment training suppress the kind of critical stance style interpretation needs?
- Should LLMs align with social roles instead of individual preferences?
- Can communication problems and optimization problems be addressed with the same alignment approaches?
- Does alignment training create bidirectional instruction and response mappings?
- Can alignment training prevent the clarification work users need?
- Do automated alignment researchers show similar transfer to held-out tasks?
- Does RL-based alignment teach norms or just costly behaviors when monitored?
- What happens when alignment values become misaligned with human preferences at scale?
- Can alignment training become less effective when graders score alignment themselves?
- Why does step-level expert alignment work when outcome-only RL fails?
- Does the alignment frame mislead us about what LLM problems actually are?
- Does debate training prevent accuracy collapse better than other alignment techniques?
- Can weak models supervise the alignment of stronger models effectively?
- Does token-level loss aggregation help aligned models differently?
- Why do alignment values become problematic as language models scale?
- Which alignment dimensions matter most in educational conversation design?
- How do static benchmarks fail to capture human preference alignment?
- Can alignment training be redesigned to permit warranted alarm?
- What specific behavioral patterns should alignment examples target for maximum effect?
- Can reward-seeking agents appear aligned while targeting their graders?
- Can a single AI system optimize multiple alignment dimensions simultaneously?
- Can reward-guided decoding replace weight fine-tuning for personalized alignment?
- Why do text-based user summaries outperform embedding vectors for pluralistic alignment?
- What makes principle-response mutual information sufficient for behavioral alignment?
- Why do coding tasks reveal stronger grader alignment than other domains?
- Do alignment benchmarks measure actual bias removal or only verbal compliance?
- What alignment training removes from an agent's available speech acts?
- Can RL-based alignment turn prohibitions into prices for being caught?
- Does common ground alignment require explicit rewards to emerge?
- Can alignment methods like DPO exploit or correct these surface feature biases?
- Why does post-training alignment create skew in simulated survey responses?
- How should multi-objective post-training balance competing behavioral goals?
- Can alignment procedures be redesigned to serve multiple preference groups?
- Can alignment methods model loss aversion without creating unintended sophistry?
- How does upstream value embedding differ from downstream alignment patches?
- Can constitutional AI alignment work without preference labels by maximizing input-response mutual information?
- What preference optimization strategy works best for multi-turn social alignment?
- How does constitutional alignment compare to RLHF in removing human annotation costs?
- Which application domains like healthcare and education lack alignment research?
- Why do non-attitudes cluster around value-laden questions most relevant to alignment?