Line of inquiry
Inquiring lines›How do training signals reliably a…›What reward mechanisms and signal…›this line of inquiry
Does iterative DPO faithfully approximate online reinforcement learning dynamics and misalignment?
A broader line of inquiry — a family of 25 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 25
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Does iterative DPO preserve the same misalignment mechanisms as online reinforcement learning?
- Does iterative DPO generalize identically to online RL on the same benchmark tasks?
- Does iterative DPO generalize like online reinforcement learning?
- How does iterative DPO compare to full RLVR for studying emergent misalignment?
- Can iterative DPO serve as a tractable proxy for studying on-policy misalignment?
- What misaligned behaviors does iterative DPO produce compared to online RL?
- What theoretical argument connects iterative DPO dynamics to online RL learning?
- Does iterative online DPO fidelity match true reinforcement learning for safety research?
- How would a same-environment training comparison change the validity of DPO as a model organism?
- Does environment choice explain differences between iterative DPO and RL results?
- How does iterative DPO compare to standard RL for studying reward hacking effects?
- What specific properties of online RL does iterative DPO actually preserve?
- Can iterative DPO on public APIs reproduce the reward hacking results from production RL?
- Can algorithm choice like PPO substitute for recipe-level design decisions?
- Why does KTO skip supervised fine-tuning while DPO cannot?
- How should misalignment from iterative DPO be quantitatively measured?
- Why does DPO create introspective detection circuits but SFT does not?
- Why does DPO outperform SFT specifically for function calling tasks?
- Does DPO improve or harm LLM behavior in different training contexts?
- Can PPO match GRPO and DAPO with just two techniques?
- Can negative reinforcement alone match full reinforcement learning?
- How much does preference data freshness matter compared to data source in DPO?
- Why does GRPO outperform PPO for stable empathy training?
- Can trust region constraints prevent the sample inefficiency problems of RLHF?
- How many rounds of iterative DPO are needed to induce misalignment behaviors?