Line of inquiry

Why does reinforcement learning suppress output diversity compared to supervised fine-tuning?

A broader line of inquiry — a family of 20 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.


Questions in this line of inquiry 20

Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.