Most of What You Heard About RLHF Is Slightly Wrong
Reinforcement learning from human feedback sits at the center of almost every credible large language model deployed today, and almost every credible misconception about how those models actually work.