Learning Optimal Advantage from Preferences and Mistaking it for Reward

Learning Optimal Advantage from Preferences and Mistaking it for Reward
复制标题

DOI:
10.48550/arxiv.2310.02456
复制
发表时间:
2023-10
期刊:
ArXiv
影响因子:
--
通讯作者:
W. B. Knox;Stephane Hatgis-Kessell;Sigurdur O. Adalgeirsson;Serena Booth;Anca D. Dragan;Peter Stone;S. Niekum
W. B. Knox;Stephane Hatgis-Kessell;Sigurdur O. Adalgeirsson;Serena Booth;Anca D. Dragan;Peter Stone;S. Niekum
中科院分区:
其他
文献类型:
--
作者:
W. B. Knox;Stephane Hatgis-Kessell;Sigurdur O. Adalgeirsson;Serena Booth;Anca D. Dragan;Peter Stone;S. Niekum

文献摘要

相似文献

我们考虑从人类对轨迹段的偏好中学习奖励函数的算法,正如在人类反馈强化学习(RLHF)中使用的那样。最近的研究假设,人类的偏好仅基于在这些部分中累积的奖励或部分回报而产生。最近的研究对这一假设的有效性提出了质疑,提出了一种基于后悔的替代偏好模型。我们调查了假设偏好是基于部分回报的结果,而实际上它们是由后悔引起的。我们认为,学习函数是最优优势函数的近似值,而不是奖励函数。我们发现,如果解决了一个特定的陷阱,这种错误的假设并不是特别有害,而是导致高度成形的奖励功能。尽管如此,这种不正确地使用最优优势函数的近似值比适当的更简单的贪心最大化方法更不可取。从后悔偏好模型的角度,我们也为RLHF微调当代大型语言模型提供了更清晰的解释。这篇论文总体上提供了关于为什么在部分回报偏好模型下的学习在实践中如此有效的见解,尽管它不符合人类如何给出偏好。
We consider algorithms for learning reward functions from human preferences over pairs of trajectory segments, as used in reinforcement learning from human feedback (RLHF). Most recent work assumes that human preferences are generated based only upon the reward accrued within those segments, or their partial return. Recent work casts doubt on the validity of this assumption, proposing an alternative preference model based upon regret. We investigate the consequences of assuming preferences are based upon partial return when they actually arise from regret. We argue that the learned function is an approximation of the optimal advantage function, not a reward function. We find that if a specific pitfall is addressed, this incorrect assumption is not particularly harmful, resulting in a highly shaped reward function. Nonetheless, this incorrect usage of the approximation of the optimal advantage function is less desirable than the appropriate and simpler approach of greedy maximization of it. From the perspective of the regret preference model, we also provide a clearer interpretation of fine tuning contemporary large language models with RLHF. This paper overall provides insight regarding why learning under the partial return preference model tends to work so well in practice, despite it conforming poorly to how humans give preferences.