Models of human preference for learning reward functions

Models of human preference for learning reward functions
复制标题

DOI:
10.48550/arxiv.2206.02231
复制
发表时间:
2022-06
期刊:
ArXiv
影响因子:
--
通讯作者:
W. B. Knox;Stephane Hatgis-Kessell;S. Booth;S. Niekum;P. Stone;A. Allievi
W. B. Knox;Stephane Hatgis-Kessell;S. Booth;S. Niekum;P. Stone;A. Allievi
中科院分区:
其他
文献类型:
--
作者:
W. B. Knox;Stephane Hatgis-Kessell;S. Booth;S. Niekum;P. Stone;A. Allievi

文献摘要

相似文献

强化学习的效用受到奖励函数与人类利益相关者利益一致性的限制。一种有前途的对齐方法是从人类生成的轨迹段对之间的偏好中学习奖励函数,这是一种基于人类反馈的强化学习(RLHF)。这些人的偏好通常被认为仅仅是由部分回报所告知的,部分回报是沿着每个部分的回报之和。我们发现这个假设是有缺陷的,并建议建模人类的偏好,而不是告知每个部分的遗憾,一个部分的偏离最佳决策的措施。鉴于无限多的偏好产生的遗憾,我们证明,我们可以确定一个奖励功能相当于奖励功能,产生这些偏好,我们证明,以前的部分回报模型缺乏这种可识别性属性在多个上下文中。我们的实证研究表明,我们提出的后悔偏好模型优于部分回报偏好模型与有限的训练数据,在其他相同的设置。此外,我们发现,我们提出的后悔偏好模型更好地预测了真实的人类偏好,并从这些偏好中学习奖励函数,从而制定出更符合人类的政策。总的来说,这项工作表明偏好模型的选择是有影响力的,并且我们提出的后悔偏好模型对最近研究的核心假设进行了改进。我们已经开源了我们的实验代码,我们收集的人类偏好数据集,以及我们收集这样一个数据集的训练和偏好获取接口。
The utility of reinforcement learning is limited by the alignment of reward functions with the interests of human stakeholders. One promising method for alignment is to learn the reward function from human-generated preferences between pairs of trajectory segments, a type of reinforcement learning from human feedback (RLHF). These human preferences are typically assumed to be informed solely by partial return, the sum of rewards along each segment. We find this assumption to be flawed and propose modeling human preferences instead as informed by each segment's regret, a measure of a segment's deviation from optimal decision-making. Given infinitely many preferences generated according to regret, we prove that we can identify a reward function equivalent to the reward function that generated those preferences, and we prove that the previous partial return model lacks this identifiability property in multiple contexts. We empirically show that our proposed regret preference model outperforms the partial return preference model with finite training data in otherwise the same setting. Additionally, we find that our proposed regret preference model better predicts real human preferences and also learns reward functions from these preferences that lead to policies that are better human-aligned. Overall, this work establishes that the choice of preference model is impactful, and our proposed regret preference model provides an improvement upon a core assumption of recent research. We have open sourced our experimental code, the human preferences dataset we gathered, and our training and preference elicitation interfaces for gathering a such a dataset.