Monte Carlo preference elicitation for learning additive reward functions

Monte Carlo preference elicitation for learning additive reward functions
复制标题

用于学习加性奖励函数的蒙特卡洛偏好启发

DOI:
10.1109/roman.2012.6343863
复制
发表时间:
2012
期刊:
2012 IEEE RO-MAN: The 21st IEEE International Symposium on Robot and Human Interactive Communication
影响因子:
--
通讯作者:
M. Veloso
M. Veloso
中科院分区:
--
文献类型:
--
作者:
Stephanie Rosenthal;M. Veloso

文献摘要

被引文献

相似文献

包括机器人在内的人工智能代理经常使用奖励函数来评估不同状态和行为之间的权衡,并确定最佳策略。我们对奖励函数特别感兴趣,这些奖励函数可以分解成在独立的子问题或状态空间的特征上计算的子奖励的相加和。如果这些子奖励捕获不同的奖励指标,如用户满意度和任务完成时间,则不清楚如何在奖励函数中缩放子奖励以产生适当的策略。在这项工作中,我们提出并评估了一种新的蒙特卡罗方法来学习子奖励的缩放因子,其中训练引起人类在两种状态-行动场景之间的偏好。由于该算法引出了对明确场景的偏好,因此与以前的引出方法相比,它更不容易受到人为错误的影响。这些偏好被用来在比例因子上生成一组不等式,我们用线性程序有效地解决了这些不等式。我们表明,我们的算法要求与蒙特卡罗方法中使用的比例因子假设数量的对数成比例的许多偏好。
AI agents including robots often use reward functions to evaluate tradeoffs between different states and actions and to determine optimal policies. We are particularly interested in reward functions that can be decomposed into an additive sum of subrewards that are computed on independent subproblems or features of the state space. If these subrewards capture different reward metrics, such as user satisfaction and task completion time, it is unclear how to scale the subrewards in the reward function to produce an appropriate policy. In this work, we propose and evaluate a novel Monte Carlo method for learning the scaling factors of subrewards, in which the training elicits humans' preferences between two state-action scenarios. Because the algorithm elicits preferences over explicit scenarios, it is less susceptible to human error than previous elicitation approaches. The preferences are used to generate a set of inequalities over the scaling factors that we solve efficiently using a linear program. We show that our algorithm asks for a number of preferences proportional to log of the number of scaling factor hypotheses used in the Monte Carlo method.