Monte Carlo preference elicitation for learning additive reward functions
Monte Carlo preference elicitation for learning additive reward functions
复制标题
用于学习加性奖励函数的蒙特卡洛偏好启发
DOI:
10.1109/roman.2012.6343863
复制
发表时间:
2012
期刊:
影响因子:
--
通讯作者:
M. Veloso
中科院分区:
文献类型:
--
作者:
Stephanie Rosenthal;M. Veloso
AI agents including robots often use reward functions to evaluate tradeoffs between different states and actions and to determine optimal policies. We are particularly interested in reward functions that can be decomposed into an additive sum of subrewards that are computed on independent subproblems or features of the state space. If these subrewards capture different reward metrics, such as user satisfaction and task completion time, it is unclear how to scale the subrewards in the reward function to produce an appropriate policy. In this work, we propose and evaluate a novel Monte Carlo method for learning the scaling factors of subrewards, in which the training elicits humans' preferences between two state-action scenarios. Because the algorithm elicits preferences over explicit scenarios, it is less susceptible to human error than previous elicitation approaches. The preferences are used to generate a set of inequalities over the scaling factors that we solve efficiently using a linear program. We show that our algorithm asks for a number of preferences proportional to log of the number of scaling factor hypotheses used in the Monte Carlo method.