Transforming Policy via Reward Advancement

Transforming Policy via Reward Advancement
复制标题

DOI:
10.1109/cdc40024.2019.9029286
复制
发表时间:
2019-12
期刊:
2019 IEEE 58th Conference on Decision and Control (CDC)
影响因子:
--
通讯作者:
Guojun Wu;Yanhua Li;Jun Luo
Guojun Wu;Yanhua Li;Jun Luo
中科院分区:
其他
文献类型:
--
作者:
Guojun Wu;Yanhua Li;Jun Luo

文献摘要

相似文献

真实的世界中的许多人类行为都可以被描述为连续的决策过程,例如城市出行者对交通方式和路线的选择[1]。与机器控制的选择不同,机器通常遵循完美理性以采用具有最高回报的策略,研究表明人类代理在有限理性下做出次优决策[2]。这种行为可以使用最大因果熵(MCE)原理进行建模[3]。在本文中,我们定义并研究了一个新的奖励转换问题(即奖励推进):在MCE原则下,恢复将Agent的策略从πo转换为预定义的目标策略πt的额外奖励函数的范围。我们表明,给定一个MDP和一个目标策略πt,有无限多的额外的奖励函数,可以实现所需的政策转换。此外,我们提出了一个算法,以进一步提取额外的奖励最小的“成本”,以实现政策的转变。我们使用合成数据和来自中国深圳的大规模(6个月)城市级公共交通数据证明了我们的奖励提升解决方案的正确性和准确性。
Many real world human behaviors can be characterized as sequential decision making processes, such as urban travelers’ choices of transport modes and routes [1]. Differing from choices controlled by machines, which in general follows perfect rationality to adopt the policy with highest reward, studies have revealed that human agents make sub-optimal decisions under bounded rationality [2]. Such behaviors can be modeled using maximum causal entropy (MCE) principle [3]. In this paper, we define and investigate a novel reward transformation problem (namely, reward advancement): Recovering the range of additional reward functions that transform the agent’s policy from πo to a predefined target policy πt under MCE principle. We show that given an MDP and a target policy πt, there are infinite many additional reward functions that can achieve the desired policy transformation. Moreover, we propose an algorithm to further extract the additional rewards with minimum "cost" to implement the policy transformation. We demonstrated the correctness and accuracy of our reward advancement solution using both synthetic data and a largescale (6 months) passenger-level public transit data from Shenzhen, China.