Multi-objective deep inverse reinforcement learning for weight estimation of objectives

Multi-objective deep inverse reinforcement learning for weight estimation of objectives
复制标题

用于目标权重估计的多目标深度逆强化学习

DOI:
10.1007/s10015-022-00773-8
复制
发表时间:
2022
影响因子:
0.9
通讯作者:
Arai Sachiyo
Arai Sachiyo
中科院分区:
--
文献类型:
--
作者:
Takayama Naoya;Arai Sachiyo

文献摘要

参考文献

被引文献

相似文献

权重是一个参数,用于在线性缩放每个目标的奖励向量时衡量多目标强化学习中的优先级。权重需要提前设定;然而,大多数现实世界的问题都有许多目标。因此,调整权重需要设计者进行多次试验和错误。此外,还需要一种自动估算权重的方法,以减轻设计者设置权重的负担。在本文中,我们提出了一种使用逆强化学习(IRL)框架基于每个目标的奖励向量和专家轨迹来估计权重的新方法。特别是,我们采用深度 IRL 与深度强化学习和乘法权重学徒学习,以实现连续状态空间中的快速权重估计。通过在连续状态空间中的多目标顺序决策问题的基准环境中的实验,我们验证了我们的新颖的权重估计方法优于投影方法和贝叶斯优化。
Weight is a parameter used for measuring the priority in multi-objective reinforcement learning when linearly scalarizing the reward vector for each objective. The weights need to be set in advance; however, most real-world problems have numerous objectives. Therefore, adjusting the weights requires many trials and errors by the designer. In addition, a method to automatically estimate weights is needed to reduce the burden on designers to set weights. In this paper, we propose a novel method for estimating the weights based on the reward vector for each objective and the expert trajectories using the framework of inverse reinforcement learning (IRL). In particular, we adopt deep IRL with deep reinforcement learning and multiplicative weights apprenticeship learning for fast weight estimation in a continuous state space. Through experiments in a benchmark environment for multi-objective sequential decision-making problems in a continuous state space, we verified that our novel weight estimation method is superior to the projection method and Bayesian optimization.
具有内存保留的可扩展贝叶斯优化
DOI: --
发表时间: 2019
期刊:
影响因子: --
作者:
Hidetaka Ito
通讯作者: Hidetaka Ito