Dynamic Weights in Multi-Objective Deep Reinforcement Learning

Dynamic Weights in Multi-Objective Deep Reinforcement Learning
复制标题

多目标深度强化学习中的动态权重

DOI:
--
复制
发表时间:
2018
期刊:
International Conference on Machine Learning
影响因子:
--
通讯作者:
Denis Steckelmacher
Denis Steckelmacher
中科院分区:
--
文献类型:
--
作者:
Axel Abels;D. Roijers;T. Lenaerts;A. Nowé;Denis Steckelmacher

文献摘要

被引文献

相似文献

许多真实的决策问题的特点是多个相互冲突的目标,必须平衡基于它们的相对重要性。在动态权重设置中,相对重要性随时间变化,需要处理这种变化的专门算法,例如Natarajan & Tadepalli(2005)的表格强化学习(RL)算法。然而,这种早期的工作是不可行的RL设置,需要使用函数逼近。我们通过提出一个多目标Q网络来概括权重变化和高维输入,该网络的输出取决于目标的相对重要性,并引入多样性经验重放(DER)来对抗动态权重设置的固有非平稳性。我们进行了广泛的实验评估,并将我们的方法与深度多任务/多目标RL的自适应算法进行了比较,结果表明,我们提出的网络与DER相结合,在权重变化场景和问题域中主导了这些自适应算法。
Many real world decision problems are characterized by multiple conflicting objectives which must be balanced based on their relative importance. In the dynamic weights setting the relative importance changes over time and specialized algorithms that deal with such change, such as the tabular Reinforcement Learning (RL) algorithm by Natarajan & Tadepalli (2005), are required. However, this earlier work is not feasible for RL settings that necessitate the use of function approximators. We generalize across weight changes and high-dimensional inputs by proposing a multi-objective Q-network whose outputs are conditioned on the relative importance of objectives, and introduce Diverse Experience Replay (DER) to counter the inherent non-stationarity of the dynamic weights setting. We perform an extensive experimental evaluation and compare our methods to adapted algorithms from Deep Multi-Task/Multi-Objective RL and show that our proposed network in combination with DER dominates these adapted algorithms across weight change scenarios and problem domains.