Manifold-based multi-objective policy search with sample reuse

Manifold-based multi-objective policy search with sample reuse
复制标题

DOI:
10.1016/j.neucom.2016.11.094
复制
发表时间:
2017-11
期刊:
影响因子:
6
通讯作者:
Simone Parisi;Matteo Pirotta;Jan Peters
Simone Parisi;Matteo Pirotta;Jan Peters
中科院分区:
计算机科学2区
文献类型:
--
作者:
Simone Parisi;Matteo Pirotta;Jan Peters

文献摘要

被引文献

相似文献

许多现实世界的应用程序的特点是多个相互冲突的目标。在这些问题中,最优性被帕累托最优性所取代,目标是找到帕累托边界,一组代表目标之间不同妥协的解决方案。尽管最近在多目标优化方面取得了进展,但实现帕累托边界的准确表示仍然是一个重要的挑战。基于强化学习和多目标策略搜索的最新进展,我们提出了两种新的基于流形的算法来解决多目标马尔可夫决策过程。这些算法结合了联合收割机的情节探索策略和重要性采样,以有效地学习策略参数空间中的流形,使其在目标空间中的图像准确地逼近帕累托前沿。我们表明,基于情节的方法和重要性抽样可以在多目标强化学习的背景下产生更好的结果。在三个多目标问题上进行评估,我们的算法在学习的帕累托前沿和样本效率方面都优于最先进的方法。
Many real-world applications are characterized by multiple conflicting objectives. In such problems optimality is replaced by Pareto optimality and the goal is to find the Pareto frontier, a set of solutions representing different compromises among the objectives. Despite recent advances in multi-objective optimization, achieving an accurate representation of the Pareto frontier is still an important challenge. Building on recent advances in reinforcement learning and multi-objective policy search, we present two novel manifold-based algorithms to solve multi-objective Markov decision processes. These algorithms combine episodic exploration strategies and importance sampling to efficiently learn a manifold in the policy parameter space such that its image in the objective space accurately approximates the Pareto frontier. We show that episode-based approaches and importance sampling can lead to significantly better results in the context of multi-objective reinforcement learning. Evaluated on three multi-objective problems, our algorithms outperform state-of-the-art methods both in terms of quality of the learned Pareto frontier and sample efficiency.