Mean-Variance Policy Iteration for Risk-Averse Reinforcement Learning

Mean-Variance Policy Iteration for Risk-Averse Reinforcement Learning
复制标题

DOI:
10.1609/aaai.v35i12.17302
复制
发表时间:
2020-04
期刊:
--
影响因子:
--
通讯作者:
Shangtong Zhang;Bo Liu;Shimon Whiteson
Shangtong Zhang;Bo Liu;Shimon Whiteson
中科院分区:
其他
文献类型:
--
作者:
Shangtong Zhang;Bo Liu;Shimon Whiteson

文献摘要

相似文献

我们提出了一种用于风险厌恶控制的均值-方差策略迭代(MVPI)框架,该框架优化了每步奖励随机变量的方差。MVPI具有很大的灵活性,任何政策评估方法和风险中性控制方法都可以加入现成的风险厌恶控制,无论是在政策设置上还是在政策设置下。这种灵活性减少了风险中性控制和风险厌恶控制之间的差距,并通过直接开发新的增强MDP来实现。我们提出风险厌恶型TD3作为MVPI的实例,在风险感知性能指标下,在挑战Mujoco机器人仿真任务时,它优于香草TD3和许多以前的风险厌恶控制方法。这个规避风险的TD3是第一个将确定性策略和非策略学习引入规避风险的强化学习的,这两者都是我们在Mujoco域中展示的性能提升的关键。
We present a mean-variance policy iteration (MVPI) framework for risk-averse control in a discounted infinite horizon MDP optimizing the variance of a per-step reward random variable. MVPI enjoys great flexibility in that any policy evaluation method and risk-neutral control method can be dropped in for risk-averse control off the shelf, in both on- and off-policy settings. This flexibility reduces the gap between risk-neutral control and risk-averse control and is achieved by working on a novel augmented MDP directly. We propose risk-averse TD3 as an example instantiating MVPI, which outperforms vanilla TD3 and many previous risk-averse control methods in challenging Mujoco robot simulation tasks under a risk-aware performance metric. This risk-averse TD3 is the first to introduce deterministic policies and off-policy learning into risk-averse reinforcement learning, both of which are key to the performance boost we show in Mujoco domains.