Actor-critic is implicitly biased towards high entropy optimal policies

Actor-critic is implicitly biased towards high entropy optimal policies
复制标题

DOI:
--
复制
发表时间:
2021-10
期刊:
ArXiv
影响因子:
--
通讯作者:
Yuzheng Hu;Ziwei Ji;Matus Telgarsky
Yuzheng Hu;Ziwei Ji;Matus Telgarsky
中科院分区:
其他
文献类型:
--
作者:
Yuzheng Hu;Ziwei Ji;Matus Telgarsky

文献摘要

相似文献

我们证明了最简单的行为者-批评家方法——通过与线性MDP交互使用TD更新的线性softmax策略,但没有明确的正则化或探索——不仅找到了最优策略,而且更倾向于高熵最优策略。为了证明这种偏差的强度,该算法不仅没有正则化,没有预测,没有像$\epsilon$-greedy那样的探索,而且在没有重置的单一轨迹上进行训练。高熵偏差的关键结果是,在所有先前的工作中以某种形式存在的MDP上的均匀混合假设可以被丢弃:高熵偏差的隐式正则化足以确保所有链混合并以高概率达到最优策略。作为辅助贡献,这项工作通过将参与者更新写为显式镜像下降来解耦参与者和评论家之间的关注点,提供了在策略空间的KL球内统一绑定混合时间的工具,并提供了具有自身隐式偏差的无投影TD分析,该分析可以从未混合的初始分布运行。
We show that the simplest actor-critic method -- a linear softmax policy updated with TD through interaction with a linear MDP, but featuring no explicit regularization or exploration -- does not merely find an optimal policy, but moreover prefers high entropy optimal policies. To demonstrate the strength of this bias, the algorithm not only has no regularization, no projections, and no exploration like $\epsilon$-greedy, but is moreover trained on a single trajectory with no resets. The key consequence of the high entropy bias is that uniform mixing assumptions on the MDP, which exist in some form in all prior work, can be dropped: the implicit regularization of the high entropy bias is enough to ensure that all chains mix and an optimal policy is reached with high probability. As auxiliary contributions, this work decouples concerns between the actor and critic by writing the actor update as an explicit mirror descent, provides tools to uniformly bound mixing times within KL balls of policy space, and provides a projection-free TD analysis with its own implicit bias which can be run from an unmixed starting distribution.