Global optimality of softmax policy gradient with single hidden layer neural networks in the mean-field regime

Global optimality of softmax policy gradient with single hidden layer neural networks in the mean-field regime
复制标题

DOI:
--
复制
发表时间:
2020-10
期刊:
ArXiv
影响因子:
--
通讯作者:
A. Agazzi;Jianfeng Lu
A. Agazzi;Jianfeng Lu
中科院分区:
其他
文献类型:
--
作者:
A. Agazzi;Jianfeng Lu

文献摘要

相似文献

研究了具有softmax策略和非线性函数逼近策略梯度算法训练的无限期折扣马尔可夫决策过程的策略优化问题。我们专注于平均场机制中的训练动态,例如建模,宽单隐层神经网络的行为,当通过熵正则化鼓励探索时。这些模型的动力学被建立为参数空间中分布的Wasserstein梯度流。我们进一步证明了全局最优性的不动点的动力学在温和的条件下,他们的初始化。
We study the problem of policy optimization for infinite-horizon discounted Markov Decision Processes with softmax policy and nonlinear function approximation trained with policy gradient algorithms. We concentrate on the training dynamics in the mean-field regime, modeling e.g., the behavior of wide single hidden layer neural networks, when exploration is encouraged through entropy regularization. The dynamics of these models is established as a Wasserstein gradient flow of distributions in parameter space. We further prove global optimality of the fixed points of this dynamics under mild conditions on their initialization.