Feynman-Kac Neural Network Architectures for Stochastic Control Using Second-Order FBSDE Theory

Feynman-Kac Neural Network Architectures for Stochastic Control Using Second-Order FBSDE Theory
复制标题

DOI:
--
复制
发表时间:
2020
期刊:
--
影响因子:
--
通讯作者:
M. Pereira;Ziyi Wang;T. Chen;Emily A. Reed;Evangelos A. Theodorou
M. Pereira;Ziyi Wang;T. Chen;Emily A. Reed;Evangelos A. Theodorou
中科院分区:
其他
文献类型:
--
作者:
M. Pereira;Ziyi Wang;T. Chen;Emily A. Reed;Evangelos A. Theodorou

文献摘要

相似文献

针对一类完全非线性的Hamilton Jacobi Bellman偏微分方程组所描述的随机最优控制问题,提出了一种深度递归神经网络结构。当考虑具有可加性、状态依赖和控制乘性的不确定性的随机动力学时,这类偏微分方程组就会出现。具有这些特征的随机模型在计算神经科学、生物学、fiNance和航空航天系统中非常重要,并且比只具有加性不确定性的模型提供了更准确的激励表示。以前的文献已经证明了线性HJB理论对于这类问题的不足,因此,人们提出了依赖于推广的Feynman-Kac引理的方法,从而得到了一个二阶向前向后向随机微分方程组。然而,到目前为止,这些方法存在组合错误,导致缺乏可扩展性。在本文中,我们提出了一种基于深度学习的算法,该算法利用二阶FBSDE表示和基于LSTM的递归神经网络不仅解决了这类随机最优控制问题,而且克服了传统方法所面临的可扩展性问题。在一个高维线性系统和三个来自机器人和生物力学的非线性系统上进行了仿真测试,验证了该控制算法的可行性和优于以往方法的性能。
We present a deep recurrent neural network architecture to solve a class of stochastic optimal control problems described by fully nonlinear Hamilton Jacobi Bellman partial differential equations. Such PDEs arise when considering stochastic dynamics characterized by uncertainties that are additive, state dependent, and control multiplicative. Stochastic models with these characteristics are important in computational neuroscience, biology, finance, and aerospace systems and provide a more accurate representation of actuation than models with only additive uncertainty. Previous literature has established the inadequacy of the linear HJB theory for such problems, so instead, methods relying on the generalized version of the Feynman-Kac lemma have been proposed resulting in a system of second-order Forward-Backward SDEs. However, so far, these methods suffer from compounding errors resulting in lack of scalability. In this paper, we propose a deep learning based algorithm that leverages the second-order FBSDE representation and LSTM-based recurrent neural networks to not only solve such stochastic optimal control problems but also overcome the problems faced by traditional approaches, including scalability. The resulting control algorithm is tested on a high-dimensional linear system and three nonlinear systems from robotics and biomechanics in simulation to demonstrate feasibility and out-performance against previous methods.