Neural networks and differential dynamic programming for reinforcement learning problems

Neural networks and differential dynamic programming for reinforcement learning problems
复制标题

DOI:
10.1109/icra.2016.7487755
复制
发表时间:
2016-05
期刊:
2016 IEEE International Conference on Robotics and Automation (ICRA)
影响因子:
--
通讯作者:
Akihiko Yamaguchi;C. Atkeson
Akihiko Yamaguchi;C. Atkeson
中科院分区:
其他
文献类型:
--
作者:
Akihiko Yamaguchi;C. Atkeson

文献摘要

相似文献

我们探索了一种基于模型的强化学习方法,其中学习了部分或完全未知的动态,并执行了明确的规划。我们用神经网络学习动态,用微分动态规划(DDP)规划行为。为了处理复杂的动力学,如操纵液体(倾倒),我们考虑时间分解动力学。我们从最近的工作开始[1],其中我们使用局部加权回归(LWR)来建模动态。本文的主要贡献是以具有随机DDP的神经网络的形式使用深度学习,并展示了神经网络相对于LWR的优势。为此,我们扩展了神经网络:(1)建模预测误差和输出噪声,(2)计算给定输入分布的输出概率分布,以及(3)计算输出期望相对于输入的梯度。由于神经网络具有非线性激活函数,这些扩展并不容易。我们提供了一个分析的解决方案,这些扩展使用一些简化的假设。通过浇注模拟实验验证了该方法的有效性。神经网络的学习性能优于LWR。减少了溢出材料的数量。我们还提出了使用PR2的机器人实验的早期结果。相关视频:https://youtu.be/aM3hE1J5W98。
We explore a model-based approach to reinforcement learning where partially or totally unknown dynamics are learned and explicit planning is performed. We learn dynamics with neural networks, and plan behaviors with differential dynamic programming (DDP). In order to handle complicated dynamics, such as manipulating liquids (pouring), we consider temporally decomposed dynamics. We start from our recent work [1] where we used locally weighted regression (LWR) to model dynamics. The major contribution of this paper is making use of deep learning in the form of neural networks with stochastic DDP, and showing the advantages of neural networks over LWR. For this purpose, we extend neural networks for: (1) modeling prediction error and output noise, (2) computing an output probability distribution for a given input distribution, and (3) computing gradients of output expectation with respect to an input. Since neural networks have nonlinear activation functions, these extensions were not easy. We provide an analytic solution for these extensions using some simplifying assumptions. We verified this method in pouring simulation experiments. The learning performance with neural networks was better than that of LWR. The amount of spilled materials was reduced. We also present early results of robot experiments using a PR2. Accompanying video: https://youtu.be/aM3hE1J5W98.