Looking Back and Ahead: Adaptation and Planning by Gradient Descent

Looking Back and Ahead: Adaptation and Planning by Gradient Descent
复制标题

回顾与展望:梯度下降的适应和规划

DOI:
10.1109/devlrn.2019.8850693
复制
发表时间:
2019
期刊:
Proceedings of the Ninth Joint IEEE International Conference on Development and Learning and on Epigenetic Robotics (ICDL-EpiRob 2019)
影响因子:
--
通讯作者:
Murata Shingo,Sawa Hiroki,Sugano Shigeki,Ogata Tetsuya
Murata Shingo,Sawa Hiroki,Sugano Shigeki,Ogata Tetsuya
中科院分区:
--
文献类型:
--
作者:
Ryo Karakida;Shotaro Akaho;Shun-ichi Amari;Murata Shingo,Sawa Hiroki,Sugano Shigeki,Ogata Tetsuya

文献摘要

参考文献

相似文献

适应和规划对生物和人工媒介都至关重要。在这项研究中,我们把这些作为一个推理问题,我们使用基于梯度的优化方法来解决。我们提出了适应和规划梯度下降(APGrade),一个基于梯度的计算框架与层次递归神经网络(RNN)的适应和规划。该框架通过基于实际观察回顾过去的情况和基于首选观察(或目标)展望未来的情况来计算(反事实)预测误差。在最小化这些误差的方向上优化RNN的更高级别的内部状态。过去的错误有助于适应,而未来的错误有助于规划。建议的APGrade框架中实现的人形机器人和机器人执行一个球的操作任务与人类实验者。实验结果表明,给定一个特定的偏好,机器人可以适应意外的情况,同时追求自己的偏好,通过规划未来的行动。
Adaptation and planning are crucial for both biological and artificial agents. In this study, we treat these as an inference problem that we solve using a gradient-based optimization approach. We propose adaptation and planning by gradient descent (APGraDe), a gradient-based computational framework with a hierarchical recurrent neural network (RNN) for adaptation and planning. This framework computes (counterfactual) prediction errors by looking back on past situations based on actual observations and by looking ahead to future situations based on preferred observations (or goal). The internal state of the higher level of the RNN is optimized in the direction of minimizing these errors. The errors for the past contribute to the adaptation while errors for the future contribute to the planning. The proposed APGraDe framework is implemented in a humanoid robot and the robot performs a ball manipulation task with a human experimenter. Experimental results show that given a particular preference, the robot can adapt to unexpected situations while pursuing its own preference through the planning of future actions.
DOI: 10.1007/978-3-319-11179-7_46
发表时间: 2014-09
期刊: --
影响因子: --
作者:
K. Takahashi;T. Ogata;Hadi Tjandra;Shingo Murata;H. Arie;S. Sugano
通讯作者: K. Takahashi;T. Ogata;Hadi Tjandra;Shingo Murata;H. Arie;S. Sugano
通过预测误差最小化机制出现两个机器人之间的交互行为
DOI: 10.1109/devlrn.2016.7846838
发表时间: 2016
期刊: 2016 Joint IEEE International Conference on Development and Learning and Epigenetic Robotics (ICDL-EpiRob)
影响因子: --
作者:
Yiwen Chen;Shingo Murata;H. Arie;T. Ogata;J. Tani;S. Sugano
通讯作者: S. Sugano
DOI: --
发表时间: 2018-04
期刊: ArXiv
影响因子: --
作者:
A. Srinivas;A. Jabri;P. Abbeel;S. Levine;Chelsea Finn
通讯作者: A. Srinivas;A. Jabri;P. Abbeel;S. Levine;Chelsea Finn
学习通过自下而上和自上而下的交互过程生成明确的行为
DOI: 10.1016/s0893-6080(02)00214-9
发表时间: 2003
期刊: Neural networks : the official journal of the International Neural Network Society
影响因子: --
作者:
J. Tani
通讯作者: J. Tani
使用自适应世界模型进行规划
DOI: --
发表时间: 1990
期刊: Neural Information Processing Systems
影响因子: --
作者:
S. Thrun;K. Möller;A. Linden
通讯作者: A. Linden