Reinforcement learning for stabilizing an inverted pendulum naturally leads to intermittent feedback control as in human quiet standing

Reinforcement learning for stabilizing an inverted pendulum naturally leads to intermittent feedback control as in human quiet standing
复制标题

用于稳定倒立摆的强化学习自然会导致间歇性反馈控制,就像人类安静站立一样

DOI:
10.1109/embc.2016.7590634
复制
发表时间:
2016
期刊:
Proceedings of the Annual International Conference of the IEEE Engineering in Medicine and Biology Society, EMBS 2016
影响因子:
--
通讯作者:
Taishin Nomura
Taishin Nomura
中科院分区:
--
文献类型:
--
作者:
Kenjiro Michimoto;Yasuyuki Suzuki;Ken Kiyono;Yasushi Kobayashi;Pietro Morasso;Taishin Nomura

文献摘要

相似文献

间歇反馈控制用于稳定人体直立姿态是一种很有前途的策略,替代标准的时间连续刚度控制。在这里,我们表明,这样的间歇控制器可以通过强化学习自然建立。为此,我们使用了一个单倒立摆模型的直立姿势和一个非常简单的奖励功能,给予一定数量的惩罚时,倒立摆福尔斯下降或改变其在状态空间中的位置。我们发现,所获得的反馈控制器表现出间歇反馈控制策略的特征,即当摆的状态位于无主动控制的倒立摆的不稳定鞍型直立平衡的稳定流形附近时,反馈控制器的动作间歇地关闭:该动作提供了在没有主动反馈控制的帮助下利用朝向不稳定直立位置的瞬时收敛动力学的机会。然后,我们推测这种强化学习可能的生理机制,并建议它可能与脑干的脚桥被盖核(PPN)的神经活动。最近的证据表明,PPN可能在帕金森病患者的姿势紧张,奖励预测以及姿势不稳定的产生和调节中发挥关键作用,这一假设得到了支持。
Intermittent feedback control for stabilizing human upright stance is a promising strategy, alternative to the standard time-continuous stiffness control. Here we show that such an intermittent controller can be established naturally through reinforcement learning. To this end, we used a single inverted pendulum model of the upright posture and a very simple reward function that gives a certain amount of punishments when the inverted pendulum falls or changes its position in the state space. We found that the acquired feedback controller exhibits hallmarks of the intermittent feedback control strategy, namely the action of the feedback controller is switched-off intermittently when the state of the pendulum is located near the stable manifold of the unstable saddle-type upright equilibrium of the inverted pendulum with no active control: this action provides an opportunity to exploit transiently converging dynamics toward the unstable upright position with no help of the active feedback control. We then speculate about a possible physiological mechanism of such reinforcement learning, and suggest that it may be related to the neural activity in the pedunculopontine tegmental nucleus (PPN) of the brainstem. This hypothesis is supported by recent evidence indicating that PPN might play critical roles for generation and regulation of postural tonus, reward prediction, as well as postural instability in patients with Parkinson's disease.