Integrating Reinforcement Learning, Equilibrium Points, and Minimum Variance to Understand the Development of Reaching: A Computational Model

Integrating Reinforcement Learning, Equilibrium Points, and Minimum Variance to Understand the Development of Reaching: A Computational Model
复制标题

DOI:
10.1037/a0037016
复制
发表时间:
2014-07-01
影响因子:
5.4
通讯作者:
Baldassarre, Gianluca
Baldassarre, Gianluca
中科院分区:
心理学1区
文献类型:
--
作者:
Caligiore, Daniele;Parisi, Domenico;Baldassarre, Gianluca

文献摘要

被引文献

相似文献

尽管有大量关于触达行为的文献,但关于婴儿触达行为发展背后的运动控制过程,仍然缺乏一个清晰的概念。本文通过提出一个基于三个关键假设的计算模型来克服这一差距:(A)试错学习过程驱动到达行为的渐进发展;(B)基于平衡点的运动控制使模型能够快速找到与目标对象获得接触问题的初始近似解;(C)在肌肉噪声存在的情况下对末端运动的精度的要求驱动到达行为的渐进精细化。基于两自由度模拟动态手臂的模型测试表明,该模型能够再现大量的经验结果,其中大部分来自对儿童的纵向研究:几个到达动作的动力学和运动学变量的发展轨迹,组成到达的子动作的时间演变,钟形速度分布的渐进发展,以及冗余自由度管理的演变。该模型还对其中几种现象做出了可检验的预测。这些经验数据中的大多数从未被以前的计算模型研究过,更重要的是,从未被一个独特的模型所解释。在这方面,对模型功能的分析表明,所有这些结果最终都可以通过上述三个假设的相互作用中出现的相同发展轨迹来解释,有时是以意想不到的方式:模型首先快速学习执行确保手与目标接触的粗略动作(这是一项具有很大适应价值的成就),然后缓慢地细化对运动的动力学方面的详细控制,以提高准确性。
Despite the huge literature on reaching behavior, a clear idea about the motor control processes underlying its development in infants is still lacking. This article contributes to overcoming this gap by proposing a computational model based on three key hypotheses: (a) trial-and-error learning processes drive the progressive development of reaching; (b) the control of the movements based on equilibrium points allows the model to quickly find the initial approximate solution to the problem of gaining contact with the target objects; (c) the request of precision of the end movement in the presence of muscular noise drives the progressive refinement of the reaching behavior. The tests of the model, based on a two degrees of freedom simulated dynamical arm, show that it is capable of reproducing a large number of empirical findings, most deriving from longitudinal studies with children: the developmental trajectory of several dynamical and kinematic variables of reaching movements, the time evolution of submovements composing reaching, the progressive development of a bell-shaped speed profile, and the evolution of the management of redundant degrees of freedom. The model also produces testable predictions on several of these phenomena. Most of these empirical data have never been investigated by previous computational models and, more important, have never been accounted for by a unique model. In this respect, the analysis of the model functioning reveals that all these results are ultimately explained, sometimes in unexpected ways, by the same developmental trajectory emerging from the interplay of the three mentioned hypotheses: The model first quickly learns to perform coarse movements that assure a contact of the hand with the target (an achievement with great adaptive value) and then slowly refines the detailed control of the dynamical aspects of movement to increase accuracy.