Reach adaption to a visuomotor gain with terminal error feedback involves reinforcement learning.

Reach adaption to a visuomotor gain with terminal error feedback involves reinforcement learning.
复制标题

DOI:
10.1371/journal.pone.0269297
复制
发表时间:
2022
期刊:
影响因子:
3.7
通讯作者:
--
中科院分区:
综合性期刊3区
文献类型:
--
作者:

文献摘要

参考文献

相似文献

运动适应可以通过由感觉预测错误驱动的基于错误的学习或由奖励预测错误驱动的强化学习来实现。最近关于视觉适应的研究表明,与提供运动的连续视觉反馈的基于错误的学习相比,当视觉反馈被移除时,强化学习会导致更持久的适应。然而,有证据表明,基于错误的学习与终端视觉反馈的运动(在运动结束时提供)可能是由感官和奖励预测错误。在这里,我们研究了反馈对学习的影响,使用视觉适应任务,参与者移动光标到一个单一的目标,而手和光标移动位移之间的增益逐渐改变。不同的组接受连续错误反馈(EC),终端错误反馈(ET),或二元强化反馈(成功/失败)在运动结束时(R)。适应后,我们测试了泛化到位于不同方向的目标,发现在ET组的泛化是EC和R组之间的中间。然后,我们研究了持久性的适应EC和ET组时,光标熄灭,只提供二进制奖励反馈。尽管ET组的表现保持不变,但EC组的表现迅速恶化。这些结果表明,终端错误反馈导致一个更强大的学习形式比连续错误反馈。此外,我们的研究结果是一致的,认为基于错误的学习与终端反馈涉及基于错误和强化学习。
Motor adaptation can be achieved through error-based learning, driven by sensory prediction errors, or reinforcement learning, driven by reward prediction errors. Recent work on visuomotor adaptation has shown that reinforcement learning leads to more persistent adaptation when visual feedback is removed, compared to error-based learning in which continuous visual feedback of the movement is provided. However, there is evidence that error-based learning with terminal visual feedback of the movement (provided at the end of movement) may be driven by both sensory and reward prediction errors. Here we examined the influence of feedback on learning using a visuomotor adaptation task in which participants moved a cursor to a single target while the gain between hand and cursor movement displacement was gradually altered. Different groups received either continuous error feedback (EC), terminal error feedback (ET), or binary reinforcement feedback (success/fail) at the end of the movement (R). Following adaptation we tested generalization to targets located in different directions and found that generalization in the ET group was intermediate between the EC and R groups. We then examined the persistence of adaptation in the EC and ET groups when the cursor was extinguished and only binary reward feedback was provided. Whereas performance was maintained in the ET group, it quickly deteriorated in the EC group. These results suggest that terminal error feedback leads to a more robust form of learning than continuous error feedback. In addition our findings are consistent with the view that error-based learning with terminal feedback involves both error-based and reinforcement learning.
DOI: 10.1038/s41562-018-0324-5
发表时间: 2018-04
影响因子: 29.9
作者:
Heald JB;Ingram JN;Flanagan JR;Wolpert DM
通讯作者: Wolpert DM
DOI: 10.1007/s00221-010-2209-3
发表时间: 2010-05
影响因子: 2
作者:
Shabbott, Britne A.;Sainburg, Robert L.
通讯作者: Sainburg, Robert L.
DOI: 10.1038/s41586-021-04129-3
发表时间: 2021-12
期刊: Nature
影响因子: 64.8
作者:
Heald JB;Lengyel M;Wolpert DM
通讯作者: Wolpert DM
DOI: 10.7554/elife.71627
发表时间: 2021-11-19
期刊: eLife
影响因子: 7.7
作者:
Cesanek E;Zhang Z;Ingram JN;Wolpert DM;Flanagan JR
通讯作者: Flanagan JR
DOI: 10.1152/jn.2000.84.2.853
发表时间: 2000-08-01
影响因子: 2.5
作者:
Scheidt, RA;Reinkensmeyer, DJ;Mussa-Ivaldi, FA
通讯作者: Mussa-Ivaldi, FA