Credit assignment during movement reinforcement learning.

Credit assignment during movement reinforcement learning.
复制标题

运动强化学习期间的学分分配

DOI:
10.1371/journal.pone.0055352
复制
发表时间:
2013
期刊:
影响因子:
3.7
通讯作者:
Wei K
Wei K
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Dam G;Kording K;Wei K

文献摘要

参考文献

被引文献

相似文献

我们经常需要学习如何根据反映我们运动整体成功的单一性能指标进行运动。然而,运动具有许多属性,例如它们的轨迹,速度和终点的定时,因此大脑需要决定运动的哪些属性应该被改进;它需要解决信用分配问题。目前,人们对人类如何在强化学习的背景下解决信用分配问题知之甚少。在这里,我们测试了人类参与者如何在自动学习任务中解决这些问题。在没有明确定义的目标运动的情况下,参与者在逐个试验的基础上进行手部接触并获得金钱奖励作为反馈。所尝试的到达轨迹的曲率和方向以可以通过实验操纵的方式确定所接收的金钱奖励。根据行动-奖励对的历史,参与者很快解决了信用分配问题,并学习了隐含的支付函数。一个内置遗忘的贝叶斯学分分配模型准确地预测了他们的逐个尝试学习。
We often need to learn how to move based on a single performance measure that reflects the overall success of our movements. However, movements have many properties, such as their trajectories, speeds and timing of end-points, thus the brain needs to decide which properties of movements should be improved; it needs to solve the credit assignment problem. Currently, little is known about how humans solve credit assignment problems in the context of reinforcement learning. Here we tested how human participants solve such problems during a trajectory-learning task. Without an explicitly-defined target movement, participants made hand reaches and received monetary rewards as feedback on a trial-by-trial basis. The curvature and direction of the attempted reach trajectories determined the monetary rewards received in a manner that can be manipulated experimentally. Based on the history of action-reward pairs, participants quickly solved the credit assignment problem and learned the implicit payoff function. A Bayesian credit-assignment model with built-in forgetting accurately predicts their trial-by-trial learning.
DOI: 10.1152/jn.2002.88.4.1685
发表时间: 2002-10-01
影响因子: 2.5
作者:
Martin, TA;Norris, SA;Thach, WT
通讯作者: Thach, WT
DOI: 10.1371/journal.pcbi.1002159
发表时间: 2011-09
影响因子: 4.3
作者:
Sternad D;Abe MO;Hu X;Müller H
通讯作者: Müller H
DOI: 10.1037/0096-1523.30.1.212
发表时间: 2004-02-01
影响因子: 2.1
作者:
Müller, H;Sternad, D
通讯作者: Sternad, D
DOI: 10.1038/nn963
发表时间: 2002-11-01
影响因子: 25
作者:
Todorov, E;Jordan, MI
通讯作者: Jordan, MI
DOI: 10.1080/00222890009601384
发表时间: 2000-12-01
影响因子: 1.4
作者:
Kudo, K;Ito, T;Ishikura, T
通讯作者: Ishikura, T