Cooperative update of beliefs and state-transition functions in human reinforcement learning

Cooperative update of beliefs and state-transition functions in human reinforcement learning
复制标题

DOI:
10.1038/s41598-019-53600-9
复制
发表时间:
2019-11
期刊:
影响因子:
4.6
通讯作者:
Hiroshi Higashi;T. Minami;S. Nakauchi
Hiroshi Higashi;T. Minami;S. Nakauchi
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Hiroshi Higashi;T. Minami;S. Nakauchi

文献摘要

相似文献

众所周知,大脑中的强化学习系统通过与环境的相互作用来促进学习。这些系统能够解决多维问题,其中一些维度与奖励相关,而另一些则无关。为了解决这些问题,计算模型使用贝叶斯学习,这是一种由人类行为和神经证据支持的策略。贝叶斯学习考虑了信念,它代表了学习者对与奖励相关的特定维度的信心。信念是作为状态转换(奖励)函数的后验概率给出的,该函数将最佳行为映射到每个维度的状态。然而,当涉及到实施这种学习策略时,信念和状态转换函数更新的顺序仍然不清楚。本研究在学习者必须识别奖励相关维度的任务中,使用对人类行为和脑电图信号的逐次分析来研究这种更新顺序。我们的行为和神经结果揭示了一种合作更新——在结果反馈后300毫秒内,状态转换函数被更新,随后是每个维度的信念。
It is widely known that reinforcement learning systems in the brain contribute to learning via interactions with the environment. These systems are capable of solving multidimensional problems, in which some dimensions are relevant to a reward, while others are not. To solve these problems, computational models use Bayesian learning, a strategy supported by behavioral and neural evidence in human. Bayesian learning takes into account beliefs, which represent a learner’s confidence in a particular dimension being relevant to the reward. Beliefs are given as a posterior probability of the state-transition (reward) function that maps the optimal actions to the states in each dimension. However, when it comes to implementing this learning strategy, the order in which beliefs and state-transition functions update remains unclear. The present study investigates this update order using a trial-by-trial analysis of human behavior and electroencephalography signals during a task in which learners have to identify the reward-relevant dimension. Our behavioral and neural results reveal a cooperative update—within 300 ms after the outcome feedback, the state-transition functions are updated, followed by the beliefs for each dimension.