Deconstructing the human algorithms for exploration.

Deconstructing the human algorithms for exploration.
复制标题

DOI:
10.1016/j.cognition.2017.12.014
复制
发表时间:
2018-04
期刊:
影响因子:
3.4
通讯作者:
Gershman SJ
Gershman SJ
中科院分区:
心理学2区
文献类型:
--
作者:
Gershman SJ

文献摘要

参考文献

被引文献

相似文献

信息收集(探索)和奖励寻求(利用)之间的困境是强化学习代理的一个基本问题。人类如何解决这一困境仍然是一个悬而未决的问题,因为实验提供了关于人类使用的底层算法的模棱两可的证据。我们表明,两个家庭的算法可以区分方面的不确定性如何影响探索。基于不确定性奖金的算法预测响应偏差的变化作为不确定性的函数,而基于采样的算法预测响应斜率的变化。两个实验提供了证据的偏见和斜率的变化,和计算建模证实,混合模型是最好的定量帐户的数据。
The dilemma between information gathering (exploration) and reward seeking (exploitation) is a fundamental problem for reinforcement learning agents. How humans resolve this dilemma is still an open question, because experiments have provided equivocal evidence about the underlying algorithms used by humans. We show that two families of algorithms can be distinguished in terms of how uncertainty affects exploration. Algorithms based on uncertainty bonuses predict a change in response bias as a function of uncertainty, whereas algorithms based on sampling predict a change in response slope. Two experiments provide evidence for both bias and slope changes, and computational modeling confirms that a hybrid model is the best quantitative account of the data.
DOI: 10.1038/nature04766
发表时间: 2006-06-15
期刊: NATURE
影响因子: 64.8
作者:
Daw, Nathaniel D.;O'Doherty, John P.;Dayan, Peter;Seymour, Ben;Dolan, Raymond J.
通讯作者: Dolan, Raymond J.
DOI: 10.1098/rstb.2007.2098
发表时间: 2007-05-29
影响因子: 6.3
作者:
Cohen, Jonathan D.;McClure, Samuel M.;Yu, Angela J.
通讯作者: Yu, Angela J.
DOI: 10.1016/j.visres.2015.03.004
发表时间: 2016-09-01
期刊: VISION RESEARCH
影响因子: 1.8
作者:
Gershman, Samuel J.;Tenenbaum, Joshua B.;Jaekel, Frank
通讯作者: Jaekel, Frank
DOI: 10.1023/a:1013689704352
发表时间: 2002-01-01
期刊: MACHINE LEARNING
影响因子: 7.5
作者:
Auer, P;Cesa-Bianchi, N;Fischer, P
通讯作者: Fischer, P
DOI: 10.1016/j.cogsys.2010.07.007
发表时间: 2011-06-01
影响因子: 3.9
作者:
Lee, Michael D.;Zhang, Shunan;Steyvers, Mark
通讯作者: Steyvers, Mark