Reward Rate Optimization in Two-Alternative Decision Making: Empirical Tests of Theoretical Predictions

Reward Rate Optimization in Two-Alternative Decision Making: Empirical Tests of Theoretical Predictions
复制标题

DOI:
10.1037/a0016926
复制
发表时间:
2009-12-01
影响因子:
2.1
通讯作者:
Cohen, Jonathan D.
Cohen, Jonathan D.
中科院分区:
心理学3区
文献类型:
--
作者:
Simen, Patrick;Contreras, David;Cohen, Jonathan D.

文献摘要

被引文献

相似文献

漂移扩散模型(DDM)实现了一个最佳的决策过程,固定的,2-可选的被迫选择任务。决策阈值的高度应用于积累信息的每次试验确定的DDM的速度-准确性权衡(SAT),从而占一个普遍存在的功能,在快速响应任务的人的表现。然而,很少有人知道参与者如何解决特定的权衡。一种可能性是,他们选择的SAT最大化的主观奖励率的表现。对于DDM,在奖励正确反应的自由反应任务中,其阈值和起始点参数存在唯一的、奖励率最大化的值(R。Bogacz,E. Brown,J. Moehlis,P. Holmes,& J. D. Cohen,2006)。这些最佳值作为响应-刺激间隔、先前刺激概率和正确响应的相对奖励幅度的函数而变化。我们测试了这些任务操作下的响应时间,准确性和响应偏差的定量预测,发现分组数据符合最佳参数化DDM的预测。
The drift-diffusion model (DDM) implements an optimal decision procedure for stationary, 2-alternative forced-choice tasks. The height of a decision threshold applied to accumulating information on each trial determines a speed-accuracy tradeoff (SAT) for the DDM, thereby accounting for a ubiquitous feature of human performance in speeded response tasks. However, little is known about how participants settle on particular tradeoffs. One possibility is that they select SATs that maximize a subjective rate of reward earned for performance. For the DDM, there exist unique, reward-rate-maximizing values for its threshold and starting point parameters in free-response tasks that reward correct responses (R. Bogacz, E. Brown, J. Moehlis, P. Holmes, & J. D. Cohen, 2006). These optimal values vary as a function of response-stimulus interval, prior stimulus probability, and relative reward magnitude for correct responses. We tested the resulting quantitative predictions regarding response time, accuracy, and response bias under these task manipulations and found that grouped data conformed well to the predictions of an optimally parameterized DDM.