When does reward maximization lead to matching law?

When does reward maximization lead to matching law?
复制标题

DOI:
10.1371/journal.pone.0003795
复制
发表时间:
2008
期刊:
影响因子:
3.7
通讯作者:
Fukai T
Fukai T
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Sakai Y;Fukai T

文献摘要

参考文献

被引文献

相似文献

在不同的行为环境中,被试采取何种策略一直是决策研究的核心问题。特别是,哪种行为策略,最大化或匹配,是更根本的动物的决策行为一直是一个争论的问题。在这里,我们证明了任何算法,以实现最大化的平均奖励的静态条件下,应导致匹配时,它忽略了依赖的预期结果对受试者的过去的选择。我们可以把这种部分报酬最大化策略称为“匹配策略”。然后,将该策略应用于受试者的决策系统更新用于做出决策的信息的情况。这些信息包括主体过去的动作或感官刺激,这些信息的内部存储通常被称为“状态变量”。我们证明,当与正确代表奖励最大化关键信息的状态变量的探索相结合时,匹配策略提供了一种最大化奖励的简单方法。我们的研究结果首次揭示了实现匹配行为的策略如何有利于奖励最大化,实现了对最大化和匹配之间关系的新见解。
What kind of strategies subjects follow in various behavioral circumstances has been a central issue in decision making. In particular, which behavioral strategy, maximizing or matching, is more fundamental to animal's decision behavior has been a matter of debate. Here, we prove that any algorithm to achieve the stationary condition for maximizing the average reward should lead to matching when it ignores the dependence of the expected outcome on subject's past choices. We may term this strategy of partial reward maximization “matching strategy”. Then, this strategy is applied to the case where the subject's decision system updates the information for making a decision. Such information includes subject's past actions or sensory stimuli, and the internal storage of this information is often called “state variables”. We demonstrate that the matching strategy provides an easy way to maximize reward when combined with the exploration of the state variables that correctly represent the crucial information for reward maximization. Our results reveal for the first time how a strategy to achieve matching behavior is beneficial to reward maximization, achieving a novel insight into the relationship between maximizing and matching.
DOI: 10.1901/jeab.2005.23-05
发表时间: 2005-11-01
影响因子: 2.7
作者:
Corrado, GS;Sugrue, LP;Newsome, WT
通讯作者: Newsome, WT
DOI: 10.1901/jeab.1989.51-215
发表时间: 1989-03-01
影响因子: 2.7
作者:
DAVISON, M;KERR, A
通讯作者: KERR, A
DOI: 10.1016/j.neunet.2006.05.034
发表时间: 2006-10-01
期刊: NEURAL NETWORKS
影响因子: 7.8
作者:
Sakai, Yutaka;Okamoto, Hiroshi;Fukai, Tomoki
通讯作者: Fukai, Tomoki
DOI: 10.1137/s0363012901385691
发表时间: 2003-01-01
影响因子: 2.2
作者:
Konda, VR;Tsitsiklis, JN
通讯作者: Tsitsiklis, JN
DOI: 10.1126/science.7292017
发表时间: 1981-01-01
期刊: SCIENCE
影响因子: 56.9
作者:
MAZUR, JE
通讯作者: MAZUR, JE