Operant Matching as a Nash Equilibrium of an Intertemporal Game

Operant Matching as a Nash Equilibrium of an Intertemporal Game
复制标题

DOI:
10.1162/neco.2009.09-08-854
复制
发表时间:
2009-10-01
期刊:
影响因子:
2.9
通讯作者:
Seung, H. Sebastian
Seung, H. Sebastian
中科院分区:
计算机科学4区
文献类型:
--
作者:
Loewenstein, Yonatan;Prelec, Drazen;Seung, H. Sebastian

文献摘要

被引文献

相似文献

在过去的几十年里,经济学家、心理学家和神经科学家进行了一些实验,让受试者(无论是人类还是动物)在不同的行动之间反复选择,并根据选择历史获得奖励。虽然个人的选择是不可预测的,但总体行为通常遵循Herrnstein的匹配定律:每个选择的平均回报对于所有选择的替代品都是相等的。一般来说,匹配行为不会最大化传递给主体的总体奖励,因此匹配似乎与效用最大化原则不一致。在这里,我们表明,匹配可以与最大化一致,考虑一个单一的主题的选择是由一个序列的多个自我的每一个时刻的时间。如果每个自我都对世界的状态视而不见,并完全低估了未来的回报,那么最终的博弈至少有一个纳什均衡满足赫恩斯坦匹配定律和个人选择的不可预测性。这种均衡一般来说是帕累托次优的,可以理解为跨期囚徒困境中多重自我的相互背叛,关于多重自我的数学假设不应该被解释为心理学假设。人类和动物都记得过去的选择,关心未来的回报。然而,他们可能无法理解或考虑到过去和未来之间的关系。当考虑到一种收敛于均衡的机制(如强化学习)时,这一点可以更加明确。使用具体的例子,我们表明存在满足匹配律但不是纳什均衡的行为。我们预计这些行为不会在动物和人类的实验中观察到。如果是这样的话,纳什均衡公式可以被看作是Herrnstein匹配定律的一个改进。
Over the past several decades, economists, psychologists, and neuroscientists have conducted experiments in which a subject, human or animal, repeatedly chooses between alternative actions and is rewarded based on choice history. While individual choices are unpredictable, aggregate behavior typically follows Herrnstein's matching law: the average reward per choice is equal for all chosen alternatives. In general, matching behavior does not maximize the overall reward delivered to the subject, and therefore matching appears inconsistent with the principle of utility maximization. Here we show that matching can be made consistent with maximization by regarding the choices of a single subject as being made by a sequence of multiple selves-one for each instant of time. If each self is blind to the state of the world and discounts future rewards completely, then the resulting game has at least one Nash equilibrium that satisfies both Herrnstein's matching law and the unpredictability of individual choices. This equilibrium is, in general, Pareto suboptimal, and can be understood as a mutual defection of the multiple selves in an intertemporal prisoner's dilemma.The mathematical assumptions about the multiple selves should not be interpreted literally as psychological assumptions. Human and animals do remember past choices and care about future rewards. However, they may be unable to comprehend or take into account the relationship between past and future. This can be made more explicit when a mechanism that converges on the equilibrium, such as reinforcement learning, is considered.Using specific examples, we show that there exist behaviors that satisfy the matching law but are not Nash equilibria. We expect that these behaviors will not be observed experimentally in animals and humans. If this is the case, the Nash equilibrium formulation can be regarded as a refinement of Herrnstein's matching law.