A stochastic policy search model for matching behavior

A stochastic policy search model for matching behavior
复制标题

DOI:
10.1007/s11432-011-4304-x
复制
发表时间:
2011-06
期刊:
Science China Information Sciences
影响因子:
--
通讯作者:
Zhenbo Cheng;Yu Zhang;Zhidong Deng
Zhenbo Cheng;Yu Zhang;Zhidong Deng
中科院分区:
其他
文献类型:
--
作者:
Zhenbo Cheng;Yu Zhang;Zhidong Deng

文献摘要

被引文献

相似文献

匹配定律是决策理论中的基本经验定律之一,它指出受试者对可选目标的偏好取决于哪些选择得到了强化。在本文中,我们研究了解释为什么被试的决策经常遵守这一定律的可能机制。在强化学习理论的基础上,提出了一种通过策略参数更新策略的决策模型,该模型可以通过前额叶皮层和基底神经节神经回路在大脑中实现。基于该模型,在一些简单的假设下,得到了满足匹配律的算法。理论分析和仿真结果表明,该算法的决策行为服从匹配规律。此外,在两个经典的实验中的匹配行为再现使用该算法。我们的结果提供了一个合理的匹配律的策略和奖励决策任务的一个有用的计算工具。
The matching law is one of the basic empirical laws in decision theory, and it states that a subject’s preference to optional targets depends on which choices are reinforced. In this paper, we study the possible mechanisms that explain why subjects’ decisions often obey this law. On the basis of reinforcement learning theory, we put forward a decision-making model in which the policy is updated by a policy parameter, and the model might be implemented in the brain through the prefrontal cortex and the basal ganglia neural circuit. Based on this model, an algorithm that satisfies the matching law is derived under some simple assumptions. Theoretical analysis and simulation results show that the decision behavior achieved by the algorithm obeys the matching law. In addition, the matching behaviors in two classical experiments are reproduced using the algorithm. Our results provide a reasonable strategy for the matching law and a useful computational tool for rewarded decision-making tasks.