Spike-based decision learning of Nash equilibria in two-player games.

Spike-based decision learning of Nash equilibria in two-player games.
复制标题

DOI:
10.1371/journal.pcbi.1002691
复制
发表时间:
2012
影响因子:
4.3
通讯作者:
Senn W
Senn W
中科院分区:
生物学2区
文献类型:
--
作者:
Friedrich J;Senn W

文献摘要

参考文献

被引文献

相似文献

人类和动物在不确定的多智能体环境中面临决策任务,其中智能体的策略可能由于其他策略的共同适应而随时间变化。然而,这种自适应决策背后的神经元基底和计算算法在很大程度上是未知的。我们提出了一个群体编码模型的尖峰神经元的政策梯度过程,成功地获得经典的博弈论任务的最佳策略。建议的群体强化学习再现了21点和检查员游戏的人类行为实验数据。它分别根据纯(确定性)和混合(随机)纳什均衡进行优化。相比之下,时间差(TD)学习,协方差学习和基本的强化学习无法为随机策略提供最佳性能。基于尖峰的人口强化学习,遵循随机奖励梯度,因此是一个可行的候选人来解释自动决策学习的纳什均衡在两个玩家的游戏。社会经济的互动被捕获在一个博弈论框架中,由多个代理人对一个商品池采取行动,以最大限度地提高自己的回报。神经经济学试图用神经元的术语来解释主体的行为。神经经济学中的经典模型使用时间差异(TD)学习。该算法以增量方式更新状态-动作对的值,并根据基于值的策略选择动作。相比之下,策略梯度方法不引入值作为中间步骤,而是直接导出使总期望回报最大化的动作选择策略。我们考虑一个决策网络的神经元人口,提出的时空尖峰模式,编码二进制行动的人口输出尖峰列车和随后的多数表决。动作选择策略由投射到群体神经元的突触的强度参数化。梯度学习规则的推导修改这些突触的强度,这取决于四个因素,突触前和突触后活动,行动和奖励。我们发现,对于经典的博弈论任务,我们的决策网络赋予的四因素学习规则导致纳什最优的行动选择。它还模仿人类对这些相同任务的决策学习。
Humans and animals face decision tasks in an uncertain multi-agent environment where an agent's strategy may change in time due to the co-adaptation of others strategies. The neuronal substrate and the computational algorithms underlying such adaptive decision making, however, is largely unknown. We propose a population coding model of spiking neurons with a policy gradient procedure that successfully acquires optimal strategies for classical game-theoretical tasks. The suggested population reinforcement learning reproduces data from human behavioral experiments for the blackjack and the inspector game. It performs optimally according to a pure (deterministic) and mixed (stochastic) Nash equilibrium, respectively. In contrast, temporal-difference(TD)-learning, covariance-learning, and basic reinforcement learning fail to perform optimally for the stochastic strategy. Spike-based population reinforcement learning, shown to follow the stochastic reward gradient, is therefore a viable candidate to explain automated decision learning of a Nash equilibrium in two-player games. Socio-economic interactions are captured in a game theoretic framework by multiple agents acting on a pool of goods to maximize their own reward. Neuroeconomics tries to explain the agent's behavior in neuronal terms. Classical models in neuroeconomics use temporal-difference(TD)-learning. This algorithm incrementally updates values of state-action pairs, and actions are selected according to a value-based policy. In contrast, policy gradient methods do not introduce values as intermediate steps, but directly derive an action selection policy which maximizes the total expected reward. We consider a decision making network consisting of a population of neurons which, upon presentation of a spatio-temporal spike pattern, encodes binary actions by the population output spike trains and a subsequent majority vote. The action selection policy is parametrized by the strengths of synapses projecting to the population neurons. A gradient learning rule is derived which modifies these synaptic strengths and which depends on four factors, the pre- and postsynaptic activities, the action and the reward. We show that for classical game-theoretical tasks our decision making network endowed with the four-factor learning rule leads to Nash-optimal action selections. It also mimics human decision learning for these same tasks.
DOI: 10.3389/fncom.2010.00017
发表时间: 2010
影响因子: 3.2
作者:
Loewenstein Y
通讯作者: Loewenstein Y
DOI: 10.1162/1532443041827880
发表时间: 2004-08-15
影响因子: 6
作者:
Hu, JL;Wellman, MP
通讯作者: Wellman, MP
DOI: 10.1523/jneurosci.6249-09.2010
发表时间: 2010-10-06
影响因子: 5.3
作者:
Fremaux, Nicolas;Sprekeler, Henning;Gerstner, Wulfram
通讯作者: Gerstner, Wulfram
DOI: 10.1162/neco.2010.05-09-1010
发表时间: 2010-07-01
期刊: NEURAL COMPUTATION
影响因子: 2.9
作者:
Friedrich, Johannes;Urbanczik, Robert;Senn, Walter
通讯作者: Senn, Walter
DOI: 10.1103/physrevlett.97.048104
发表时间: 2006-07-28
影响因子: 8.6
作者:
Fiete, Ila R.;Seung, H. Sebastian
通讯作者: Seung, H. Sebastian