Human and Machine Learning in Non-Markovian Decision Making

Human and Machine Learning in Non-Markovian Decision Making
复制标题

DOI:
10.1371/journal.pone.0123105
复制
发表时间:
2015-04-21
期刊:
影响因子:
3.7
通讯作者:
Herzog, Michael H.
Herzog, Michael H.
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Clarke, Aaron Michael;Friedrich, Johannes;Herzog, Michael H.

文献摘要

被引文献

相似文献

人类可以在各种反馈条件下学习。强化学习(RL)是一种特别重要的学习类型,其中必须做出一系列有奖励的决策。RL的计算和行为研究主要集中在马尔可夫决策过程,其中下一个状态仅取决于当前状态和动作。人们对非马尔可夫决策知之甚少,在非马尔可夫决策中,下一个状态依赖于比当前状态和动作更多的东西。例如,当行为和反馈之间没有唯一的映射时,学习就是非马尔可夫的。我们已经建立了一个基于尖峰神经元的模型,可以通过执行策略梯度下降来处理这些非马尔可夫条件[1]。在这里,我们检查了模型的性能,并将其与人类学习和贝叶斯最优参考进行了比较,贝叶斯最优参考提供了性能的上限。我们发现,在所有情况下,我们的尖峰神经元模型很好地描述了人类的表现。
Humans can learn under a wide variety of feedback conditions. Reinforcement learning (RL), where a series of rewarded decisions must be made, is a particularly important type of learning. Computational and behavioral studies of RL have focused mainly on Markovian decision processes, where the next state depends on only the current state and action. Little is known about non-Markovian decision making, where the next state depends on more than the current state and action. Learning is non-Markovian, for example, when there is no unique mapping between actions and feedback. We have produced a model based on spiking neurons that can handle these non-Markovian conditions by performing policy gradient descent [1]. Here, we examine the model's performance and compare it with human learning and a Bayes optimal reference, which provides an upper-bound on performance. We find that in all cases, our population of spiking neurons model well-describes human performance.