Learning optimal decisions with confidence

Learning optimal decisions with confidence
复制标题

DOI:
10.1073/pnas.1906787116
复制
发表时间:
2019-12-03
影响因子:
11.1
通讯作者:
Pouget, Alexandre
Pouget, Alexandre
中科院分区:
综合性期刊1区
文献类型:
--
作者:
Drugowitsch, Jan;Mendonca, Andre G.;Pouget, Alexandre

文献摘要

被引文献

相似文献

扩散决策模型(DDM)是在不确定性和时间压力下进行决策的非常成功的模型。在感知决策的背景下,这些模型通常从两个输入单元开始,以神经元-反神经元对的形式组织。相比之下,在大脑中,感觉输入是通过大型神经元群体的活动编码的。此外,虽然DDM是手工布线的,但神经系统必须通过试错来学习网络的权重。目前还没有关于DDM学习的规范理论,因此也没有关于决策者如何在这种情况下学习做出最佳决策的理论。在这里,我们推导出这样一个规则,用于学习一个接近最优的线性组合的DDM输入的基础上试验的试验反馈。该规则是贝叶斯的,因为它不仅学习权重的平均值,而且还学习协方差矩阵形式的平均值周围的不确定性。在这个规则中,学习的速率与不正确(正确)决策的置信度成正比(分别成反比)。此外,我们表明,在不稳定的环境中,该规则预测的偏见重复相同的选择后,正确的决定,与偏见的强度,是由以前的选择的难度调制。最后,我们将我们的学习规则扩展到其中一个选择更有可能是先验的情况,这提供了对这种偏见如何调节扩散模型中导致最优决策的机制的见解。
Diffusion decision models (DDMs) are immensely successful models for decision making under uncertainty and time pressure. In the context of perceptual decision making, these models typically start with two input units, organized in a neuron-antineuron pair. In contrast, in the brain, sensory inputs are encoded through the activity of large neuronal populations. Moreover, while DDMs are wired by hand, the nervous system must learn the weights of the network through trial and error. There is currently no normative theory of learning in DDMs and therefore no theory of how decision makers could learn to make optimal decisions in this context. Here, we derive such a rule for learning a near-optimal linear combination of DDM inputs based on trial-by-trial feedback. The rule is Bayesian in the sense that it learns not only the mean of the weights but also the uncertainty around this mean in the form of a covariance matrix. In this rule, the rate of learning is proportional (respectively, inversely proportional) to confidence for incorrect (respectively, correct) decisions. Furthermore, we show that, in volatile environments, the rule predicts a bias toward repeating the same choice after correct decisions, with a bias strength that is modulated by the previous choice's difficulty. Finally, we extend our learning rule to cases for which one of the choices is more likely a priori, which provides insights into how such biases modulate the mechanisms leading to optimal decisions in diffusion models.