Implicit Value Updating Explains Transitive Inference Performance: The Betasort Model.

Implicit Value Updating Explains Transitive Inference Performance: The Betasort Model.
复制标题

DOI:
10.1371/journal.pcbi.1004523
复制
发表时间:
2015
影响因子:
4.3
通讯作者:
Terrace HS
Terrace HS
中科院分区:
生物学2区
文献类型:
--
作者:
Jensen G;Muñoz F;Alkan Y;Ferrera VP;Terrace HS

文献摘要

被引文献

相似文献

传递推理(在已知B > C和C > D的情况下推断出B b> D的能力)是连续学习的一个普遍特征,在几十个物种中都有观察到。尽管有这些强大的行为效应,但依赖于奖励预测误差或联想强度的强化学习模型通常无法执行这些推断。我们提出了一种被称为betasort的算法,该算法受认知过程的启发,以低计算成本执行传递推理。这是通过以下方式实现的:(1)使用beta分布沿单位跨度表示刺激位置,(2)不对称地处理正反馈和负反馈,以及(3)在每次试验中更新每个刺激的位置,无论该刺激是否可见。比较了恒河猴、人类、betasort算法以及Q-learning(一种已建立的奖励预测误差(RPE)模型)的性能。在这些测试中,只有Q-learning在关键测试中没有做出高于随机的反应。Betasort的成功(与RPE模型相比)和它的计算效率(与完全马尔可夫决策过程实现相比)表明,生物强化学习的研究最好是通过特征驱动的方法来比较正式模型。尽管机器学习系统可以解决各种各样的问题,但它们在逻辑推理方面的能力仍然有限。我们开发了一种新的计算模型,称为betasort,它解决了一类问题的这些限制:在这些问题中,算法必须通过试错来推断一组项目的顺序。与现有的机器学习系统不同(但与儿童和许多非人类动物一样),betasort能够对一组图像的顺序执行“传递推理”。betasort的错误模式类似于儿童和非人类动物的错误模式,并且由此产生的学习以较低的计算成本实现。此外,根据机器学习文献中分类的正式规范,betasort很难被分类为“无模型”或“基于模型”。这些结果的一个更广泛的含义是,要想更全面地理解大脑是如何学习的,就需要分析人员考虑其他候选学习模型。
Transitive inference (the ability to infer that B > D given that B > C and C > D) is a widespread characteristic of serial learning, observed in dozens of species. Despite these robust behavioral effects, reinforcement learning models reliant on reward prediction error or associative strength routinely fail to perform these inferences. We propose an algorithm called betasort, inspired by cognitive processes, which performs transitive inference at low computational cost. This is accomplished by (1) representing stimulus positions along a unit span using beta distributions, (2) treating positive and negative feedback asymmetrically, and (3) updating the position of every stimulus during every trial, whether that stimulus was visible or not. Performance was compared for rhesus macaques, humans, and the betasort algorithm, as well as Q-learning, an established reward-prediction error (RPE) model. Of these, only Q-learning failed to respond above chance during critical test trials. Betasort’s success (when compared to RPE models) and its computational efficiency (when compared to full Markov decision process implementations) suggests that the study of reinforcement learning in organisms will be best served by a feature-driven approach to comparing formal models. Although machine learning systems can solve a wide variety of problems, they remain limited in their ability to make logical inferences. We developed a new computational model, called betasort, which addresses these limitations for a certain class of problems: Those in which the algorithm must infer the order of a set of items by trial and error. Unlike extant machine learning systems (but like children and many non-human animals), betasort is able to perform “transitive inferences” about the ordering of a set of images. The patterns of error made by betasort resemble those made by children and non-human animals, and the resulting learning achieved at low computational cost. Additionally, betasort is difficult to classify as either “model-free” or “model-based” according to the formal specifications of those classifications in the machine learning literature. One of the broader implications of these results is that achieving a more comprehensive understanding of how the brain learns will require analysts to entertain other candidate learning models.