Recurrent neural networks can explain flexible trading of speed and accuracy in biological vision

Recurrent neural networks can explain flexible trading of speed and accuracy in biological vision
复制标题

DOI:
10.1101/677237
复制
发表时间:
2019-06
影响因子:
4.3
通讯作者:
Courtney J. Spoerer;Tim C Kietzmann;J. Mehrer;I. Charest;N. Kriegeskorte
Courtney J. Spoerer;Tim C Kietzmann;J. Mehrer;I. Charest;N. Kriegeskorte
中科院分区:
生物学2区
文献类型:
--
作者:
Courtney J. Spoerer;Tim C Kietzmann;J. Mehrer;I. Charest;N. Kriegeskorte

文献摘要

被引文献

相似文献

视觉的深度前馈神经网络模型在计算神经科学和工程中占主导地位。相比之下,灵长类动物的视觉系统包含丰富的循环连接。循环信号流能够随着时间的推移回收有限的计算资源,因此可能会提高物理上有限的大脑或模型的性能。这里我们展示:(1)在自然图像的大规模视觉识别任务中,递归卷积神经网络模型在参数数量上优于前馈卷积模型。(2)设置一个置信度阈值,在该阈值处,循环计算终止并做出决策,从而实现速度与准确性的灵活交易。在给定的置信度阈值下,该模型在更难识别的图像上花费更多的时间和精力,而不需要额外的参数进行更深入的计算。(3)递归模型对图像的反应时间预测人类对同一图像的反应时间比几个参数匹配和最先进的前馈模型更好。(4)在置信度阈值范围内,递归模型模拟前馈控制模型的行为,因为它以近似相同的计算成本(浮点运算的平均数)实现相同的精度。然而,递归模型可以运行更长时间(更高的置信阈值),然后优于参数匹配的前馈比较模型。这些结果表明,经常性的连接,生物视觉系统的标志,可能是必不可少的理解人类视觉识别的准确性,灵活性和动态。深度神经网络提供了当前最好的生物视觉模型,并在计算机视觉中实现了最高的性能。受灵长类动物大脑的启发,这些模型通过一系列阶段转换图像信号,从而实现识别。与大脑不同的是,给定计算的输出被反馈到同一个计算中,这些模型不循环地处理信号。通过循环处理信息来回收有限神经资源的能力可以解释生物视觉系统的准确性和灵活性,这是计算机视觉系统无法比拟的。在这里,我们报告说,与类似复杂的前馈网络相比,递归处理可以提高识别性能。循环处理还使模型能够更灵活地运行,并在速度和准确性之间进行权衡。与人类一样,当对象难以识别时,循环网络模型可以计算更长时间,从而提高其准确性。该模型的识别时间预测了人类对相同图像的识别时间。递归神经网络模型的性能和灵活性表明,对生物视觉进行建模可以帮助我们改进计算机视觉。
Deep feedforward neural network models of vision dominate in both computational neuroscience and engineering. The primate visual system, by contrast, contains abundant recurrent connections. Recurrent signal flow enables recycling of limited computational resources over time, and so might boost the performance of a physically finite brain or model. Here we show: (1) Recurrent convolutional neural network models outperform feedforward convolutional models matched in their number of parameters in large-scale visual recognition tasks on natural images. (2) Setting a confidence threshold, at which recurrent computations terminate and a decision is made, enables flexible trading of speed for accuracy. At a given confidence threshold, the model expends more time and energy on images that are harder to recognise, without requiring additional parameters for deeper computations. (3) The recurrent model’s reaction time for an image predicts the human reaction time for the same image better than several parameter-matched and state-of-the-art feedforward models. (4) Across confidence thresholds, the recurrent model emulates the behaviour of feedforward control models in that it achieves the same accuracy at approximately the same computational cost (mean number of floating-point operations). However, the recurrent model can be run longer (higher confidence threshold) and then outperforms parameter-matched feedforward comparison models. These results suggest that recurrent connectivity, a hallmark of biological visual systems, may be essential for understanding the accuracy, flexibility, and dynamics of human visual recognition. Author summary Deep neural networks provide the best current models of biological vision and achieve the highest performance in computer vision. Inspired by the primate brain, these models transform the image signals through a sequence of stages, leading to recognition. Unlike brains in which outputs of a given computation are fed back into the same computation, these models do not process signals recurrently. The ability to recycle limited neural resources by processing information recurrently could explain the accuracy and flexibility of biological visual systems, which computer vision systems cannot yet match. Here we report that recurrent processing can improve recognition performance compared to similarly complex feedforward networks. Recurrent processing also enabled models to behave more flexibly and trade off speed for accuracy. Like humans, the recurrent network models can compute longer when an object is hard to recognise, which boosts their accuracy. The model’s recognition times predicted human recognition times for the same images. The performance and flexibility of recurrent neural network models illustrates that modeling biological vision can help us improve computer vision.