A high-throughput screening approach to discovering good forms of biologically inspired visual representation.

A high-throughput screening approach to discovering good forms of biologically inspired visual representation.
复制标题

DOI:
10.1371/journal.pcbi.1000579
复制
发表时间:
2009-11
影响因子:
4.3
通讯作者:
Cox DD
Cox DD
中科院分区:
生物学2区
文献类型:
--
作者:
Pinto N;Doukhan D;DiCarlo JJ;Cox DD

文献摘要

参考文献

被引文献

相似文献

尽管生物对象识别的许多模型共享一组公共的“大致”属性,但任何一个模型的性能都强烈依赖于在该模型的特定实例化中的参数的选择--例如,每层单元的数量、汇集核的大小、归一化操作中的指数等。由于这种参数(显式或隐式)的数量通常很大,并且评估一个特定参数集的计算成本很高,所以可能的模型实例化的空间在很大程度上没有被探索。因此,当一个模型无法接近生物视觉系统的能力时,我们就不确定这种失败是因为我们缺少一个基本的想法,还是因为正确的“部件”没有被正确地调整,没有以足够的规模组装,或者没有得到足够的训练。在这里,我们提出了一种利用流处理硬件(高端NVIDIA显卡和PlayStation 3‘S IBM Cell处理器)的最新进展来探索此类参数集的高吞吐量方法。类似于分子生物学和遗传学中的高通量筛选方法,我们探索了数千种潜在的网络结构和参数实例,筛选出具有良好对象识别性能的网络结构和参数实例,以供进一步分析。我们表明,这种方法可以在一系列基本对象识别任务中产生显著的、可重复的性能提升,始终优于文献中各种最先进的专门构建的视觉系统。随着可用计算能力的规模继续扩大,我们认为这种方法有可能极大地加快人工视觉和我们对生物视觉的计算基础的理解的进展。理解生物视觉的计算基础的主要障碍之一是其庞大的规模--视觉系统是一台由数十亿个元素组成的大规模并行计算机。虽然这种规模在历史上甚至超出了最快的超级计算系统的能力范围,但最近商用图形处理器(如PlayStation3和高端NVIDIA显卡中的处理器)的进步使前所未有的计算资源得以广泛使用。在这里,我们描述了一种高通量的方法,它利用现代图形硬件的能力来搜索大规模的、受生物启发的视觉系统候选模型的巨大空间。这些模型中最好的是从数千名候选人中挑选出来的,在一系列物体和人脸识别任务中表现优于各种最先进的视觉系统。我们认为,这些实验指出了一条新的前进道路,无论是在创建机器视觉系统方面,还是在提供对生物视觉的计算基础的洞察方面。
While many models of biological object recognition share a common set of “broad-stroke” properties, the performance of any one model depends strongly on the choice of parameters in a particular instantiation of that model—e.g., the number of units per layer, the size of pooling kernels, exponents in normalization operations, etc. Since the number of such parameters (explicit or implicit) is typically large and the computational cost of evaluating one particular parameter set is high, the space of possible model instantiations goes largely unexplored. Thus, when a model fails to approach the abilities of biological visual systems, we are left uncertain whether this failure is because we are missing a fundamental idea or because the correct “parts” have not been tuned correctly, assembled at sufficient scale, or provided with enough training. Here, we present a high-throughput approach to the exploration of such parameter sets, leveraging recent advances in stream processing hardware (high-end NVIDIA graphic cards and the PlayStation 3's IBM Cell Processor). In analogy to high-throughput screening approaches in molecular biology and genetics, we explored thousands of potential network architectures and parameter instantiations, screening those that show promising object recognition performance for further analysis. We show that this approach can yield significant, reproducible gains in performance across an array of basic object recognition tasks, consistently outperforming a variety of state-of-the-art purpose-built vision systems from the literature. As the scale of available computational power continues to expand, we argue that this approach has the potential to greatly accelerate progress in both artificial vision and our understanding of the computational underpinning of biological vision. One of the primary obstacles to understanding the computational underpinnings of biological vision is its sheer scale—the visual system is a massively parallel computer, comprised of billions of elements. While this scale has historically been beyond the reach of even the fastest super-computing systems, recent advances in commodity graphics processors (such as those found in the PlayStation 3 and high-end NVIDIA graphics cards) have made unprecedented computational resources broadly available. Here, we describe a high-throughput approach that harnesses the power of modern graphics hardware to search a vast space of large-scale, biologically inspired candidate models of the visual system. The best of these models, drawn from thousands of candidates, outperformed a variety of state-of-the-art vision systems across a range of object and face recognition tasks. We argue that these experiments point a new way forward, both in the creation of machine vision systems and in providing insights into the computational underpinnings of biological vision.
DOI: 10.1023/b:visi.0000029664.99615.94
发表时间: 2004-11-01
影响因子: 19.5
作者:
Lowe, DG
通讯作者: Lowe, DG
DOI: 10.1007/bf00344251
发表时间: 1980-01-01
影响因子: 1.9
作者:
FUKUSHIMA, K
通讯作者: FUKUSHIMA, K
DOI: 10.1162/neco.1991.3.2.194
发表时间: 1991-06-01
期刊: NEURAL COMPUTATION
影响因子: 2.9
作者:
Foldiak, Peter
通讯作者: Foldiak, Peter
DOI: 10.1038/nn1519
发表时间: 2005-09-01
影响因子: 25
作者:
Cox, DD;Meier, P;DiCarlo, JJ
通讯作者: DiCarlo, JJ
DOI: 10.1007/s00422-005-0585-8
发表时间: 2005-07-01
影响因子: 1.9
作者:
Einhäuser, W;Hipp, J;König, P
通讯作者: König, P