Objects and scenes classification with selective use of central and peripheral image content

Objects and scenes classification with selective use of central and peripheral image content
复制标题

DOI:
10.1016/j.jvcir.2019.102698
复制
发表时间:
2020-01-01
影响因子:
2.6
通讯作者:
Nazarpour, Kianoush
Nazarpour, Kianoush
中科院分区:
计算机科学3区
文献类型:
--
作者:
Alameer, Ali;Degenaar, Patrick;Nazarpour, Kianoush

文献摘要

被引文献

相似文献

人类视觉识别系统比当前任何机器人视觉设置都更有效。这种优越性的原因之一是人类根据识别任务使用不同的视野。例如,对人类受试者的实验表明,在识别场景时,周边视觉比中央视觉更有用。我们在识别物体和场景方面测试了我们最近开发的模型,即弹性网络正则化分层 MAX (En-HMAX)。在各种实验条件下,图像被不同大小的窗口和暗点遮挡。使用该模型,物体和场景的分类准确率可以达到 90%。 Modelling human experiments, window and scotoma analysis with the En-HMAX model revealed that object and scene recognition are sensitive to the availability of data in the centre and the periphery of the images, respectively. Similarly, results of deep learning models have shown that the classification accuracy diminishes dramatically in the absence of the peripheral vision. These differences led us to further analyse the performance of the En-HMAX model with the parafoveal versus peripheral areas of vision, in a second study. Results of the second study show that approximately 50% of the visual field would be sufficient to achieve 96% accuracy in the classification of unseen images. En-HMAX 模型根据图像类别采用类似于人类视觉系统的相对重要性顺序。我们表明,利用相关的视觉区域可以显着减少图像处理时间和尺寸。 (C) 2019 Elsevier Inc. 保留所有权利。
The human visual recognition system is more efficient than any current robotic vision setting. One reason for this superiority is that humans utilize different fields of vision, depending on the recognition task. For instance, experiments on human subjects show that the peripheral vision is more useful than the central vision in recognizing scenes. We tested our recently-developed model, that is, the elastic net-regularized hierarchical MAX (En-HMAX), in recognizing objects and scenes. In various experimental conditions, images were occluded with windows and scotomas of varying sizes. With this model, classification accuracies of up to 90% for objects and scenes were possible. Modelling human experiments, window and scotoma analysis with the En-HMAX model revealed that object and scene recognition are sensitive to the availability of data in the centre and the periphery of the images, respectively. Similarly, results of deep learning models have shown that the classification accuracy diminishes dramatically in the absence of the peripheral vision. These differences led us to further analyse the performance of the En-HMAX model with the parafoveal versus peripheral areas of vision, in a second study. Results of the second study show that approximately 50% of the visual field would be sufficient to achieve 96% accuracy in the classification of unseen images. The En-HMAX model adopts a relative order of importance, similar to the human visual system, depending on the image category. We showed that utilizing the relevant regions of vision can significantly reduce the image processing time and size. (C) 2019 Elsevier Inc. All rights reserved.