High-Resolution Neural Network for Driver Visual Attention Prediction

High-Resolution Neural Network for Driver Visual Attention Prediction
复制标题

DOI:
10.3390/s20072030
复制
发表时间:
2020-04-01
期刊:
影响因子:
3.9
通讯作者:
Lee, Yeejin
Lee, Yeejin
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Kang, Byeongkeun;Lee, Yeejin

文献摘要

被引文献

相似文献

驾驶是一项对视觉信息有很高要求的任务,因此人类视觉系统在做出正确的安全驾驶决策方面起着至关重要的作用。了解驾驶员的视觉注意力和相关行为信息是高级驾驶员辅助系统(ADAS)和高效自动驾驶汽车(AV)中一项具有挑战性但至关重要的任务。具体来说,从图像中对驾驶员注意力的鲁棒预测可能是帮助智能车辆系统的关键,在智能车辆系统中,自动驾驶汽车需要与周围环境进行安全互动。因此,在本文中,我们研究了人类驾驶员的视觉行为的计算机视觉来估计驾驶员的注意力在图像中的位置。首先,我们表明,在高分辨率的特征表示,提高视觉注意力预测精度和定位性能时,融合在低分辨率的功能。为了证明这一点,我们采用了一个深度卷积神经网络框架,该框架可以在多个分辨率下学习和提取特征表示。特别地,网络在原始图像分辨率下保持具有最高分辨率的特征表示。其次,当使用典型的视觉注意力数据集训练神经网络时,注意力预测往往偏向图像中心。为了避免过度拟合到中心偏置的解决方案,网络使用不同的图像区域进行训练。最后,实验结果验证了我们提出的框架提高了驾驶员注意力位置的预测精度。
Driving is a task that puts heavy demands on visual information, thereby the human visual system plays a critical role in making proper decisions for safe driving. Understanding a driver's visual attention and relevant behavior information is a challenging but essential task in advanced driver-assistance systems (ADAS) and efficient autonomous vehicles (AV). Specifically, robust prediction of a driver's attention from images could be a crucial key to assist intelligent vehicle systems where a self-driving car is required to move safely interacting with the surrounding environment. Thus, in this paper, we investigate a human driver's visual behavior in terms of computer vision to estimate the driver's attention locations in images. First, we show that feature representations at high resolution improves visual attention prediction accuracy and localization performance when being fused with features at low-resolution. To demonstrate this, we employ a deep convolutional neural network framework that learns and extracts feature representations at multiple resolutions. In particular, the network maintains the feature representation with the highest resolution at the original image resolution. Second, attention prediction tends to be biased toward centers of images when neural networks are trained using typical visual attention datasets. To avoid overfitting to the center-biased solution, the network is trained using diverse regions of images. Finally, the experimental results verify that our proposed framework improves the prediction accuracy of a driver's attention locations.