A Refined Non-Driving Activity Classification Using a Two-Stream Convolutional Neural Network

A Refined Non-Driving Activity Classification Using a Two-Stream Convolutional Neural Network
复制标题

DOI:
10.1109/jsen.2020.3005810
复制
发表时间:
2020-06
影响因子:
4.3
通讯作者:
Lichao Yang;Tingyu Yang;Haochen Liu;Xiaocai Shan;J. Brighton;L. Skrypchuk;Alexandros;Mouzakitis;Yifan Zhao
Lichao Yang;Tingyu Yang;Haochen Liu;Xiaocai Shan;J. Brighton;L. Skrypchuk;Alexandros;Mouzakitis;Yifan Zhao
中科院分区:
综合性期刊2区
文献类型:
--
作者:
Lichao Yang;Tingyu Yang;Haochen Liu;Xiaocai Shan;J. Brighton;L. Skrypchuk;Alexandros;Mouzakitis;Yifan Zhao

文献摘要

被引文献

相似文献

对驾驶员状态进行监测,对于实现三级自动驾驶车辆的智能安全交接过渡具有重要意义。我们提出了一种基于相机的系统来识别非驾驶活动(NDA),这些非驾驶活动可能导致基于空间和时间信息融合的不同认知能力的接管。基于所提取的驾驶员和与之交互的对象/设备的掩码自动选择感兴趣区域(ROI)。然后,将感兴趣区域(空间流)的RGB图像及其关联的当前和历史光流帧(时间流)馈送到用于NDA分类的双流卷积神经网络(CNN)。这种方法不仅能够识别对象/设备,而且能够识别对象和驱动程序之间的交互模式,这使得能够进行精细的NDA分类。在这篇文章中,我们评估了10名参与者使用两种类型的设备(平板电脑和手机)和5种类型的任务(电子邮件、阅读、视频、网页浏览和游戏)对10个NDA进行分类的性能。实验结果表明,该系统的平均分类正确率从单一空间流的61.0%提高到90.5%。
It is of great importance to monitor the driver’s status to achieve an intelligent and safe take-over transition in the level 3 automated driving vehicle. We present a camera-based system to recognise the non-driving activities (NDAs) which may lead to different cognitive capabilities for take-over based on a fusion of spatial and temporal information. The region of interest (ROI) is automatically selected based on the extracted masks of the driver and the object/device interacting with. Then, the RGB image of the ROI (the spatial stream) and its associated current and historical optical flow frames (the temporal stream) are fed into a two-stream convolutional neural network (CNN) for the classification of NDAs. Such an approach is able to identify not only the object/device but also the interaction mode between the object and the driver, which enables a refined NDA classification. In this paper, we evaluated the performance of classifying 10 NDAs with two types of devices (tablet and phone) and 5 types of tasks (emailing, reading, watching videos, web-browsing and gaming) for 10 participants. Results show that the proposed system improves the averaged classification accuracy from 61.0% when using a single spatial stream to 90.5%.