Machine Learning and End-to-End Deep Learning for Monitoring Driver Distractions From Physiological and Visual Signals

Machine Learning and End-to-End Deep Learning for Monitoring Driver Distractions From Physiological and Visual Signals
复制标题

DOI:
10.1109/access.2020.2986810
复制
发表时间:
2020-01-01
期刊:
影响因子:
3.9
通讯作者:
Hassan, Teena
Hassan, Teena
中科院分区:
计算机科学3区
文献类型:
--
作者:
Gjoreski, Martin;Gams, Matja S.;Hassan, Teena

文献摘要

被引文献

相似文献

自动驾驶汽车变得无处不在只是时间问题;然而,人类驾驶监督将在几十年内仍然是必要的。为了评估驾驶员在关键场景中控制车辆的能力,可以使用可穿戴传感器或嵌入车辆中的传感器(例如摄像机)来监测驾驶员分心。驾驶分心的类型,可以感觉到与各种传感器是一个开放的研究问题,这项研究试图回答。这项研究比较了来自生理传感器(手掌皮肤电活动(pEDA),心率和呼吸率)和视觉传感器(眼动跟踪,瞳孔直径,鼻EDA(nEDA),情绪激活和面部动作单元(AU))的数据,用于检测四种类型的分心。该数据集是在以前的驾驶模拟研究中收集的。统计测试表明,用于检测驾驶员分心的信息量最大的特征/模态取决于分心的类型,其中情绪激活和AU是最有希望的。对7种经典机器学习(ML)和7种端到端深度学习(DL)方法进行了实验比较,这些方法在10个受试者的单独测试集上进行了评估,结果表明,当将窗口分类为分心或不分心时,极端梯度增强(XGB)分类器使用60秒的AU窗口作为输入,实现了最高的F1分数79%。当对完整的驾驶课程进行分类时,XGB的F1分数为94%。性能最好的DL模型是光谱-时间ResNet,在对片段进行分类时,其F1得分为75%,在对完整驾驶会话进行分类时,其F1得分为87%。最后,本研究确定并讨论了可能对相关ML方法产生不利影响的问题,如标签抖动,场景过拟合和泛化性能不理想。
It is only a matter of time until autonomous vehicles become ubiquitous; however, human driving supervision will remain a necessity for decades. To assess the drive's ability to take control over the vehicle in critical scenarios, driver distractions can be monitored using wearable sensors or sensors that are embedded in the vehicle, such as video cameras. The types of driving distractions that can be sensed with various sensors is an open research question that this study attempts to answer. This study compared data from physiological sensors (palm electrodermal activity (pEDA), heart rate and breathing rate) and visual sensors (eye tracking, pupil diameter, nasal EDA (nEDA), emotional activation and facial action units (AUs)) for the detection of four types of distractions. The dataset was collected in a previous driving simulation study. The statistical tests showed that the most informative feature/modality for detecting driver distraction depends on the type of distraction, with emotional activation and AUs being the most promising. The experimental comparison of seven classical machine learning (ML) and seven end-to-end deep learning (DL) methods, which were evaluated on a separate test set of 10 subjects, showed that when classifying windows into distracted or not distracted, the highest F1-score of 79%; was realized by the extreme gradient boosting (XGB) classifier using 60-second windows of AUs as input. When classifying complete driving sessions, XGB's F1-score was 94%. The best-performing DL model was a spectro-temporal ResNet, which realized an F1-score of 75%; when classifying segments and an F1-score of 87%; when classifying complete driving sessions. Finally, this study identified and discussed problems, such as label jitter, scenario overfitting and unsatisfactory generalization performance, that may adversely affect related ML approaches.