SPEECH TASKS RELEVANT TO SLEEPINESS DETERMINED WITH DEEP TRANSFER LEARNING.

SPEECH TASKS RELEVANT TO SLEEPINESS DETERMINED WITH DEEP TRANSFER LEARNING.
复制标题

通过深度迁移学习确定与睡意相关的语音任务。

DOI:
10.1109/icassp43922.2022.9747000
复制
发表时间:
2022
期刊:
Proceedings of the ... IEEE International Conference on Acoustics, Speech, and Signal Processing. ICASSP (Conference)
影响因子:
--
通讯作者:
Warrenburg,LindsayA
Warrenburg,LindsayA
中科院分区:
--
文献类型:
--
作者:
Tran,Bang;Zhu,Youxiang;Liang,Xiaohui;Schwoebel,JamesW;Warrenburg,LindsayA

文献摘要

相似文献

在注意力紧张的情况下过度嗜睡可能会导致不良事件,如车祸。检测和监测嗜睡状态有助于防止这些不良事件的发生。在本文中,我们使用Voiceome数据集从1828名参与者中提取语音来建立一个深度迁移学习模型,该模型使用HUBERT(HUBERT)语音表示来检测个体的困倦。在睡眠检测中,语音是一种未得到充分利用的数据源,但由于语音采集简单、经济、非侵入性,它为睡眠检测提供了一种有前途的资源。为了寻求关于个别言语任务重要性的一致证据,研究人员采用了两种互补的技术。我们的第一项技术,掩蔽,通过组合所有语音任务,掩蔽语音中选定的响应,并观察模型精度的系统变化来评估任务重要性。我们的第二种技术,单独训练,比较了多个模型的准确性,每个模型使用相同的架构,但在不同的语音任务子集上进行训练。我们的评估表明,性能最好的模型利用了波士顿命名测验中的记忆回忆任务和范畴命名任务,准确率分别为80.07%(F1分数为0.85)和81.13%(F1分数为0.89)。
Excessive sleepiness in attention-critical contexts can lead to adverse events, such as car crashes. Detecting and monitoring sleepiness can help prevent these adverse events from happening. In this paper, we use the Voiceome dataset to extract speech from 1,828 participants to develop a deep transfer learning model using Hidden-Unit BERT (HuBERT) speech representations to detect sleepiness from individuals. Speech is an under-utilized source of data in sleep detection, but as speech collection is easy, cost-effective, and non-invasive, it provides a promising resource for sleepiness detection. Two complementary techniques were conducted in order to seek converging evidence regarding the importance of individual speech tasks. Our first technique, masking, evaluated task importance by combining all speech tasks, masking selected responses in the speech, and observing systematic changes in model accuracy. Our second technique, separate training, compared the accuracy of multiple models, each of which used the same architecture, but was trained on a different subset of speech tasks. Our evaluation shows that the best-performing model utilizes the memory recall task and categorical naming task from the Boston Naming Test, which achieved an accuracy of 80.07% (F1-score of 0.85) and 81.13% (F1-score of 0.89), respectively.