Real-time Context-Aware Multimodal Network for Activity and Activity-Stage Recognition from Team Communication in Dynamic Clinical Settings

Real-time Context-Aware Multimodal Network for Activity and Activity-Stage Recognition from Team Communication in Dynamic Clinical Settings
复制标题

实时上下文感知多模态网络,用于动态临床环境中团队沟通的活动和活动阶段识别

DOI:
10.1145/3580798
复制
发表时间:
2022
期刊:
Wearable and Ubiquitous Technologies
影响因子:
--
通讯作者:
Burd, Randall S.
Burd, Randall S.
中科院分区:
--
文献类型:
--
作者:
Gao, Chenyang;Marsic, Ivan;Sarcevic, Aleksandra;Gestrich-Thompson, Waverly;Burd, Randall S.

文献摘要

参考文献

相似文献

在临床环境中,大多数自动识别系统使用视觉或感官数据来识别活动。这些系统无法识别依赖于口头评估、缺乏视觉提示或不使用医疗设备的活动。我们在临床领域研究了基于语音的活动和活动阶段识别,做出了以下贡献。(1)我们收集了一个高质量的数据集,代表了在实际创伤复苏事件中的常见活动和活动阶段,即危重患者的初步评估和治疗。(2)我们引入了一种基于音频信号和一组关键词的新型多模式网络,该网络不需要高性能自动语音识别(ASR)引擎。(3)我们设计了新颖的上下文模块,以捕获复杂工作流程中有关活动和阶段的团队对话中的动态依赖关系。(4)我们介绍了一种数据增强方法,该方法通过结合选定的话语及其音频片段来模拟团队通信,并表明该方法有助于在我们的数据有限的情况下提高性能。在离线实验中,我们提出的上下文感知多模态模型在活动和活动阶段识别方面的F1分数分别为73.2±0.8%和78.1±1.1%。在在线实验中,当使用ASR输出的话语级分割时,两种识别类型的性能下降约10%。当我们忽略话语级分割时,性能下降了约15%。我们的实验表明,基于语音的活动和活动阶段识别在动态临床事件的可行性。
In clinical settings, most automatic recognition systems use visual or sensory data to recognize activities. These systems cannot recognize activities that rely on verbal assessment, lack visual cues, or do not use medical devices. We examined speech-based activity and activity-stage recognition in a clinical domain, making the following contributions. (1) We collected a high-quality dataset representing common activities and activity stages during actual trauma resuscitation events-the initial evaluation and treatment of critically injured patients. (2) We introduced a novel multimodal network based on audio signal and a set of keywords that does not require a high-performing automatic speech recognition (ASR) engine. (3) We designed novel contextual modules to capture dynamic dependencies in team conversations about activities and stages during a complex workflow. (4) We introduced a data augmentation method, which simulates team communication by combining selected utterances and their audio clips, and showed that this method contributed to performance improvement in our data-limited scenario. In offline experiments, our proposed context-aware multimodal model achieved F1-scores of 73.2±0.8% and 78.1±1.1% for activity and activity-stage recognition, respectively. In online experiments, the performance declined about 10% for both recognition types when using utterance-level segmentation of the ASR output. The performance declined about 15% when we omitted the utterance-level segmentation. Our experiments showed the feasibility of speech-based activity and activity-stage recognition during dynamic clinical events.
DOI: 10.1145/3512920
发表时间: 2022-03
影响因子: --
作者:
Swathi Jagannath;Neha Kamireddi;K. A. Zellner;R. Burd;I. Marsic;Aleksandra Sarcevic
通讯作者: Swathi Jagannath;Neha Kamireddi;K. A. Zellner;R. Burd;I. Marsic;Aleksandra Sarcevic
DOI: 10.1109/ichi48887.2020.9374372
发表时间: 2020-11
期刊: Proceedings. IEEE International Conference on Healthcare Informatics
影响因子: --
作者:
Abdulbaqi J;Gu Y;Xu Z;Gao C;Marsic I;Burd RS
通讯作者: Burd RS
使用视频审核来改善创伤复苏——是时候采取新方法了。
DOI: --
发表时间: 2006
期刊: Canadian journal of surgery. Journal canadien de chirurgie
影响因子: --
作者:
M. Fitzgerald;Robert A Gocentas;L. Dziukas;P. Cameron;C. Mackenzie;N. Farrow
通讯作者: N. Farrow
DOI: 10.1109/taslp.2021.3122291
发表时间: 2021-01-01
影响因子: 5.4
作者:
Hsu, Wei-Ning;Bolte, Benjamin;Mohamed, Abdelrahman
通讯作者: Mohamed, Abdelrahman
DOI: 10.1097/00005373-200206000-00009
发表时间: 2002-06-01
影响因子: --
作者:
Holcomb, JB;Dumire, RD;Mattox, KL
通讯作者: Mattox, KL