A Comparative Study of Speaker Role Identification in Air Traffic Communication Using Deep Learning Approaches

A Comparative Study of Speaker Role Identification in Air Traffic Communication Using Deep Learning Approaches
复制标题

DOI:
10.1145/3572792
复制
发表时间:
2021-11
影响因子:
2
通讯作者:
Dongyue Guo;Jianwei Zhang;Bo Yang;Yi Lin
Dongyue Guo;Jianwei Zhang;Bo Yang;Yi Lin
中科院分区:
计算机科学4区
文献类型:
--
作者:
Dongyue Guo;Jianwei Zhang;Bo Yang;Yi Lin

文献摘要

相似文献

空中交通管制(ATC)中的飞行员-驾驶员对话的自动口语指令理解(SIU)不仅需要识别语音的单词和语义,还需要确定说话者的角色。然而,在空中交通通信中的自动理解系统的研究中,很少有人关注说话人角色识别(SRI)。在这篇文章中,我们制定了SRI任务的飞行员导频通信作为一个二元分类问题。此外,基于文本的,基于语音的,语音和文本的多模态方法被提出来实现SRI任务的综合比较。为了消除比较方法的影响,应用各种先进的神经网络架构来优化基于文本和基于语音的方法的实现。最重要的是,一个多模态说话人角色识别网络(MMSRINet)的设计,以实现SRI任务,同时考虑语音和文本的模态特征。为了聚合模态特征,提出了模态融合模块,分别通过模态注意机制和自注意池层来融合和挤压声学和文本表示。最后,比较方法进行了验证的ATC语音语料库收集从现实世界的ATC环境。实验结果表明,所有的比较方法工作的SRI任务,所提出的MMSRINet表现出竞争力的性能和鲁棒性相比,与其他方法的可见和不可见的数据,分别达到98.56%和98.08%的准确率。
Automatic spoken instruction understanding (SIU) of the controller-pilot conversations in the air traffic control (ATC) requires not only recognizing the words and semantics of the speech but also determining the role of the speaker. However, few of the published works on the automatic understanding systems in air traffic communication focus on speaker role identification (SRI). In this article, we formulate the SRI task of controller-pilot communication as a binary classification problem. Furthermore, the text-based, speech-based, and speech-and-text-based multi-modal methods are proposed to achieve a comprehensive comparison of the SRI task. To ablate the impacts of the comparative approaches, various advanced neural network architectures are applied to optimize the implementation of text-based and speech-based methods. Most importantly, a multi-modal speaker role identification network (MMSRINet) is designed to achieve the SRI task by considering both the speech and textual modality features. To aggregate modality features, the modal fusion module is proposed to fuse and squeeze acoustic and textual representations by modal attention mechanism and self-attention pooling layer, respectively. Finally, the comparative approaches are validated on the ATCSpeech corpus collected from a real-world ATC environment. The experimental results demonstrate that all the comparative approaches worked for the SRI task, and the proposed MMSRINet shows competitive performance and robustness compared with the other methods on both seen and unseen data, achieving 98.56% and 98.08% accuracy, respectively.