ATCSpeechNet: A multilingual end-to-end speech recognition framework for air traffic control systems

ATCSpeechNet: A multilingual end-to-end speech recognition framework for air traffic control systems
复制标题

DOI:
10.1016/j.asoc.2021.107847
复制
发表时间:
2021-09-05
影响因子:
8.7
通讯作者:
Zhang, Yi
Zhang, Yi
中科院分区:
计算机科学2区
文献类型:
--
作者:
Lin, Yi;Yang, Bo;Zhang, Yi

文献摘要

被引文献

相似文献

本文提出了一个多语种的端到端框架ATCSpeechNet,用于解决空中交通管制(ATC)系统中将通信语音翻译成人类可读文本的问题。在所提出的框架中,我们专注于将多语言自动语音识别(ASR)集成到一个模型中,在该模型中,开发了一个端到端的范例,将语音波形直接转换为文本,而无需任何特征工程或词典。为了弥补ATC挑战(包括多语言,多说话人对话和不稳定的语音速率)导致的手工特征工程的不足,提出了一种语音表示学习(SRL)网络,以从原始波中捕获鲁棒和有区别的语音表示。采用自监督训练策略,从未标记数据中优化SRL网络,并进一步预测语音特征,即,波到特征改进了一种端到端的架构来完成ASR任务,其中应用基于图素的建模单元来解决多语言ASR问题。针对ATC领域中转录样本较少的问题,采用基于掩码预测的无监督方法,通过特征到特征的过程,在未标记数据上对ASR模型的主干网络进行预训练.最后,通过将SRL与ASR集成,以监督的方式制定了端到端的多语言ASR框架,该框架能够在一个模型中将原始波转换为文本,即,波形转文字在ATCPeech语料库上的实验结果表明,该方法在标注语料量很小、资源消耗较少的情况下取得了很高的性能,在58小时的转录语料库上标注错误率仅为4.20%。与基线模型相比,该方法获得了超过100%的相对性能改善,可以进一步提高转录样本的大小。研究结果也证实,所提出的自我学习与训练策略对提升最终绩效有显著的贡献。此外,所提出的框架的有效性也验证了常见的语料库(AISHELL,LibriSpeech,和cv-fr)。更重要的是,所提出的多语言框架,不仅降低了系统的复杂性,但也获得了更高的准确性相比,独立的单语ASR模型。该方法还可以大大降低标注样本的成本,这有利于推进ASR技术的工业应用。(C)2021年由Elsevier B. V.出版
In this paper, a multilingual end-to-end framework, called ATCSpeechNet, is proposed to tackle the issue of translating communication speech into human-readable text in air traffic control (ATC) sys-tems. In the proposed framework, we focus on integrating multilingual automatic speech recognition (ASR) into one model, in which an end-to-end paradigm is developed to convert speech waveforms into text directly, without any feature engineering or lexicon. To compensate the deficiency of handcrafted feature engineering caused by ATC challenges, including multilingual, multispeaker dialog and unstable speech rates, a speech representation learning (SRL) network is proposed to capture robust and discriminative speech representations from raw waves. The self-supervised training strategy is adopted to optimize the SRL network from unlabeled data, and to further predict the speech features, i.e., wave-to-feature. An end-to-end architecture is improved to complete the ASR task, in which a grapheme-based modeling unit is applied to address the multilingual ASR issue. Facing the problem of small transcribed samples in the ATC domain, an unsupervised approach with mask prediction is applied to pretrain the backbone network of the ASR model on unlabeled data by a feature-to -feature process. Finally, by integrating the SRL with ASR, an end-to-end multilingual ASR framework is formulated in a supervised manner, which is able to translate the raw wave into text in one model, i.e., wave-to-text. Experimental results on the ATCSpeech corpus demonstrate that the proposed approach achieves high performance with a very small labeled corpus and less resource consumption, only a 4.20% label error rate on the 58-hour transcribed corpus. Compared to the baseline model, the proposed approach obtains over 100% relative performance improvement which can be further enhanced with increasing size of the transcribed samples. It is also confirmed that the proposed SRL and training strategies make significant contributions to improving the final performance. In addition, the effectiveness of the proposed framework is also validated on common corpora (AISHELL, LibriSpeech, and cv-fr). More importantly, the proposed multilingual framework not only reduces the system complexity but also obtains higher accuracy compared to that of the independent monolingual ASR models. The proposed approach can also greatly reduce the cost of annotating samples, which benefits to advance the ASR technique to industrial applications. (C) 2021 Published by Elsevier B.V.