The Effect of Different Occupational Background Noises on Voice Recognition Accuracy

The Effect of Different Occupational Background Noises on Voice Recognition Accuracy
复制标题

不同职业背景噪声对语音识别准确度的影响

DOI:
10.1115/1.4053521
复制
发表时间:
2022
影响因子:
3.1
通讯作者:
Hu, Boyi
Hu, Boyi
中科院分区:
工程技术4区
文献类型:
--
作者:
Song, Li;Ozkan Yerebakan, Mustafa;Luo, Yue;Amaba, Ben;Swope, William;Hu, Boyi

文献摘要

被引文献

相似文献

语音识别已成为我们生活中不可或缺的一部分,通常用于呼叫中心和虚拟助理。然而,语音识别越来越多地应用于更多的工业用途。这些用例中的每一个都具有独特的特征,可能会影响语音识别的有效性,从而影响工业生产力,性能甚至安全性。其中最突出的是在每个行业中占主导地位的独特背景噪音。不同的机器和不同的工作布局的存在是主要因素。另一个重要的特征是在这些环境中存在的通信类型。日常交流通常涉及在相对安静的条件下说出的较长的句子,而工业环境中的交流通常很短,并且在大声的条件下进行。在这项研究中,我们通过比较两种语音识别算法在几种背景噪声条件下的性能,证明了考虑这两个因素的重要性:基于常规卷积神经网络(CNN)的语音识别算法与基于自动语音识别(ASR)的模型具有去噪模块。我们的研究结果表明,有一个显着的性能下降之间的典型背景噪声使用(白色噪声)和其余的背景噪声。此外,我们的自定义ASR模型与去噪模块的性能优于基于CNN的模型,在所有背景噪声中的整体性能提高了14-35%。这两个结果都证明,需要为这些环境开发专门的语音识别算法,以可靠地将其部署为控制机制。
Voice recognition has become an integral part of our lives, commonly used in call centers and as part of virtual assistants. However, voice recognition is increasingly applied to more industrial uses. Each of these use cases has unique characteristics that may impact the effectiveness of voice recognition, which could impact industrial productivity, performance, or even safety. One of the most prominent among them is the unique background noises that are dominant in each industry. The existence of different machinery and different work layouts are primary contributors to this. Another important characteristic is the type of communication that is present in these settings. Daily communication often involves longer sentences uttered under relatively silent conditions, whereas communication in industrial settings is often short and conducted in loud conditions. In this study, we demonstrated the importance of taking these two elements into account by comparing the performances of two voice recognition algorithms under several background noise conditions: a regular Convolutional Neural Network (CNN)-based voice recognition algorithm to an Auto Speech Recognition (ASR)-based model with a denoising module. Our results indicate that there is a significant performance drop between the typical background noise use (white noise) and the rest of the background noises. Also, our custom ASR model with the denoising module outperformed the CNN-based model with an overall performance increase between 14–35% across all background noises. Both results give proof that specialized voice recognition algorithms need to be developed for these environments to reliably deploy them as control mechanisms.