Block-Based High Performance CNN Architectures for Frame-Level Overlapping Speech Detection

Block-Based High Performance CNN Architectures for Frame-Level Overlapping Speech Detection
复制标题

DOI:
10.1109/taslp.2020.3036237
复制
发表时间:
2021-01-01
影响因子:
5.4
通讯作者:
Hansen, John H. L.
Hansen, John H. L.
中科院分区:
计算机科学2区
文献类型:
--
作者:
Yousefi, Midia;Hansen, John H. L.

文献摘要

被引文献

相似文献

由于深度学习技术的出现,自动语音识别 (ASR)、说话人分类、说话人识别和语音合成等语音技术系统取得了显着进步。然而,这些支持语音的系统在自然环境条件下都表现不佳,特别是在涉及一个或多个潜在干扰说话者的情况下。因此,重叠语音检测已成为语音技术应用的重要前端分流步骤。这对于无法进行手动标记的大规模数据集至关重要。提出了基于块的 CNN 架构来解决帧短至 25 ms 的音频流中重叠语音的建模问题。所提出的架构对于以下两方面都具有鲁棒性:(i)由于训练期间网络参数的变化而导致的网络激活分布的变化,(ii)由特征提取、环境噪声或房间干扰引起的输入特征的局部变化。我们还研究了替代输入特征(包括谱幅度、MFCC、MFB 和比重图)对计算时间和分类性能的影响。基于GRID语料库对模拟重叠语音信号进行评估。实验结果突显了该系统在检测重叠语音帧方面的能力,对同性别重叠语音的准确度为 90.5%,精确度为 93.5%,召回率为 92.7%,Fscore 为 92.8%。对于异性案例,所有分类指标的网络得分均超过 95%。
Speech technology systems such as Automatic Speech Recognition (ASR), speaker diarization, speaker recognition, and speech synthesis have advanced significantly by the emergence of deep learning techniques. However, none of these voice-enabled systems perform well in natural environmental circumstances, specifically in situations where one or more potential interfering talkers are involved. Therefore, overlapping speech detection has become an important front-end triage step for speech technology applications. This is crucial for large-scale datsets where manual labeling in not possible. A block-based CNN architecture is proposed to address modeling overlapping speech in audio streams with frames as short as 25 ms. The proposed architecture is robust to both: (i) shifts in distribution of network activations due to the change in network parameters during training, (ii) local variations from the input features caused by feature extraction, environmental noise, or room interference. We also investigate the effect of alternate input features including spectral magnitude, MFCC, MFB, and pyknogram on both computational time and classification performance. Evaluation is performed on simulated overlapping speech signals based on the GRID corpus. The experimental results highlight the capability of the proposed system in detecting overlapping speech frames with 90.5% accuracy, 93.5% precision, 92.7% recall, and 92.8% Fscore on same gender overlapped speech. For opposite gender cases, the network scores exceed 95% in all the classification metrics.