Multi-Stream Acoustic Modelling Using Raw Real and Imaginary Parts of the Fourier Transform

Multi-Stream Acoustic Modelling Using Raw Real and Imaginary Parts of the Fourier Transform
复制标题

DOI:
10.1109/taslp.2023.3237167
复制
发表时间:
2023
期刊:
IEEE/ACM Transactions on Audio, Speech, and Language Processing
影响因子:
--
通讯作者:
Erfan Loweimi;Zhengjun Yue;P. Bell;S. Renals;Z. Cvetković
Erfan Loweimi;Zhengjun Yue;P. Bell;S. Renals;Z. Cvetković
中科院分区:
其他
文献类型:
--
作者:
Erfan Loweimi;Zhengjun Yue;P. Bell;S. Renals;Z. Cvetković

文献摘要

被引文献

相似文献

在本文中,我们研究了使用语音信号的傅立叶变换的原始实部和虚部的多流声学建模。使用原始星等谱或由其导出的特征作为实部和虚部的代理会导致不可逆的信息丢失和次优的信息融合。我们讨论并量化这些信息在语音质量和可理解性方面的重要性。在提出的框架中,实部和虚部被视为两个信息流,通过单独的卷积网络进行预处理,然后在最佳抽象级别进行组合,随后通过循环和全连接层进行进一步的后处理。分析了不同体系结构中信息融合的最佳水平、交叉熵损失、帧分类精度和WER方面的训练动态以及在单流和多流模型的第一卷积层学习到的滤波器的形状和性质。我们研究了所提出的系统在各种任务中的有效性:TIMIT/NTIMIT(电话识别),Aurora-4(噪声鲁棒性),WSJ(阅读语音),AMI(会议)和TORGO(困难语音)。在所有任务中,我们都取得了具有竞争力的表现:在极光4中,平均WER降至4.6%,在WSJ中,Eval-92和Eval-93的WER降至4.6%和6.2%,在AMI-IHM的Dev/Eval集中,WER降至23.3%/23.8%,在AMI-SDM中,WER降至43.7%/47.6%。在TORGO中,对于困难言语和典型言语,我们分别实现了31.7%和10.2%的wer。
In this paper, we investigate multi-stream acoustic modelling using the raw real and imaginary parts of the Fourier transform of speech signals. Using the raw magnitude spectrum, or features derived from it, as a proxy for the real and imaginary parts leads to irreversible information loss and suboptimal information fusion. We discuss and quantify the importance of such information in terms of speech quality and intelligibility. In the proposed framework, the real and imaginary parts are treated as two streams of information, pre-processed via separate convolutional networks, and then combined at an optimal level of abstraction, followed by further post-processing via recurrent and fully-connected layers. The optimal level of information fusion in various architectures, training dynamics in terms of cross-entropy loss, frame classification accuracy and WER as well as the shape and properties of the filters learned in the first convolutional layer of single- and multi-stream models are analysed. We investigated the effectiveness of the proposed systems in various tasks: TIMIT/NTIMIT (phone recognition), Aurora-4 (noise robustness), WSJ (read speech), AMI (meeting) and TORGO (dysarthric speech). Across all tasks we achieved competitive performance: in Aurora-4, down to 4.6% WER on average, in WSJ down to 4.6% and 6.2% WERs for Eval-92 and Eval-93, for Dev/Eval sets of the AMI-IHM down to 23.3%/23.8% WERs and in the AMI-SDM down to 43.7%/47.6% WERs have been achieved. In TORGO, for dysarthric and typical speech we achieved down to 31.7% and 10.2% WERs, respectively.