Speech Acoustic Modelling from Raw Phase Spectrum

Speech Acoustic Modelling from Raw Phase Spectrum
复制标题

DOI:
10.1109/icassp39728.2021.9413727
复制
发表时间:
2021-06
期刊:
ICASSP 2021 - 2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
影响因子:
--
通讯作者:
Erfan Loweimi;Z. Cvetković;P. Bell;S. Renals
Erfan Loweimi;Z. Cvetković;P. Bell;S. Renals
中科院分区:
其他
文献类型:
--
作者:
Erfan Loweimi;Z. Cvetković;P. Bell;S. Renals

文献摘要

被引文献

相似文献

基于幅度谱的特征是自动语音识别 (ASR) 系统中声学建模中使用最广泛的前端。在本文中,我们研究了使用原始短时相位谱进行声学建模的可能性和有效性。特别是,我们研究了原始包裹、展开和最小相位相位谱以及源和滤波器组件的相位在声学建模中的有用性。此外,我们还探索了使用多头 CNN 同时部署声道和原始相位谱的激励分量的有效性,并研究了多种信息融合方案。这为开发用于语音识别的有效的基于相位的多流信息处理系统铺平了道路。即使对于具有类似噪声形状的包裹相位,其性能也可以与基于幅度的经典特征相媲美或更好,并且在 WSJ (Eval-92) 任务中实现了高达 4.8% 的 WER。
Magnitude spectrum-based features are the most widely employed front-ends for acoustic modelling in automatic speech recognition (ASR) systems. In this paper, we investigate the possibility and efficacy of acoustic modelling using the raw short-time phase spectrum. In particular, we study the usefulness of the raw wrapped, unwrapped and minimum-phase phase spectra as well as the phase of the source and filter components for acoustic modelling. Furthermore, we explore the effectiveness of simultaneous deployment of the vocal tract and excitation components of the raw phase spectrum using multi-head CNNs and investigate multiple information fusion schemes. This paves the way for developing an effective phase-based multi-stream information processing systems for speech recognition. The performance, even for wrapped phase with a noise-like shape, is comparable to or better than the magnitude-based classic features, and up to 4.8% WER has been achieved in the WSJ (Eval-92) task.