Speech Acoustic Modelling Using Raw Source and Filter Components

Speech Acoustic Modelling Using Raw Source and Filter Components
复制标题

DOI:
10.21437/interspeech.2021-53
复制
发表时间:
2021-08
期刊:
--
影响因子:
--
通讯作者:
Erfan Loweimi;Z. Cvetković;P. Bell;S. Renals
Erfan Loweimi;Z. Cvetković;P. Bell;S. Renals
中科院分区:
其他
文献类型:
--
作者:
Erfan Loweimi;Z. Cvetković;P. Bell;S. Renals

文献摘要

被引文献

相似文献

源滤波器建模是语音处理中的基本技术之一,具有广泛的应用。在声学建模中,广泛采用参数化滤波器组件的MFCC和PLP等特征。在本文中,我们研究了从原始滤波器和源组件构建声学模型的效率。原始幅度谱作为原始信息流,通过倒谱提升分解为激励和声道信息流。然后,通过多头CNN构建声学模型,其中包括允许通过一系列定制变换处理每个单独的流,并将它们融合在最佳抽象级别。我们讨论了这种信息分解和重组的可能优势,研究这些模型的动力学,并探索最佳融合水平。此外,我们还说明了CNN的学习滤波器,并为捕获的模式提供了一些解释。在WSJ和Aurora-4任务中,所提出的方法与最佳融合方案的相对WER降低高达14%和7%。
Source-filter modelling is among the fundamental techniques in speech processing with a wide range of applications. In acoustic modelling, features such as MFCC and PLP which parametrise the filter component are widely employed. In this paper, we investigate the efficacy of building acoustic models from the raw filter and source components. The raw magnitude spectrum, as the primary information stream, is decomposed into the excitation and vocal tract information streams via cepstral liftering. Then, acoustic models are built via multi-head CNNs which, among others, allow for processing each individual stream via a sequence of bespoke transforms and fusing them at an optimal level of abstraction. We discuss the possible advantages of such information factorisation and recombination, investigate the dynamics of these models and explore the optimal fusion level. Furthermore, we illustrate the CNN’s learned filters and provide some interpretation for the captured patterns. The proposed approach with optimal fusion scheme results in up to 14% and 7% relative WER reduction in WSJ and Aurora-4 tasks.