Raw Sign and Magnitude Spectra for Multi-Head Acoustic Modelling

Raw Sign and Magnitude Spectra for Multi-Head Acoustic Modelling
复制标题

DOI:
10.21437/interspeech.2020-0018
复制
发表时间:
2020-10
期刊:
--
影响因子:
--
通讯作者:
Erfan Loweimi;P. Bell;S. Renals
Erfan Loweimi;P. Bell;S. Renals
中科院分区:
其他
文献类型:
--
作者:
Erfan Loweimi;P. Bell;S. Renals

文献摘要

被引文献

相似文献

在本文中,我们研究了符号谱及其与原始幅度谱的组合在自动语音识别(ASR)声学建模中的有用性。符号谱是±1秒的序列,捕获相位谱的一位。它对幅度谱忽略的信息进行编码,从而实现独特的信号表征和重建。特别是,我们证明它携带与信号时间结构以及语音源成分相关的信息。此外,我们还研究了通过多头 CNN 在不同融合级别将其与原始震级谱相结合对于 ASR 的有用性。虽然从信息角度来看,这两个信息流相当于原始波形信号,但整体性能明显高于原始波形和 MFCC 和滤波器组等经典功能。这已在 TIMIT、NTIMT、Aurora-4 和 WSJ 任务中观察到并得到验证,并且相对 WER 降低了高达 14.5%。
In this paper we investigate the usefulness of the sign spectrum and its combination with the raw magnitude spectrum in acoustic modelling for automatic speech recognition (ASR). The sign spectrum is a sequence of ± 1 s, capturing one bit of the phase spectrum. It encodes information overlooked by the magnitude spectrum enabling unique signal characterisation and reconstruction. In particular, we demonstrate it carries information related to the temporal structure of the signal as well as the speech’s source component. Furthermore, we investigate the usefulness of combining it with the raw magnitude spectrum via multi-head CNNs at different fusion levels for ASR. While information-wise these two streams of information are together equivalent to the raw waveform signal the overall performance is noticeably higher than raw waveform and classic features such as MFCC and filterbank. This has been observed and verified in TIMIT, NTIMT, Aurora-4 and WSJ tasks and up to 14.5% relative WER reduction has been achieved.