Static and Dynamic Variance Compensation for Recognition of Reverberant Speech With Dereverberation Preprocessing

Static and Dynamic Variance Compensation for Recognition of Reverberant Speech With Dereverberation Preprocessing
复制标题

DOI:
10.1109/tasl.2008.2010214
复制
发表时间:
2009-02
期刊:
IEEE Transactions on Audio, Speech, and Language Processing
影响因子:
--
通讯作者:
Marc Delcroix;T. Nakatani;Shinji Watanabe
Marc Delcroix;T. Nakatani;Shinji Watanabe
中科院分区:
其他
文献类型:
--
作者:
Marc Delcroix;T. Nakatani;Shinji Watanabe

文献摘要

被引文献

相似文献

自动语音识别的性能在噪声或混响的存在下会严重下降。对噪声鲁棒性进行了大量的研究。相比之下,混响语音的识别问题受到的关注要少得多,仍然非常具有挑战性。在本文中,我们使用一个去混响的方法来减少混响识别之前。这样的预处理器可以去除大多数混响效果。然而,它经常引入失真,导致语音特征与用于识别的声学模型之间的动态失配。模型适应可以用来减少这种不匹配。然而,传统的模型自适应技术假设静态失配,因此可能无法很好地科普由去混响引起的动态失配。本文提出了一种新的自适应方案,能够管理静态和动态失配。我们引入了一个参数模型方差适应,包括静态和动态组件,以实现适当的去混响和语音识别器之间的互连。使用期望最大化算法实现的自适应训练来优化模型参数。实验表明,使用所提出的方法与混响语音的混响时间为0.5秒,它是可能的,以实现80%的相对错误率减少与去混响语音识别(字错误率为31%)相比,和最终的错误率为5.4%,这是通过结合所提出的方差补偿和MLLR适应。
The performance of automatic speech recognition is severely degraded in the presence of noise or reverberation. Much research has been undertaken on noise robustness. In contrast, the problem of the recognition of reverberant speech has received far less attention and remains very challenging. In this paper, we use a dereverberation method to reduce reverberation prior to recognition. Such a preprocessor may remove most reverberation effects. However, it often introduces distortion, causing a dynamic mismatch between speech features and the acoustic model used for recognition. Model adaptation could be used to reduce this mismatch. However, conventional model adaptation techniques assume a static mismatch and may therefore not cope well with a dynamic mismatch arising from dereverberation. This paper proposes a novel adaptation scheme that is capable of managing both static and dynamic mismatches. We introduce a parametric model for variance adaptation that includes static and dynamic components in order to realize an appropriate interconnection between dereverberation and a speech recognizer. The model parameters are optimized using adaptive training implemented with the expectation maximization algorithm. An experiment using the proposed method with reverberant speech for a reverberation time of 0.5 s revealed that it was possible to achieve an 80% reduction in the relative error rate compared with the recognition of dereverberated speech (word error rate of 31%), and the final error rate was 5.4%, which was obtained by combining the proposed variance compensation and MLLR adaptation.