RASTA Processing of Speech

RASTA Processing of Speech
复制标题

DOI:
10.1109/89.326616
复制
发表时间:
1994-10-01
期刊:
IEEE TRANSACTIONS ON SPEECH AND AUDIO PROCESSING
影响因子:
--
通讯作者:
Morgan, Nelson
Morgan, Nelson
中科院分区:
其他
文献类型:
--
作者:
Hermansky, Hynek;Morgan, Nelson

文献摘要

被引文献

相似文献

即使是目前最好的随机识别器的性能在意外的通信环境中也会严重下降。在某些情况下,环境效应可以通过一组简单的变换来建模,特别是通过与环境脉冲响应和一些环境噪声的卷积来建模。通常,这些环境影响的时间特性与言语的时间特性大不相同。我们一直在试验滤波方法,试图利用这些差异来产生用于语音识别和增强的鲁棒表示,并将这类表示称为相对光谱(RASTA)。在本文中,我们回顾了该方法的理论和实验基础,讨论了与人类听觉感知的关系,并将原始方法扩展到加性噪声和卷积噪声的组合。我们讨论了RASTA特征与所需识别模型的性质之间的关系,以及这些特征与δ特征和倒谱均值减法之间的关系。最后,我们展示了RASTA技术在语音增强中的应用。
Performance of even the best current stochastic recognizers severely degrades in an unexpected communications environment. In some cases, the environmental effect can he modeled by a set of simple transformations and, in particular, by convolution with an environmental impulse response and the addition of some environmental noise. Often, the temporal properties of these environmental effects are quite different from the temporal properties of speech. We have been experimenting with filtering approaches that attempt to exploit these differences to produce robust representations for speech recognition and enhancement and have called this class of representations relative spectra (RASTA). In this paper, we review the theoretical and experimental foundations of the method, discuss the relationship with human auditory perception, and extend the original method to combinations of additive noise and convolutional noise. We discuss the relationship between RASTA features and the nature of the recognition models that are required and the relationship of these features to delta features and to cepstral mean subtraction. Finally, we show an application of the RASTA technique to speech enhancement.