Speaker and Noise Factorization for Robust Speech Recognition

Speaker and Noise Factorization for Robust Speech Recognition
复制标题

DOI:
10.1109/tasl.2012.2198059
复制
发表时间:
2012-09
期刊:
IEEE Transactions on Audio, Speech, and Language Processing
影响因子:
--
通讯作者:
Yongqiang Wang;M. Gales
Yongqiang Wang;M. Gales
中科院分区:
其他
文献类型:
--
作者:
Yongqiang Wang;M. Gales

文献摘要

被引文献

相似文献

语音识别系统需要在广泛的条件下运行。因此,它们应该对由各种声学因素(例如扬声器差异、传输信道和背景噪声)引起的外部变化具有鲁棒性。对于许多场景,多个因素同时影响潜在的"干净"语音信号。本文探讨了处理扬声器和背景噪声差异的技术。采用声学因子分解方法。这里,分配单独的变换来表示扬声器[最大似然线性回归(MLLR)]以及噪声和信道[基于模型的向量泰勒级数(VTS)]因素。与对扬声器和噪声因素的组合影响进行建模的标准方法相比,这是一个高度灵活的框架。例如,因子分解允许将在一种噪声条件下获得的扬声器特征应用于不同的环境。为了获得这种因式分解的MLLR和VTS的训练和应用程序的修改版本。所提出的方案进行评估的适应和因式分解的AURORA4数据。
Speech recognition systems need to operate in a wide range of conditions. Thus they should be robust to extrinsic variability caused by various acoustic factors, for example speaker differences, transmission channel and background noise. For many scenarios, multiple factors simultaneously impact the underlying “clean” speech signal. This paper examines techniques to handle both speaker and background noise differences. An acoustic factorization approach is adopted. Here, separate transforms are assigned to represent the speaker [maximum-likelihood linear regression (MLLR)], and noise and channel [model-based vector Taylor series (VTS)] factors. This is a highly flexible framework compared to the standard approaches of modeling the combined impact of both speaker and noise factors. For example factorization allows the speaker characteristics obtained in one noise condition to be applied to a different environment. To obtain this factorization modified versions of MLLR and VTS training and application are derived. The proposed scheme is evaluated for both adaptation and factorization on the AURORA4 data.