MAXIMUM-LIKELIHOOD LINEAR-REGRESSION FOR SPEAKER ADAPTATION OF CONTINUOUS DENSITY HIDDEN MARKOV-MODELS

MAXIMUM-LIKELIHOOD LINEAR-REGRESSION FOR SPEAKER ADAPTATION OF CONTINUOUS DENSITY HIDDEN MARKOV-MODELS
复制标题

DOI:
10.1006/csla.1995.0010
复制
发表时间:
1995-04-01
影响因子:
4.3
通讯作者:
WOODLAND, PC
WOODLAND, PC
中科院分区:
计算机科学3区
文献类型:
--
作者:
LEGGETTER, CJ;WOODLAND, PC

文献摘要

被引文献

相似文献

提出了一种基于连续密度隐马尔可夫模型(HHMM)的说话人自适应方法。一个初始的说话人无关系统是适合于通过更新HMM参数来改善一个新的说话人的建模。从可用的适应数据中收集统计数据,并用于计算平均向量的基于线性回归的变换。计算变换矩阵以最大化自适应数据的可能性,并且可以使用前向-后向算法来实现。通过在多个分布之间绑定变换,可以对训练数据中未表示的分布执行自适应。该方法的一个重要特点是可以使用任意的自适应数据-不需要特殊的enrollments.Experiments已经在ARPA RM 1数据库上进行了使用HMM系统与跨词三音子和混合高斯输出分布。结果表明,自适应可以使用少至11 s的自适应数据来执行,并且随着使用更多的数据,自适应性能提高。例如,使用40个自适应话语,在有监督的自适应下,与说话者无关的系统的错误减少了37%,在无监督模式下减少了32%。
A method of speaker adaptation for continuous density hidden Markov models (HMMs) is presented. An initial speaker-independent system is adapted to improve the modelling of a new speaker by updating the HMM parameters. Statistics are gathered from the available adaptation data and used to calculate a linear regression-based transformation for the mean vectors. The transformation matrices ale calculated to maximize the likelihood of the adaptation data and can be implemented using the forward-backward algorithm. By tying the transformations among a number of distributions, adaptation can be performed for distributions which are not represented in the training data. An important feature of the method is that arbitrary adaptation data can be used-no special enrolment sentences are needed.Experiments have been performed on the ARPA RM1 database using an HMM system with cross-word triphones and mixture Gaussian output distributions. Results show that adaptation can be performed using as little as 11 s of adaptation data, and that as more data is used the adaptation performance improves. For example, using 40 adaptation utterances, a 37% reduction in error from the speaker-independent system was achieved with supervised adaptation and a 32% reduction in unsupervised mode.