Speaker background models for connected digit password speaker verification

Speaker background models for connected digit password speaker verification
复制标题

DOI:
10.1109/icassp.1996.540295
复制
发表时间:
1996-05
期刊:
1996 IEEE International Conference on Acoustics, Speech, and Signal Processing Conference Proceedings
影响因子:
--
通讯作者:
A. Rosenberg;S. Parthasarathy
A. Rosenberg;S. Parthasarathy
中科院分区:
其他
文献类型:
--
作者:
A. Rosenberg;S. Parthasarathy

文献摘要

被引文献

相似文献

似然比或群组标准化评分已被证明是有效的,以提高说话人确认系统的性能。在这方面的一个重要问题是建立原则,用于构建扬声器背景或队列模型,提供最有效的归一化分数。研究了几种说话人背景模型。这些模型包括个体说话者模型、从不同数量说话者的汇集话语构建的模型、基于与客户模型的相似性选择的模型、从说话者的随机选择构建的模型、以及从在与客户模型不同的条件下记录的数据库构建的模型。实验结果表明,基于与参考说话人相似性的混合模型比来自相同相似说话人集合的单个群组模型表现得更好。基于相似性的来自少量说话者的池化背景模型表现最好,但并不显著优于随机选择40个或更多性别平衡的说话者,其训练条件与参考说话者相匹配。
Likelihood ratio or cohort normalized scoring has been shown to be effective for improving the performance of speaker verification systems. An important problem in this connection is the establishment of principles for constructing speaker background or cohort models which provide the most effective normalized scores. Several kinds of speaker background models are studied. These include individual speaker models, models constructed from the pooled utterances of different numbers of speakers, models selected on the basis of similarity with customer models, models constructed from random selections of speakers, and models constructed from databases recorded under different conditions than the customer models. The results of experiments show that pooled models based on similarity to the reference speaker perform better than individual cohort models from the same similar set of speakers. Pooled background models from a small number of speakers based on similarity perform about the best, but not significantly better than a random selection of 40 or more gender balanced speakers with training conditions matched to the reference speakers.