Variable frame rate-based data augmentation to handle speaking-style variability for automatic speaker verification

Variable frame rate-based data augmentation to handle speaking-style variability for automatic speaker verification
复制标题

DOI:
10.21437/interspeech.2020-3006
复制
发表时间:
2020-08
期刊:
ArXiv
影响因子:
--
通讯作者:
Amber Afshan;Jinxi Guo;S. Park;Vijay Ravi;A. McCree;A. Alwan
Amber Afshan;Jinxi Guo;S. Park;Vijay Ravi;A. McCree;A. Alwan
中科院分区:
其他
文献类型:
--
作者:
Amber Afshan;Jinxi Guo;S. Park;Vijay Ravi;A. McCree;A. Alwan

文献摘要

被引文献

相似文献

使用UCLA Speaker Variability数据库研究了说话风格变异对说话人自动确认的影响,该数据库包含每个说话人的多种说话风格。一个x-向量/PLDA(概率线性判别分析)系统进行了训练与SRE和交换机数据库与标准的增强技术,并与UCLA数据库的话语进行评估。当登记和测试话语具有相同风格(例如,0.98%和0.57%的阅读和会话的讲话,分别),但大幅增加时,注册和测试话语之间的风格不匹配。例如,当注册与会话话语,EER增加到3.03%,2.96%和22.12%,分别测试时,阅读,叙事和宠物导向的讲话。为了减少风格不匹配的影响,我们提出了一种基于熵的可变帧速率技术来人为地生成用于PLDA适应的风格归一化表示。所提出的系统显著提高了性能。在上述条件下,EER提高到2.69%(对话-阅读),2.27%(对话-叙述)和18.75%(宠物指导-阅读)。总的来说,所提出的技术执行多风格PLDA适应,而不需要每个扬声器在不同的说话风格的训练数据。
The effects of speaking-style variability on automatic speaker verification were investigated using the UCLA Speaker Variability database which comprises multiple speaking styles per speaker. An x-vector/PLDA (probabilistic linear discriminant analysis) system was trained with the SRE and Switchboard databases with standard augmentation techniques and evaluated with utterances from the UCLA database. The equal error rate (EER) was low when enrollment and test utterances were of the same style (e.g., 0.98% and 0.57% for read and conversational speech, respectively), but it increased substantially when styles were mismatched between enrollment and test utterances. For instance, when enrolled with conversation utterances, the EER increased to 3.03%, 2.96% and 22.12% when tested on read, narrative, and pet-directed speech, respectively. To reduce the effect of style mismatch, we propose an entropy-based variable frame rate technique to artificially generate style-normalized representations for PLDA adaptation. The proposed system significantly improved performance. In the aforementioned conditions, the EERs improved to 2.69% (conversation -- read), 2.27% (conversation -- narrative), and 18.75% (pet-directed -- read). Overall, the proposed technique performed comparably to multi-style PLDA adaptation without the need for training data in different speaking styles per speaker.