Perceptual-Similarity-Aware Deep Speaker Representation Learning for Multi-Speaker Generative Modeling

Perceptual-Similarity-Aware Deep Speaker Representation Learning for Multi-Speaker Generative Modeling
复制标题

DOI:
10.1109/taslp.2021.3059114
复制
发表时间:
2021
期刊:
IEEE/ACM Transactions on Audio, Speech, and Language Processing
影响因子:
--
通讯作者:
Yuki Saito;Shinnosuke Takamichi;H. Saruwatari
Yuki Saito;Shinnosuke Takamichi;H. Saruwatari
中科院分区:
其他
文献类型:
--
作者:
Yuki Saito;Shinnosuke Takamichi;H. Saruwatari

文献摘要

被引文献

相似文献

我们提出了一种新的深度说话人表示学习,该学习考虑了说话人之间的感知相似性,用于多说话人生成建模。在成功地对说话人个性进行准确的判别建模之后,深度说话人表示学习(即使用深度神经网络的说话人表示学习)的知识被引入到多说话人生成建模中。然而,传统的判别算法不一定能学习到适合这种生成建模的说话人嵌入,这可能导致合成语音的质量较低,可控性较差。我们提出了三种表征学习算法,它们利用通过对说话人对相似度的大规模感知评分获得的感知说话人相似矩阵。这些算法训练一个说话人编码器,用三种不同的矩阵表示来学习说话人嵌入:一组向量、Gram矩阵和一个图。此外,我们提出了一种主动学习算法,迭代感知评分和说话编码器训练。为了在减少评分和训练成本的同时获得准确的嵌入,该算法根据顺序训练的说话人编码器的相似性预测结果,选择未评分的说话人对进行下一步评分。实验评价结果表明:1)所提出的表征学习算法学习到与感知说话人对相似度强相关的说话人嵌入;2)所提出的嵌入比通过判别建模学习到的传统d向量更好地改善了语音自动编码任务中的合成语音质量;3)所提出的主动学习算法在降低评分和训练成本的同时获得了更高的合成语音质量。4)在提出的相似度{向量、矩阵、图}嵌入算法中,第一种算法对合成语音的说话人相似度达到最佳,第三种算法对合成语音的自然度改善最大。
We propose novel deep speaker representation learning that considers perceptual similarity among speakers for multi-speaker generative modeling. Following its success in accurate discriminative modeling of speaker individuality, knowledge of deep speaker representation learning (i.e., speaker representation learning using deep neural networks) has been introduced to multi-speaker generative modeling. However, the conventional discriminative algorithm does not necessarily learn speaker embeddings suitable for such generative modeling, which may result in lower quality and less controllability of synthetic speech. We propose three representation learning algorithms that utilize a perceptual speaker similarity matrix obtained by large-scale perceptual scoring of speaker-pair similarity. The algorithms train a speaker encoder to learn speaker embeddings with three different representations of the matrix: a set of vectors, the Gram matrix, and a graph. Furthermore, we propose an active learning algorithm that iterates the perceptual scoring and speaker encoder training. To obtain accurate embeddings while reducing costs of scoring and training, the algorithm selects unscored speaker-pairs to be scored next on the basis of the sequentially-trained speaker encoder's similarity prediction results. Experimental evaluation results show that 1) the proposed representation learning algorithms learn speaker embeddings strongly correlated with perceptual speaker-pair similarity, 2) the embeddings improve synthetic speech quality in speech autoencoding tasks better than conventional d-vectors learned by discriminative modeling, 3) the proposed active learning algorithm achieves higher synthetic speech quality while reducing costs of scoring and training, and 4) among the proposed similarity {vector, matrix, graph} embedding algorithms, the first achieves the best speaker similarity for synthetic speech and the third gives the most improvement in the synthetic speech naturalness.