A comparative study of spectral clustering for i-vector-based speaker clustering under noisy conditions

A comparative study of spectral clustering for i-vector-based speaker clustering under noisy conditions
复制标题

DOI:
10.1109/icassp.2015.7178329
复制
发表时间:
2015-04
期刊:
2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
影响因子:
--
通讯作者:
Naohiro Tawara;Tetsuji Ogawa;Tetsunori Kobayashi
Naohiro Tawara;Tetsuji Ogawa;Tetsunori Kobayashi
中科院分区:
其他
文献类型:
--
作者:
Naohiro Tawara;Tetsuji Ogawa;Tetsunori Kobayashi

文献摘要

相似文献

本文研究了受噪声干扰语音的说话人聚类问题。一般来说,说话人聚类的性能很大程度上取决于语音话语之间的相似性可以测量。最近提出的基于i向量的余弦相似度在说话人聚类系统中具有最先进的性能。然而,这种相似性往往无法捕捉到说话人的相似性在嘈杂的条件下。因此,我们试图检查基于i-向量的相似性的频谱聚类的效率,因为频谱聚类可以通过非线性投影产生对噪声的鲁棒性。实验比较表明,谱聚类在非平稳噪声条件下,比传统的方法,如凝聚聚类和k-means聚类,有显着的改善。
The present paper dealt with speaker clustering for speech corrupted by noise. In general, the performance of speaker clustering significantly depends on how well the similarities between speech utterances can be measured. The recently proposed i-vector-based cosine similarity has yielded the state-of-the-art performance in speaker clustering systems. However, this similarity often fails to capture the speaker similarity under noisy conditions. Therefore, we attempted to examine the efficiency of spectral clustering on i-vector-based similarity for speech corrupted by noise because spectral clustering can yield robustness against noise by non-linear projection. Experimental comparisons demonstrated that spectral clustering yielded significant improvement from conventional methods, such as agglomerative clustering and k-means clustering, under non-stationary noise conditions.