A pitch-based rapid speech segmentation for speaker indexing

A pitch-based rapid speech segmentation for speaker indexing
复制标题

用于说话人索引的基于音调的快速语音分割

DOI:
--
复制
发表时间:
2005
期刊:
IEEE International Symposium on Multimedia
影响因子:
--
通讯作者:
Zhaohui Wu
Zhaohui Wu
中科院分区:
--
文献类型:
--
作者:
Min Yang;Yingchun Yang;Zhaohui Wu

文献摘要

被引文献

相似文献

在许多应用中,连续音频的分割是一个重要的处理。在说话人索引中,说话人模型的可靠性很大程度上取决于分割。常用的方法是基于贝叶斯信息准则(BIC),然而,这是不那么有能力时,处理短的话语。本文提出了一种基于基音周期的语音分割方法,该方法能够准确、快速地检测出说话人频繁切换。在我们的算法中,基音被引入到说话人分割中。首先,通过基音检测出说话段。然后计算基音距离,并与自适应阈值进行比较。说话人的变化最终决定在话语段。我们应用我们的方法和三种比较方法对HUB 4-NE广播数据。对每种算法进行了说话人索引实验。在分割效果评价中,提出了虚警和漏警两个指标作为补充。实验结果表明,该算法能更快更好地检测出说话人的短时变化。该方法的话者标引等误率为10.43%,远低于其他方法的12.94%、25.84%和15.91%。
Segmentation of continuous audio is an important processing in many applications. In speaker indexing, the reliability of speaker model depends much on segmentation. Commonly used methods are based on the Bayesian information criteria (BIC), which is however not so capable when dealing with short utterances. In this paper, we present a pitch-based speech segmentation method, which can detect frequent speaker changes accurately and rapidly. In our algorithm, pitch is introduced in speaker segmentation. Firstly, utterance segments are detected by pitch. Then distances of pitch are computed, and compared with a self-adaptable threshold. Speaker changes are finally decided among utterance segments. We applied our method and three comparative methods on the HUB4-NE broadcast data. Speaker indexing experiments have been taken following each algorithm. We also suggested two indicators as complements of false alarm and missing rate in the evaluation of segmentation. The experiment results show that our algorithm works faster and better, with most of short time speaker changes detected. Speaker indexing equal error rate of our method is 10.43%, which is much lower than 12.94%, 25.84% and 15.91% of other methods.