Proposal of a New Confidence Parameter Estimating the Number of Speakers -An experimental investigation-

Proposal of a New Confidence Parameter Estimating the Number of Speakers -An experimental investigation-
复制标题

提出一种新的估计发言者数量的置信参数-实验研究-

DOI:
--
复制
发表时间:
2010
期刊:
J. Inf. Hiding Multim. Signal Process.
影响因子:
--
通讯作者:
S. Ouamour
S. Ouamour
中科院分区:
--
文献类型:
--
作者:
H. Sayoud;S. Ouamour

文献摘要

被引文献

相似文献

。在语音重叠的情况下是否可以知道有多少说话者同时说话?如果人类大脑(尚未掌握创造能力)能够做到这一点,甚至能够理解混合语音的含义,那么现有的自动说话人识别系统还无法做到这一点。实际上,这些系统在这种情况下会表现出严重的退化。对于此任务,我们提出了一种能够估计混合语音信号中说话者数量的新方法。这里开发的算法基于对语音信号频谱分析提取的第七梅尔系数的统计特征的计算。该算法使用置信度参数(我们称之为 PENS),在七个不同的 ORATOR 数据库集上进行了测试,其中每个集包含七个多说话者文件。结果表明,PENS 参数允许我们在单扬声器信号(只有一个扬声器正在讲话)和混合扬声器信号(多个扬声器同时讲话)之间做出良好的区分,没有任何歧义。此外,它允许我们在混合语音信号的情况下以良好的精度估计说话者的数量,特别是当说话者的数量少于四个时。
. Is it possible to know how many speakers are speaking simultaneously in case of speech overlap? If the human brain, creation not yet mastered, manages to do it and even to understand the mixed speech meaning, it is not yet the case for the existing systems of automatic speaker recognition. In practice, these systems present a strong degradation in such situations. For this task, we propose a new method able to estimate the number of speakers in a mixture of speech signals. The algorithm developed here is based on the computation of the statistical characteristic of the 7th Mel coefficient extracted by spectral analysis from the speech signal. This algorithm using a confidence parameter, which we called PENS, is tested on seven different sets of the ORATOR database, where each set contains seven multi-speaker files. Results show that the PENS parameter permits us to make a good discrimination, without any ambiguity, between a mono-speaker signal (only one speaker is speaking) and a mixed-speakers signal (several speakers are speaking simultaneously). Moreover, it permits us to estimate, in case of mixed speech signals, the number of speakers with a good precision, especially when the number of speakers is less than four.