Monaural voiced speech segregation based on elaborate harmonic grouping strategies

Monaural voiced speech segregation based on elaborate harmonic grouping strategies
复制标题

DOI:
10.1007/s11432-011-4506-2
复制
发表时间:
2011-12
期刊:
Science China Information Sciences
影响因子:
--
通讯作者:
Wenju Liu;Xueliang Zhang;Wei Jiang;Peng Li;Bo Xu
Wenju Liu;Xueliang Zhang;Wei Jiang;Peng Li;Bo Xu
中科院分区:
其他
文献类型:
--
作者:
Wenju Liu;Xueliang Zhang;Wei Jiang;Peng Li;Bo Xu

文献摘要

被引文献

相似文献

本文提出了一种基于谐波分组策略的单音语音分离增强算法。本文算法的主要成果体现在三个方面。首先,该算法根据载波包络能量比将时频(T-F)单元划分为已分辨和未分辨单元,分类结果比跨信道相关更准确;其次,根据人类感知中已被验证的最小振幅原理和谐波原理,对已分辨的T-F单元进行分组。最后,利用“增强型”包络自相关函数检测调幅率,大大降低了未解析单元分组时的半频误差。系统评价和比较表明,该算法极大地提高了分离性能。其中,信噪比较前一种方法提高了0.96 dB。此外,我们的算法在提高PESQ分数和主观感知分数方面也很有效。
In this paper, an enhanced algorithm based on several elaborate harmonic grouping strategies for monaural voiced speech segregation is proposed. Main achievements of the proposed algorithm lie in three aspects. Firstly, the algorithm classifies the time-frequency (T-F) units into resolved and unresolved ones by carrier-to-envelope energy ratio, which leads to more accurate classification results than by cross-channel correlation. Secondly, resolved T-F units are grouped together according to minimum amplitude principle, which has been verified to exist in human perception, as well as the harmonic principle. Finally, “enhanced” envelope autocorrelation function is employed to detect amplitude modulation rates, which helps a lot in reducing half-frequency error in grouping of unresolved units. Systematic evaluation and comparison show that performance of separation is greatly improved by the proposed algorithm. Specifically, signal-to-noise ratio (SNR) is improved by 0.96 dB compared with that of previous method. Besides, our algorithm is also effective in improving the PESQ score and subjective perception score.