Monaural speech segregation based on pitch tracking and amplitude modulation

Monaural speech segregation based on pitch tracking and amplitude modulation
复制标题

DOI:
10.1109/tnn.2004.832812
复制
发表时间:
2004-09-01
影响因子:
--
通讯作者:
Wang, DL
Wang, DL
中科院分区:
其他
文献类型:
--
作者:
Hu, GN;Wang, DL

文献摘要

被引文献

相似文献

从一个单声道录音分离语音已被证明是非常具有挑战性的。有声语音的单声道分离已经在结合听觉场景分析原理的先前系统中进行了研究。这些系统的一个主要问题是它们不能处理语音的高频部分。心理声学证据表明,不同的感知机制涉及处理解决和未解决的谐波。我们提出了一个新的系统,分离解决和未解决的谐波不同的有声语音分离。对于解析谐波,系统基于时间连续性和跨通道相关性生成分段,并根据其周期性对其进行分组。对于未分辨的谐波,它除了时间连续性之外还基于公共幅度调制(AM)生成段,并根据AM速率对它们进行分组。分离过程的基础是一个音高轮廓,它首先从根据主导音高分离的语音中估计出来,然后根据心理声学约束进行调整。我们的系统进行了系统的评估,并与以往的系统相比,它产生了更好的性能,特别是对语音的高频部分。
Segregating speech from one monaural recording has proven to be very challenging. Monaural segregation of voiced speech has been studied in previous systems that incorporate auditory scene analysis principles. A major problem for these systems is their inability to deal with the high-frequency part of speech. Psychoacoustic evidence suggests that different perceptual mechanisms are involved in handling resolved and unresolved harmonics. We propose a novel system for voiced speech segregation that segregates resolved and unresolved harmonics differently. For resolved harmonics, the system generates segments based on temporal continuity and cross-channel correlation, and groups them according to their periodicities. For unresolved harmonics, it generates segments based on common amplitude modulation (AM) in addition to temporal continuity and groups them according to AM rates. Underlying the segregation process is a pitch contour that is first estimated from speech segregated according to dominant pitch and then adjusted according to psychoacoustic constraints. Our system is systematically evaluated and compared with pervious systems, and it yields substantially better performance, especially for the high-frequency part of speech.