Monaural speech segregation

Monaural speech segregation
复制标题

单耳语音分离

DOI:
10.1121/1.4780325
复制
发表时间:
2003
期刊:
影响因子:
--
通讯作者:
Guoning Hu
Guoning Hu
中科院分区:
--
文献类型:
--
作者:
Deliang Wang;Guoning Hu

文献摘要

参考文献

被引文献

相似文献

从单声道录音中分离语音是听觉场景分析的主要任务,并且已被证明是非常具有挑战性的。我们提出了一个多阶段的任务模型。该模型从模拟听觉周边开始。随后的阶段计算中级听觉表征,包括相关图和跨通道相关性。该系统的核心在二维时频表示中执行分割和分组,该二维时频表示对频率和时间、周期性和幅度调制(AM)中的接近度进行编码。出于心理声学观察,我们的系统采用不同的机制来处理解决和未解决的谐波。对于解析的谐波,我们的系统基于时间连续性和跨通道相关性生成片段-听觉场景的基本组成部分,并根据周期性对它们进行分组。对于未分辨谐波,系统除了时间连续性之外还基于AM生成分段,并根据AM重复率对它们进行分组。我们使用正弦建模和梯度下降来获得AM重复率。分离过程的基础是音调轮廓,该轮廓首先根据全局音调从分离的语音中估计,然后根据心理声学约束进行调整。该模型已被系统地评估,它产生了比以前的系统更好的性能。
Speech segregation from a monaural recording is a primary task of auditory scene analysis, and has proven to be very challenging. We present a multistage model for the task. The model starts with simulated auditory periphery. A subsequent stage computes midlevel auditory representations, including correlograms and cross‐channel correlations. The core of the system performs segmentation and grouping in a two‐dimensional time‐frequency representation that encodes proximity in frequency and time, periodicity, and amplitude modulation (AM). Motivated by psychoacoustic observations, our system employs different mechanisms for handling resolved and unresolved harmonics. For resolved harmonics, our system generates segments—basic components of an auditory scene—based on temporal continuity and cross‐channel correlation, and groups them according to periodicity. For unresolved harmonics, the system generates segments based on AM in addition to temporal continuity and groups them according to AM repetition rates. We derive AM repetition rates using sinusoidal modeling and gradient descent. Underlying the segregation process is a pitch contour that is first estimated from speech segregated according to global pitch and then adjusted according to psychoacoustic constraints. The model has been systematically evaluated, and it yields substantially better performance than previous systems.
使用事件同步听觉图像和 STRAIGHT 进行语音分离
DOI: --
发表时间: 2005
期刊: Speech separation by human and machines
影响因子: --
作者:
Toshio Irino
通讯作者: Toshio Irino