Sound source segregation based on estimating incident angle of each frequency component of input signals acquired by multiple microphones

Sound source segregation based on estimating incident angle of each frequency component of input signals acquired by multiple microphones
复制标题

基于估计多个麦克风获取的输入信号的每个频率分量的入射角的声源分离

DOI:
10.1250/ast.22.149
复制
发表时间:
2001
影响因子:
0.7
通讯作者:
Y. Kaneda
Y. Kaneda
中科院分区:
--
文献类型:
--
作者:
M. Aoki;M. Okamoto;S. Aoki;Hiroyuki Matsui;T. Sakurai;Y. Kaneda

文献摘要

被引文献

相似文献

我们已经开发出一种方法,分离所需的语音从并发的声音接收两个麦克风。在这种方法中,我们称之为SAFIA,由两个麦克风接收的信号通过离散傅立叶变换进行分析。对于每个频率分量,计算通道之间的幅度和相位差。这些差用于选择来自期望方向的信号的频率分量,并将这些分量重构为期望源信号。为了阐明频率分辨率对所提出的方法的影响,我们进行了三个实验。首先分析了频率分辨率与功率谱累积分布的关系。我们发现,语音信号的功率集中在特定的频率成分时,频率分辨率约为10赫兹。其次,我们确定是否一个给定的频率分辨率减少了两个语音信号的频率分量之间的重叠。10 Hz的频率分辨率使重叠最小化。第三,通过主观测试,分析了音质与频率分辨率的关系。就音质而言,最佳频率分辨率对应于将语音信号功率集中在特定频率分量上并使重叠程度最小化的频率分辨率。最后,我们证明了这种方法提高了信噪比超过18dB。
We have developed a method of segregating desired speech from concurrent sounds received by two microphones. In this method, which we call SAFIA, signals received by two microphones are analyzed by discrete Fourier transformation. For each frequency component, differences in the amplitude and phase between channels are calculated. These differences are used to select frequency components of the signal that come from the desired direction and to reconstruct these components as the desired source signal. To clarify the effect of frequency resolution on the proposed method, we conducted three experiments. First, we analyzed the relationship between frequency resolition and the power spectrum’s cumulative distribution. We found that the speech-signal power was concentrated on specific frequency components when the frequency resolution was about 10-Hz. Second, we determined whether a given frequency resolution decreased the overlap between the frequency components of two speech signals. A 10-Hz frequency resolution minimized the overlap. Third, we analyzed the relationship between sound quality and frequency resolution through subjective tests. The best frequency resolution in terms of sound quality corresponded to the frequency resolutions that concentrated the speech signal power on specific frequency components and that minimized the degree of overlap. Finally, we demonstrated that this method improved the signal-to-noise ratio by over 18dB.