Auditory scene analysis based on time-frequency integration of shared FM and AM (II): Optimum time-domain integration and stream sound reconstruction

Auditory scene analysis based on time-frequency integration of shared FM and AM (II): Optimum time-domain integration and stream sound reconstruction
复制标题

基于共享FM和AM时频积分的听觉场景分析(二):最优时域积分和流声重构

DOI:
10.1002/scj.1160
复制
发表时间:
2002
期刊:
Systems and Computers in Japan
影响因子:
--
通讯作者:
S. Ando
S. Ando
中科院分区:
--
文献类型:
--
作者:
M. Abe;S. Ando

文献摘要

被引文献

相似文献

在前面的论文中,我们提出了一种听觉场景分析方法,通过投票的方法将时频空间中的瞬时频率、频率变化率和幅度变化率强化为多峰概率密度分布,并实现对混合声音流的分组。在本文中,作为该方法后半部分的要点,我们将引入流参数根据已知动态缓慢变化的假设,并提出一种时间轴上的积分方法,其中通过非参数卡尔曼滤波器在时间序列上最优估计流参数的概率密度分布。通过这样做,可以实现诸如增强流参数的准确性、流的中断的插值和连接以及将先验知识引入流选择等更高听觉场景分析的机制。此外,构建了与流相对应的声音分离和重构系统,并通过合成声音或音乐声音和语音的基础实验验证了所提出的技术。 © 2002 Wiley periodicals, Inc. Syst Comp Jpn, 33(10): 83–94, 2002;在线发表于 Wiley InterScience (www.interscience.wiley.com)。 DOI 10.1002/scj.1160
In the preceding paper, we have proposed a method for auditory scene analysis, in which the instantaneous frequency, frequency change rate, and amplitude change rate in time-frequency space are intensified into a multipeak probability density distribution by voting method and the grouping into streams of mixed sounds is realized. In this paper, as the main point of the second half of this method, we will introduce the assumption that the stream parameters vary slowly according to the known dynamics and propose an integration method on the time axis, in which the probability density distribution of the stream parameters is optimally estimated in time series by a nonparametric Kalman filter. By doing so, the mechanism of higher auditory scene analysis such as enhancement of the accuracy of the stream parameters, interpolation and connection of the breaks of the streams, and introduction of a priori knowledge into stream selection can be realized. Moreover, the separation and reconstruction system of sounds which correspond to streams is constructed, and the proposed technique is verified by fundamental experiments for synthesized sounds or musical sounds and voices. © 2002 Wiley Periodicals, Inc. Syst Comp Jpn, 33(10): 83–94, 2002; Published online in Wiley InterScience (www.interscience.wiley.com). DOI 10.1002/scj.1160