Musical Instrument Recognition in Polyphonic Audio Using Missing Feature Approach

Musical Instrument Recognition in Polyphonic Audio Using Missing Feature Approach
复制标题

DOI:
10.1109/tasl.2013.2248720
复制
发表时间:
2013-09-01
影响因子:
--
通讯作者:
Klapuri, Anssi
Klapuri, Anssi
中科院分区:
其他
文献类型:
--
作者:
Giannoulis, Dimitrios;Klapuri, Anssi

文献摘要

被引文献

相似文献

描述了一种用于在多个声源同时活动的复调音频信号中进行乐器识别的方法。该方法是基于局部谱特征和特征缺失技术。描述了一种新的掩模估计算法,其识别包含每个声源的可靠信息的频谱区域,然后使用有界边缘化来处理被确定为不可靠的特征向量元素。掩模估计技术是基于这样的假设,即音乐声音的频谱包络往往是缓慢变化的对数频率的函数,因此,不可靠的频谱分量可以被检测为从估计的平滑频谱包络的正偏差。提出了一种计算效率高的算法,用于在分类过程中边缘化掩模。在仿真中,该方法明显优于混合信号的参考方法。所提出的掩模估计技术导致的识别精度,大约是一个平凡的所有一个掩模(所有功能都假定可靠)和一个理想的“甲骨文”掩模之间的一半。
A method is described for musical instrument recognition in polyphonic audio signals where several sound sources are active at the same time. The proposed method is based on local spectral features and missing-feature techniques. A novel mask estimation algorithmis described that identifies spectral regions that contain reliable information for each sound source, and bounded marginalization is then used to treat the feature vector elements that are determined to be unreliable. The mask estimation technique is based on the assumption that the spectral envelopes of musical sounds tend to be slowly-varying as a function of log-frequency and unreliable spectral components can therefore be detected as positive deviations from an estimated smooth spectral envelope. A computationally efficient algorithm is proposed for marginalizing the mask in the classification process. In simulations, the proposed method clearly outperforms reference methods for mixture signals. The proposed mask estimation technique leads to a recognition accuracy that is approximately half-way between a trivial all-one mask (all features are assumed reliable) and an ideal "oracle" mask.