Boosted Locality Sensitive Hashing: Discriminative Binary Codes for Source Separation

Boosted Locality Sensitive Hashing: Discriminative Binary Codes for Source Separation
复制标题

DOI:
10.1109/icassp40776.2020.9053052
复制
发表时间:
2020-02
期刊:
ICASSP 2020 - 2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
影响因子:
--
通讯作者:
Sunwoo Kim;Haici Yang;Minje Kim
Sunwoo Kim;Haici Yang;Minje Kim
中科院分区:
其他
文献类型:
--
作者:
Sunwoo Kim;Haici Yang;Minje Kim

文献摘要

相似文献

随着深度学习技术的进步,语音增强任务有了显著的改进,但代价是计算复杂性增加。在这项研究中,我们提出了一种自适应Boosting方法来学习对位置敏感的哈希码,这些哈希码有效地表示了音频频谱。我们将学习的哈希码用于单通道语音去噪任务,作为复杂机器学习模型的替代方案,特别是在资源受限的环境中。我们的自适应Boosting算法将简单的Logistic回归学习为弱学习器。一旦训练完成,它们的二进制分类结果将测试噪声语音的每个频谱转换为比特串。简单的逐位运算计算汉明距离,以在训练噪声语音频谱的词典中找到数学上最接近的匹配帧,其关联的理想二进制掩码被平均以估计该测试混合的去噪掩码。我们提出的学习算法与AdaBoost的不同之处在于,投影的训练是为了最小化散列码的自相似矩阵与原始频谱的自相似矩阵之间的距离,而不是错误分类率。我们在不同噪声类型的Timit语料库上测试了我们的鉴别哈希码,并在去噪性能和复杂度方面与深度学习方法进行了比较。
Speech enhancement tasks have seen significant improvements with the advance of deep learning technology, but with the cost of increased computational complexity. In this study, we propose an adaptive boosting approach to learning locality sensitive hash codes, which represent audio spectra efficiently. We use the learned hash codes for single-channel speech denoising tasks as an alternative to a complex machine learning model, particularly to address the resource-constrained environments. Our adaptive boosting algorithm learns simple logistic regressors as the weak learners. Once trained, their binary classification results transform each spectrum of test noisy speech into a bit string. Simple bitwise operations calculate Hamming distance to find the $\mathcal{K}$-nearest matching frames in the dictionary of training noisy speech spectra, whose associated ideal binary masks are averaged to estimate the denoising mask for that test mixture. Our proposed learning algorithm differs from AdaBoost in the sense that the projections are trained to minimize the distances between the self-similarity matrix of the hash codes and that of the original spectra, rather than the misclassification rate. We evaluate our discriminative hash codes on the TIMIT corpus with various noise types, and show comparative performance to deep learning methods in terms of denoising performance and complexity.