Kernel Machines Beat Deep Neural Networks on Mask-based Single-channel Speech Enhancement

Kernel Machines Beat Deep Neural Networks on Mask-based Single-channel Speech Enhancement
复制标题

DOI:
10.21437/interspeech.2019-1344
复制
发表时间:
2018-11
期刊:
ArXiv
影响因子:
--
通讯作者:
Like Hui;Siyuan Ma;M. Belkin
Like Hui;Siyuan Ma;M. Belkin
中科院分区:
其他
文献类型:
--
作者:
Like Hui;Siyuan Ma;M. Belkin

文献摘要

相似文献

本文提出了一种快速核方法用于基于掩码的单通道语音增强。具体来说,我们的方法解决了一个核回归问题与非光滑核函数(指数幂核)与高效的迭代方法(EigenPro)。由于该方法的简单性,其超参数,如核带宽可以自动和有效地选择使用线搜索与训练数据的子样本。我们观察到的回归损失(均方误差)和语音增强的常规指标之间的经验相关性。这一观察结果证明了我们的训练目标是正确的,并激励我们通过训练每个频率子带的独立核模型来实现更低的回归损失。我们将我们的方法与基于掩码的HINT和TIMIT的最先进的深度神经网络进行了比较。实验结果表明,我们的内核方法始终优于深度神经网络,同时需要更少的训练时间。
We apply a fast kernel method for mask-based single-channel speech enhancement. Specifically, our method solves a kernel regression problem associated to a non-smooth kernel function (exponential power kernel) with a highly efficient iterative method (EigenPro). Due to the simplicity of this method, its hyper-parameters such as kernel bandwidth can be automatically and efficiently selected using line search with subsamples of training data. We observe an empirical correlation between the regression loss (mean square error) and regular metrics for speech enhancement. This observation justifies our training target and motivates us to achieve lower regression loss by training separate kernel model per frequency subband. We compare our method with the state-of-the-art deep neural networks on mask-based HINT and TIMIT. Experimental results show that our kernel method consistently outperforms deep neural networks while requiring less training time.