Bloom-Net: Blockwise Optimization for Masking Networks Toward Scalable and Efficient Speech Enhancement

Bloom-Net: Blockwise Optimization for Masking Networks Toward Scalable and Efficient Speech Enhancement
复制标题

DOI:
10.1109/icassp43922.2022.9746767
复制
发表时间:
2021-11
期刊:
ICASSP 2022 - 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
影响因子:
--
通讯作者:
Sunwoo Kim;Minje Kim
Sunwoo Kim;Minje Kim
中科院分区:
其他
文献类型:
--
作者:
Sunwoo Kim;Minje Kim

文献摘要

相似文献

在本文中,我们提出了一种基于掩蔽的网络(BLOOM-Net)的分块优化方法,用于训练可扩展的语音增强网络。在这里,我们使用残差学习方案设计网络,并顺序训练内部分隔符块,以获得用于语音增强的可扩展的基于掩蔽的深度神经网络。其可扩展性使其能够根据测试时环境动态调整运行时复杂性。为此,我们对模型进行模块化,以便它们可以灵活地适应增强性能的不同需求和资源限制,由于增加的可扩展性而产生最小的内存或训练开销。我们的语音增强实验表明,所提出的分块优化方法实现了所需的可扩展性,与相应的端到端训练模型相比,性能仅略有下降。
In this paper, we present a blockwise optimization method for masking-based networks (BLOOM-Net) for training scalable speech enhancement networks. Here, we design our network with a residual learning scheme and train the internal separator blocks sequentially to obtain a scalable masking-based deep neural network for speech enhancement. Its scalability lets it dynamically adjust the run-time complexity depending on the test time environment. To this end, we modularize our models in that they can flexibly accommodate varying needs for enhancement performance and constraints on the resources, incurring minimal memory or training overhead due to the added scalability. Our experiments on speech enhancement demonstrate that the proposed blockwise optimization method achieves the desired scalability with only a slight performance degradation compared to corresponding models trained end-to-end.