Improving Label Assignments Learning by Dynamic Sample Dropout Combined with Layer-wise Optimization in Speech Separation

Improving Label Assignments Learning by Dynamic Sample Dropout Combined with Layer-wise Optimization in Speech Separation
复制标题

DOI:
10.21437/interspeech.2023-1172
复制
发表时间:
2023-08
期刊:
--
影响因子:
--
通讯作者:
Chenyu Gao;Yue Gu;I. Marsic
Chenyu Gao;Yue Gu;I. Marsic
中科院分区:
其他
文献类型:
--
作者:
Chenyu Gao;Yue Gu;I. Marsic

文献摘要

相似文献

在有监督语音分离中,置换不变量训练(PIT)通过选择最佳的置换来更新模型,从而被广泛地用于处理标签歧义。尽管它取得了成功,但以往的研究表明,PIT在相邻的历元上受到过度标签分配切换的困扰,阻碍了模型学习更好的标签分配。为了解决这个问题,我们提出了一种新的训练策略,动态样本丢弃(DSD),它考虑了以前最好的标签分配和评估度量,以排除在训练过程中可能对学习的标签分配产生负面影响的样本。此外,我们还包括分层优化(LO),以通过解决层解耦来提高性能。我们的实验表明,DSD和LO的组合性能优于基线,并解决了过多的标签分配、切换和层分离问题。所提出的DSD和LO方法实现简单,不需要额外的训练集或步骤,对各种语音分离任务具有通用性。
In supervised speech separation, permutation invariant training (PIT) is widely used to handle label ambiguity by selecting the best permutation to update the model. Despite its success, previous studies showed that PIT is plagued by excessive label assignment switching in adjacent epochs, impeding the model to learn better label assignments. To address this issue, we propose a novel training strategy, dynamic sample dropout (DSD), which considers previous best label assignments and evaluation metrics to exclude the samples that may negatively impact the learned label assignments during training. Additionally, we include layer-wise optimization (LO) to improve the performance by solving layer-decoupling. Our experiments showed that combining DSD and LO outperforms the baseline and solves excessive label assignment switching and layer-decoupling issues. The proposed DSD and LO approach is easy to implement, requires no extra training sets or steps, and shows generality to various speech separation tasks.