Multi-Channel Talker-Independent Speaker Separation Through Location-Based Training

Multi-Channel Talker-Independent Speaker Separation Through Location-Based Training
复制标题

DOI:
10.1109/taslp.2022.3202129
复制
发表时间:
2022
期刊:
IEEE/ACM Transactions on Audio, Speech, and Language Processing
影响因子:
--
通讯作者:
H. Taherian;Ke Tan;Deliang Wang
H. Taherian;Ke Tan;Deliang Wang
中科院分区:
其他
文献类型:
--
作者:
H. Taherian;Ke Tan;Deliang Wang

文献摘要

相似文献

排列歧义是基于深度学习的独立于说话者的说话者分离的一个关键问题。深度聚类和排列不变训练(PIT)已被广泛用于解决单声道场景中的排列模糊问题。尽管这两种方法都已扩展到多麦克风场景,但我们相信通过利用多个说话者的空间关系可以自然地避免排列歧义问题。在本文中,我们提出了基于位置的训练(LBT),这是一种在多通道说话者分离中实现说话者独立性的新方法。与检查所有可能排列的 PIT 不同,LBT 根据发言者在物理空间中的位置来分配发言者。由于训练复杂度与并发说话者数量呈线性关系,LBT 在计算上比具有阶乘复杂度的 PIT 高效得多,特别是当需要分离大量重叠的说话者时。具体来说,我们提出了两个训练标准:基于方位角和基于距离的训练,分别使用扬声器相对于麦克风阵列的方位角和距离。评估结果表明,在不同阵列几何形状和各种声学条件下的两扬声器和三扬声器混合中,LBT 的性能显着优于 PIT。此外,我们提出了一种联合训练策略,将基于方位角和基于距离的训练相结合,进一步提高了分离性能。
Permutation ambiguity is a crucial issue for deep learning based talker-independent speaker separation. Deep clustering and permutation invariant training (PIT) have been widely used to address the permutation ambiguity problem in monaural scenarios. Although both approaches have been extended to multi-microphone scenarios, we believe that the permutation ambiguity problem can be naturally avoided by leveraging the spatial relations of multiple speakers. In this article, we present location-based training (LBT), a new approach to achieve talker independency in multi-channel speaker separation. Unlike PIT that examines all possible permutations, LBT assigns speakers according to their positions in physical space. With a linear training complexity to the number of concurrent speakers, LBT is computationally much more efficient than PIT with a factorial complexity, particularly when a large number of overlapping speakers needs to be separated. Specifically, we propose two training criteria: azimuth-based and distance-based training, using speaker azimuths and distances relative to a microphone array, respectively. Evaluation results show that LBT significantly outperforms PIT on two-speaker and three-speaker mixtures with different array geometries and in various acoustic conditions. In addition, we propose a joint training strategy to integrate azimuth-based and distance-based training, which further improves separation performance.