Block-Online Guided Source Separation

Block-Online Guided Source Separation
复制标题

块在线引导源分离

DOI:
--
复制
发表时间:
2020
期刊:
Spoken Language Technology Workshop
影响因子:
--
通讯作者:
Kenji Nagamatsu
Kenji Nagamatsu
中科院分区:
--
文献类型:
--
作者:
Shota Horiguchi;Yusuke Fujita;Kenji Nagamatsu

文献摘要

被引文献

相似文献

提出了一种块在线的引导源分离算法。GSS是一种使用日志信息来更新观测信号的生成模型的参数的语音分离方法。先前的研究表明,GSS在多人场景中表现良好。然而,它需要大量的计算时间,这是部署在线应用程序的障碍。还有一个问题是,离线GSS是一种基于话语的算法,因此它根据话语的长度产生延迟。与所提出的算法,分块的输入样本和相应的时间注释连接在一起,在前面的上下文中,并用于更新参数。使用上下文使算法能够准确地估计时频掩模仅从一个迭代的优化为每个块,其延迟不依赖于话语长度,但预定的块长度。它还通过仅更新每个块及其上下文中的活动扬声器的参数来降低计算成本。在CHiME-6语料库和一个会议语料库上的测试表明,该算法与传统的离线GSS算法性能相当,但计算速度提高了32倍,足以满足实时应用的要求。
We propose a block-online algorithm of guided source separation (GSS). GSS is a speech separation method that uses diarization information to update parameters of the generative model of observation signals. Previous studies have shown that GSS performs well in multi-talker scenarios. However, it requires a large amount of calculation time, which is an obstacle to the deployment of online applications. It is also a problem that the offline GSS is an utterance-wise algorithm so that it produces latency according to the length of the utterance. With the proposed algorithm, block-wise input samples and corresponding time annotations are concatenated with those in the preceding context and used to update the parameters. Using the context enables the algorithm to estimate time-frequency masks accurately only from one iteration of optimization for each block, and its latency does not depend on the utterance length but predetermined block length. It also reduces calculation cost by updating only the parameters of active speakers in each block and its context. Evaluation on the CHiME-6 corpus and a meeting corpus showed that the proposed algorithm achieved almost the same performance as the conventional offline GSS algorithm but with 32x faster calculation, which is sufficient for real-time applications.