Discriminative multiple sound source localization based on deep neural networks using independent location model

Discriminative multiple sound source localization based on deep neural networks using independent location model
复制标题

DOI:
10.1109/slt.2016.7846325
复制
发表时间:
2016-12
期刊:
2016 IEEE Spoken Language Technology Workshop (SLT)
影响因子:
--
通讯作者:
Ryu Takeda;Kazunori Komatani
Ryu Takeda;Kazunori Komatani
中科院分区:
其他
文献类型:
--
作者:
Ryu Takeda;Kazunori Komatani

文献摘要

被引文献

相似文献

提出了一种基于深度神经网络的多声源定位训练方法。这种网络根据位置标签作为声音位置的后验概率估计器,实现了高的定位正确率。由于以前的DNNS的SSL配置处理单一声源的情况,因此应该将其扩展到多个声源的情况,以便将其应用到实际环境中。然而,天真的设计导致1)标签和训练数据模式的数量增加,以及2)在不同数量的声源上缺乏标签一致性,例如一个或两个或更多声音情况。这两个问题是用我们提出的方法解决的,前者包括一个独立的位置模型,后者包括一个块一致标号和排序。实验表明,本文提出的训练方法训练的基于DNN的隐马尔可夫模型在块级正确率方面比传统的隐马尔可夫模型提高了18个百分点。
We propose a training method for multiple sound source localization (SSL) based on deep neural networks (DNNs). Such networks function as posterior probability estimator of sound location in terms of position labels and achieve high localization correctness. Since the previous DNNs' configuration for SSL handles one-sound-source cases, it should be extended to multiple-sound-source cases to apply it to real environments. However, a naïve design causes 1) an increase in the number of labels and training data patterns and 2) a lack of label consistency across different numbers of sound sources, such as one and two-or-more-sound cases. These two problems were solved using our proposed method, which involves an independent location model for the former and an block-wise consistent labeling with ordering for the latter. Our experiments indicated that the SSL based on DNNs trained by our proposed training method out-performed a conventional SSL method by a maximum of 18 points in terms of block-level correctness.