Deep multiple instance learning for foreground speech localization in ambient audio from wearable devices.

Deep multiple instance learning for foreground speech localization in ambient audio from wearable devices.
复制标题

深层实例学习可穿戴设备的环境音频中的前景语音本地化。

DOI:
10.1186/s13636-020-00194-0
复制
发表时间:
2021
期刊:
EURASIP journal on audio, speech, and music processing
影响因子:
--
通讯作者:
Narayanan S
Narayanan S
中科院分区:
其他
文献类型:
--
作者:
Hebbar R;Papadopoulos P;Reyes R;Danvers AF;Polsinelli AJ;Moseley SA;Sbarra DA;Mehl MR;Narayanan S

文献摘要

参考文献

被引文献

相似文献

近年来,机器学习技术已被用于在几个与音频相关的任务中产生最先进的结果。这些方法的成功在很大程度上归功于获得大量开放源码数据集和计算资源的增强。然而,这些方法的一个缺点是,由于领域不匹配,它们往往无法很好地概括现实生活场景中的任务。一个这样的任务是从可穿戴音频设备进行前景语音检测。一些干扰因素,如动态变化的环境条件,包括背景说话人、电视或无线电音频,使得前景语音检测成为一项具有挑战性的任务。此外,获取用于分析和模型训练的准确的音频流的时刻到时刻的注释也是耗时和昂贵的。在这项工作中,我们使用多实例学习(MIL)来促进使用较低时间分辨率(粗标记)的注释开发这样的模型。我们展示了如何应用MIL来定位粗标签音频中的前景语音,并显示了袋级和实例级的结果。我们还研究了不同的池方法,以及它们如何适应我们在应用程序中观察到的密集分布的事件。最后,我们使用语音活动检测嵌入作为前景检测的特征,给出了改进。
Over the recent years, machine learning techniques have been employed to produce state-of-the-art results in several audio related tasks. The success of these approaches has been largely due to access to large amounts of open-source datasets and enhancement of computational resources. However, a shortcoming of these methods is that they often fail to generalize well to tasks from real life scenarios, due to domain mismatch. One such task is foreground speech detection from wearable audio devices. Several interfering factors such as dynamically varying environmental conditions, including background speakers, TV, or radio audio, render foreground speech detection to be a challenging task. Moreover, obtaining precise moment-to-moment annotations of audio streams for analysis and model training is also time-consuming and costly. In this work, we use multiple instance learning (MIL) to facilitate development of such models using annotations available at a lower time-resolution (coarsely labeled). We show how MIL can be applied to localize foreground speech in coarsely labeled audio and show both bag-level and instance-level results. We also study different pooling methods and how they can be adapted to densely distributed events as observed in our application. Finally, we show improvements using speech activity detection embeddings as features for foreground detection.
DOI: 10.1109/tbme.2014.2309951
发表时间: 2014-05
期刊: IEEE transactions on bio-medical engineering
影响因子: --
作者:
Zheng YL;Ding XR;Poon CC;Lo BP;Zhang H;Zhou XL;Yang GZ;Zhao N;Zhang YT
通讯作者: Zhang YT
DOI: 10.1145/3326362
发表时间: 2019-11-01
影响因子: 6.2
作者:
Wang, Yue;Sun, Yongbin;Solomon, Justin M.
通讯作者: Solomon, Justin M.
DOI: 10.1109/jsen.2014.2357257
发表时间: 2015-06-01
影响因子: 4.3
作者:
Rodgers, Mary M.;Pai, Vinay M.;Conroy, Richard S.
通讯作者: Conroy, Richard S.
DOI: 10.1037/pspp0000272
发表时间: 2020-12-01
影响因子: 7.6
作者:
Sun, Jessie;Harris, Kelci;Vazire, Simine
通讯作者: Vazire, Simine
DOI: 10.1109/tpami.2016.2535231
发表时间: 2017-01-01
影响因子: 23.6
作者:
Cinbis, Ramazan Gokberk;Verbeek, Jakob;Schmid, Cordelia
通讯作者: Schmid, Cordelia