Conformer-Based Self-Supervised Learning For Non-Speech Audio Tasks

Conformer-Based Self-Supervised Learning For Non-Speech Audio Tasks
复制标题

非语音音频任务的基于一致性的自监督学习

DOI:
--
复制
发表时间:
2021
期刊:
IEEE International Conference on Acoustics, Speech, and Signal Processing
影响因子:
--
通讯作者:
Yatharth Saraf
Yatharth Saraf
中科院分区:
--
文献类型:
--
作者:
Sangeeta Srivastava;Yun Wang;Andros Tjandra;Anurag Kumar;Chunxi Liu;Kritika Singh;Yatharth Saraf

文献摘要

被引文献

相似文献

从未标记的数据中学习的代表性在人工智能研究中引起了重大兴趣。尽管自我监督的语音表示学习在语音研究社区中很受欢迎,但很少有作品全面分析了非语音音频任务的音频表示。在本文中,我们提出了一种自我监督的音频表示方法,并将其应用于各种下游非语音音频任务。我们将著名的WAV2VEC 2.0框架结合在一起,该框架在对语音任务的自我监督学习方面的成功以及参数效率高效的构象体架构中的成功。我们的自我监督预训练可以将标记数据的需求减少三分之二。在音频集基准上,我们获得了平均平均精度(MAP)分数为0.415,这是该数据集的新最新技术,这是仅通过音频自我监督的学习。我们的微调构象异构体还超越或匹配以前训练的以前系统在几个下游任务上预先训练的系统的性能。我们进一步讨论了预训练和微调的重要设计考虑因素。
Representation learning from unlabeled data has been of major interest in artificial intelligence research. While self-supervised speech representation learning has been popular in the speech research community, very few works have comprehensively analyzed audio representation learning for non-speech audio tasks. In this paper, we propose a self-supervised audio representation learning method and apply it to a variety of downstream non-speech audio tasks. We combine the well-known wav2vec 2.0 framework, which has shown success in self-supervised learning for speech tasks, with parameter-efficient conformer architectures. Our self-supervised pre-training can reduce the need for labeled data by two-thirds. On the AudioSet benchmark, we achieve a mean average precision (mAP) score of 0.415, which is a new state-of-the-art on this dataset through audio-only self-supervised learning. Our fine-tuned conformers also surpass or match the performance of previous systems pre-trained in a supervised way on several downstream tasks. We further discuss the important design considerations for both pre-training and fine-tuning.