Group Masked Model Learning for General Audio Representation

Group Masked Model Learning for General Audio Representation
复制标题

DOI:
10.1109/icip49359.2023.10222551
复制
发表时间:
2023-10
期刊:
2023 IEEE International Conference on Image Processing (ICIP)
影响因子:
--
通讯作者:
Sara Atito;Muhammad Awais;Tony Alex;Josef Kittler
Sara Atito;Muhammad Awais;Tony Alex;Josef Kittler
中科院分区:
其他
文献类型:
--
作者:
Sara Atito;Muhammad Awais;Tony Alex;Josef Kittler

文献摘要

相似文献

视觉转换器最近在计算机视觉和音频社区引起了极大的兴趣,因为它们在学习远程关系方面具有灵活性。然而,众所周知,变压器非常需要数据,这需要更多数量级的数据[1]来训练。这推动了音频转换器自监督预训练的研究,它减少了对大量标记数据的依赖,并专注于提取音频谱图的简明表示。在本文中,我们提出了Audio-GMML,一种基于组掩蔽模型学习(GMML)的通用音频表示的自监督转换器,以及一种块聚合策略来提高学习表示的性能并加强给定音频的全局结构。我们在几个下游任务上对我们的预训练模型进行了评估,在五个音频和语音分类任务上设置了新的最先进的性能。代码和预先训练的重量将向科学界公开。
Vision transformers have recently generated significant interest in the computer vision and audio communities due to their flexibility in learning long-range relationships. However, transformers are known to be data hungry which require orders of magnitude more data [1] to train. This has motivated the research in self-supervised pretraining of audio transformers, which reduces the dependency on large amounts of labeled data and focuses on extracting concise representation of the audio spectrograms. In this paper, we propose Audio-GMML, a self-supervised transformer for general audio representations that is based on Group Masked Model Learning (GMML) and a patch aggregation strategy to improve the performance of learned representations and enforce global structure of the given audio. We evaluate our pretrained models on several downstream tasks, setting a new state-of-the-art performance on five audio and speech classification tasks. The code and pretrained weights will be made publicly available for the scientific community.