Learning Higher Representations from Pre-Trained Deep Models with Data Augmentation for the COMPARE 2020 Challenge Mask Task

Learning Higher Representations from Pre-Trained Deep Models with Data Augmentation for the COMPARE 2020 Challenge Mask Task
复制标题

DOI:
10.21437/interspeech.2020-1552
复制
发表时间:
2020-10
期刊:
--
影响因子:
--
通讯作者:
Tomoya Koike;Kun Qian;Björn Schuller;Yoshiharu Yamamoto
Tomoya Koike;Kun Qian;Björn Schuller;Yoshiharu Yamamoto
中科院分区:
其他
文献类型:
--
作者:
Tomoya Koike;Kun Qian;Björn Schuller;Yoshiharu Yamamoto

文献摘要

相似文献

人类手工制作的特征在几乎所有与机器学习相关的任务中总是被认为是昂贵、耗时和困难的。首先,这些精心设计的功能非常依赖人类专家领域知识,这可能会限制跨领域的协作工作。其次,在这种暴力场景中提取的特征可能不容易转移到另一个任务中,这意味着需要设计一系列新的特征。为此,我们介绍了一种基于迁移学习策略与数据增强技术相结合的方法,用于C OM P AR E 2020挑战面具子挑战。与以往主要基于图像数据的预训练模型不同,我们使用了基于大规模音频数据的预训练模型,即,AudioSet。此外,SpecAugment和mixup方法用于改进深度模型的泛化。实验结果表明,最佳建议的模型可以显着(p < . 001,通过单尾z检验)将测试集上的未加权平均召回率(UAR)从71.8%(基线)提高到76.2%。最后,最好的结果,即,测试集上77.5%的UAR是通过两个最佳建议模型和基线中最佳单个模型的后期融合实现的。
Human hand-crafted features are always regarded as expensive, time-consuming, and difficult in almost all of the machine-learning-related tasks. First, those well-designed features extremely rely on human expert domain knowledge, which may restrain the collaboration work across fields. Second, the features extracted in such a brute-force scenario may not be easy to be transferred to another task, which means a series of new features should be designed. To this end, we introduce a method based on a transfer learning strategy combined with data augmentation techniques for the C OM P AR E 2020 Challenge Mask Sub-Challenge. Unlike the previous studies mainly based on pre-trained models by image data, we use a pre-trained model based on large scale audio data, i.e., AudioSet. In addition, the SpecAugment and mixup methods are used to improve the generalisation of the deep models. Experimental results demonstrate that the best-proposed model can significantly ( p < . 001 , by one-tailed z -test) improve the unweighted average recall (UAR) from 71.8% (baseline) to 76.2% on the test set. Finally, the best result, i.e., 77.5% of the UAR on the test set, is achieved by a late fusion of the two best proposed models and the best single model in the baseline.