Self-paced Data Augmentation for Training Neural Networks

Self-paced Data Augmentation for Training Neural Networks
复制标题

DOI:
10.1016/j.neucom.2021.02.080
复制
发表时间:
2020-10
期刊:
ArXiv
影响因子:
--
通讯作者:
Tomoumi Takase;Ryo Karakida;H. Asoh
Tomoumi Takase;Ryo Karakida;H. Asoh
中科院分区:
其他
文献类型:
--
作者:
Tomoumi Takase;Ryo Karakida;H. Asoh

文献摘要

被引文献

相似文献

数据增强被广泛用于机器学习;然而,尚未建立应用数据增强的有效方法,即使它包括应仔细调整的几个因素。其中一个因素是样本适用性,这涉及选择适合数据扩充的样本。将数据增强应用于所有训练样本的典型方法忽略了样本适用性,这可能会降低分类器性能。为了解决这个问题,我们提出了自步调增强(SPA),以自动和动态地选择合适的样本进行数据增强训练神经网络时。所提出的方法减轻了无效的数据增强所造成的泛化性能的恶化。我们讨论了两个原因,建议SPA工程相对于课程学习和理想的变化损失函数不稳定性。实验结果表明,所提出的集对分析可以提高泛化性能,尤其是在训练样本数量较少时。此外,所提出的SPA优于最先进的RandAugment方法。
Data augmentation is widely used for machine learning; however, an effective method to apply data augmentation has not been established even though it includes several factors that should be tuned carefully. One such factor is sample suitability, which involves selecting samples that are suitable for data augmentation. A typical method that applies data augmentation to all training samples disregards sample suitability, which may reduce classifier performance. To address this problem, we propose the self-paced augmentation (SPA) to automatically and dynamically select suitable samples for data augmentation when training a neural network. The proposed method mitigates the deterioration of generalization performance caused by ineffective data augmentation. We discuss two reasons the proposed SPA works relative to curriculum learning and desirable changes to loss function instability. Experimental results demonstrate that the proposed SPA can improve the generalization performance, particularly when the number of training samples is small. In addition, the proposed SPA outperforms the state-of-the-art RandAugment method.