Evolutionary Approach for AutoAugment Using the Thermodynamical Genetic Algorithm

Evolutionary Approach for AutoAugment Using the Thermodynamical Genetic Algorithm
复制标题

DOI:
10.1609/aaai.v35i11.17184
复制
发表时间:
2021-05
期刊:
--
影响因子:
--
通讯作者:
Akira Terauchi;N. Mori
Akira Terauchi;N. Mori
中科院分区:
其他
文献类型:
--
作者:
Akira Terauchi;N. Mori

文献摘要

相似文献

数据增强是通过提高机器学习模型的泛化能力来稳定学习的最有效方法之一。近年来,自动数据增强方法,如AutoAugment或Fast AutoAugment引起了人们的关注;这些方法改善了图像分类和目标检测任务的结果。然而,仍然存在一些问题。最值得注意的是,更大的训练数据集需要更高的计算成本。当用小数据集进行搜索以试图确定数据增强方法时,真实数据空间和采样数据空间彼此不完全对应,从而导致泛化性能恶化。此外,在现有的自动增强方法中,搜索阶段往往是由一个例外的子策略,这导致操作的多样性的损失。在这项研究中,我们通过将进化计算引入到以前的方法中来解决这些问题。如前所述,保持多样性至关重要。因此,我们采用了顺序遗传算法(TDGA),它可以控制人口的多样性与一个特定的遗传算子,称为顺序选择规则。为了验证该方法的有效性,以CIFAR-10和SVHN两个基准数据集为例进行了计算实验。实验结果表明,该方法可以获得各种有用的增强子策略的问题,同时减少了计算成本。
Data augmentation is one of the most effective ways to stabilize learning by improving the generalization of machine-learning models. In recent years, automatic data augmentation methods, such as AutoAugment or Fast AutoAugment have been attracting attention; and these methods improved the results of image classification and object detection tasks. However, several problems remain. Most notably, a larger training dataset requires higher computational costs. When searching with a small dataset in an attempt to determine the data augmentation approach, the true data space and sampling data space do not fully correspond with each other, thereby causing the generalization performance to deteriorate. Moreover, in the existing automatic augmentation methods, the search phase is often dominated by an exceptional sub-policy, which results in a loss of diversity of operations. In this study, we solved these problems by introducing evolutionary computation to previous methods. As mentioned earlier, maintaining diversity is essential. Therefore, we adopted the thermodynamical genetic algorithm (TDGA), which can control the population diversity with a specific genetic operator, known as the thermodynamical selection rule. To confirm the effectiveness of the proposed method, computational experiments were conducted using two benchmark datasets, CIFAR-10 and SVHN, as examples. The experimental results show that the proposed method can obtain various useful augmentation sub-policies for the problems while reducing the computational cost.