Adaptive Stochastic Natural Gradient Method for Optimizing Functions with Low Effective Dimensionality

Adaptive Stochastic Natural Gradient Method for Optimizing Functions with Low Effective Dimensionality
复制标题

DOI:
10.1007/978-3-030-58112-1_50
复制
发表时间:
2020
期刊:
2021 IEEE Congress on Evolutionary Computation (CEC)
影响因子:
--
通讯作者:
Teppei Yamaguchi;Kento Uchida;S. Shirakawa
Teppei Yamaguchi;Kento Uchida;S. Shirakawa
中科院分区:
其他
文献类型:
--
作者:
Teppei Yamaguchi;Kento Uchida;S. Shirakawa

文献摘要

相似文献

黑盒优化算法,如进化算法,已被认为是现实世界应用程序的有用工具。几种高效的基于概率模型的进化算法,如紧凑遗传算法(CGA)和协方差矩阵自适应进化策略(CMA-ES),可以看作是统计流形上的随机自然梯度上升。我们的基线算法是自适应随机自然梯度(ASNG)方法,它根据近似自然梯度的信噪比(SNR)自动调整学习率。自适应神经网络在实际应用中取得了很好的效果,但当目标函数的有效维度(LED)较低时,收敛速度变差,即部分设计变量失效或对目标值影响不大。本文提出了一种基于单元信噪比的近似自然梯度的单元平差方法,并将该方法引入到ASNG中。该方法抑制了低信噪比的自然梯度元素,加快了自适应神经网络的学习速率自适应。我们将该方法应用到CGA中,并在二进制优化的基准函数上验证了该方法的有效性。
Black-box optimization algorithms, such as evolutionary algorithms, have been recognized as useful tools for real-world applications. Several efficient probabilistic model-based evolutionary algorithms, such as the compact genetic algorithm (cGA) and the covariance matrix adaptation evolution strategy (CMA-ES), can be regarded as a stochastic natural gradient ascent on statistical manifolds. Our baseline algorithm is the adaptive stochastic natural gradient (ASNG) method which automatically adapts the learning rate based on the signal-to-noise ratio (SNR) of the approximated natural gradient. ASNG has shown effectiveness in a practical application, but the convergence speed of ASNG deteriorates on objective functions with low effective dimensionality (LED), where LED means that part of the design variables is ineffective or does not affect the objective value significantly. In this paper, we propose an element-wise adjustment method for the approximated natural gradient based on the element-wise SNR and introduce the proposed adjustment method into ASNG. The proposed method suppresses the natural gradient elements with the low SNRs, helping to accelerate the learning rate adaptation in ASNG. We incorporate the proposed method into the cGA and demonstrate the effectiveness of the proposed method on the benchmark functions of binary optimization.