MAB-OS: Multi-Armed Bandits Metaheuristic Optimizer Selection

MAB-OS: Multi-Armed Bandits Metaheuristic Optimizer Selection
复制标题

DOI:
10.1016/j.asoc.2022.109452
复制
发表时间:
2022-08
期刊:
Appl. Soft Comput.
影响因子:
--
通讯作者:
Kazem Meidani;S. Mirjalili;A. Farimani
Kazem Meidani;S. Mirjalili;A. Farimani
中科院分区:
其他
文献类型:
--
作者:
Kazem Meidani;S. Mirjalili;A. Farimani

文献摘要

被引文献

相似文献

元启发式算法是无导数优化器,用于估计优化问题的全局最优解。为了在挖掘和探索之间保持平衡,以及算法之间的性能互补性,引入了许多元启发式方法。在这项工作中,我们提出了一个基于多臂强盗(MAB)问题的框架,这是一种经典的强化学习(RL)方法,在优化过程中为每个优化问题智能地选择合适的优化器。这种在线算法选择技术利用算法的收敛行为,通过选择算法的更新规则来找到探索-利用的正确平衡,该更新规则具有最大的估计改进。通过对Harris Hawks Optimizer(HHO)、差分进化(DE)和Whale Optimization Algorithm(WOA)这三种武装匪徒进行实验,我们发现MAB Optimizer Selection(MAB-OS)框架在收敛速度和最终解方面对不同类型的适应度景观具有最佳的整体性能。用于本工作的数据和代码可在https://github.com/BaratiLab/MAB-OS上获得。
Metaheuristic algorithms are derivative-free optimizers designed to estimate the global optima for optimization problems. Keeping balance between exploitation and exploration and the performance complementarity between the algorithms have led to the introduction of quite a few metaheuristic methods. In this work, we propose a framework based on Multi-Armed Bandits (MAB) problem, which is a classical Reinforcement Learning (RL) method, to intelligently select a suitable optimizer for each optimization problem during the optimization process. This online algorithm selection technique leverages on the convergence behavior of the algorithms to find the right balance of exploration–exploitation by choosing the update rule of the algorithm with the most estimated improvement in the solution. By performing experiments with three armed-bandits being Harris Hawks Optimizer (HHO), Differential Evolution (DE), and Whale Optimization Algorithm (WOA), we show that the MAB Optimizer Selection (named as MAB-OS) framework has the best overall performance on different types of fitness landscapes in terms of both convergence rate and the final solution. The data and codes used for this work are available at: https://github.com/BaratiLab/MAB-OS.