Bandit-Based Search for Constraint Programming

Bandit-Based Search for Constraint Programming
复制标题

DOI:
10.1007/978-3-642-40627-0_36
复制
发表时间:
2013-09
期刊:
--
影响因子:
--
通讯作者:
Manuel Loth;M. Sebag;Y. Hamadi;Marc Schoenauer
Manuel Loth;M. Sebag;Y. Hamadi;Marc Schoenauer
中科院分区:
其他
文献类型:
--
作者:
Manuel Loth;M. Sebag;Y. Hamadi;Marc Schoenauer

文献摘要

被引文献

相似文献

约束编程 (CP) 求解器通常使用基于树搜索的启发式方法来探索解空间。蒙特卡罗树搜索(MCTS)旨在不确定性下做出最优顺序决策,根据指定的奖励函数逐渐生长搜索树以探索最有希望的区域。在 CP 和 MCTS 的交叉点,本文提出了约束规划的 Bandit 搜索 (BaSCoP) 算法,使 MCTS 适应 CP 搜索的具体情况。这一贡献依赖于 i) 适合 CP 并与多次重启策略兼容的通用奖励函数; ii) 在 Gecode 约束求解器顶部的 MCTS.BaSCoP 中使用深度优先搜索作为展开程序,可以显着改进某些 CP 基准套件上的深度优先搜索,证明其作为通用而强大的 CP 搜索方法的相关性。
Constraint Programming (CP) solvers classically explore the solution space using tree-search based heuristics. Monte-Carlo Tree Search (MCTS), aimed at optimal sequential decision making under uncertainty, gradually grows a search tree to explore the most promising regions according to a specified reward function. At the crossroad of CP and MCTS, this paper presents the Bandit Search for Constraint Programming (BaSCoP) algorithm, adapting MCTS to the specifics of the CP search. This contribution relies on i) a generic reward function suited to CP and compatible with a multiple restart strategy; ii) the use of depth-first search as roll-out procedure in MCTS.BaSCoP, on the top of the Gecode constraint solver, is shown to significantly improve on depth-first search on some CP benchmark suites, demonstrating its relevance as a generic yet robust CP search method.