ReaxFF Parameter Optimization with Monte-Carlo and Evolutionary Algorithms: Guidelines and Insights

ReaxFF Parameter Optimization with Monte-Carlo and Evolutionary Algorithms: Guidelines and Insights
复制标题

DOI:
10.1021/acs.jctc.9b00769
复制
发表时间:
2019-12-01
影响因子:
5.5
通讯作者:
Verstraelen, Toon
Verstraelen, Toon
中科院分区:
化学1区
文献类型:
--
作者:
Shchygol, Ganna;Yakovlev, Alexei;Verstraelen, Toon

文献摘要

被引文献

相似文献

ReaxFF 是一种计算高效的力场,如果感兴趣的化学物质有可靠的力场参数,可以模拟具有多种化学物质的扩展分子模型中的复杂反应动力学。如果不是,则必须通过最小化 ReaxFF 在相关训练集上产生的误差来优化它们。由于这种优化并非微不足道,因此已经开发了许多方法,特别是遗传算法(GA)来搜索参数空间中的全局最优值。最近,提出了两种替代参数校准技术,即蒙特卡罗力场优化器(MCFF)和协方差矩阵自适应进化策略(CMA-ES)。在这项工作中,使用文献中的三个训练集系统地比较了 CMA-ES、MCFF 和 GA 方法 (OGOLEM)。通过使用不同的随机种子和初始参数猜测重复优化,结果表明,不应盲目信任使用任何这些方法运行的单个优化:不可再现、收敛不良或过早收敛是常见的缺陷。 GA 显示陷入局部最小值的风险最小,而 CMA-ES 能够在三分之二的情况下达到最低错误,尽管不是系统性的。对于每种方法,我们都提供了合理的默认设置,并且我们的分析为它们在未来工作中的使用提供了有用的指导。影响参数优化的一个重要副作用是数值噪声。详细的分析表明,可以通过例如在训练集中专门使用明确的几何优化来减少它。即使没有这种噪声,也可以找到许多不同的接近最优的参数向量,这为改进训练集和检测过度拟合伪影开辟了新途径。
ReaxFF is a computationally efficient force field to simulate complex reactive dynamics in extended molecular models with diverse chemistries, if reliable force field parameters are available for the chemistry of interest. If not, they must be optimized by minimizing the error ReaxFF makes on a relevant training set. Because this optimization is far from trivial, many methods, in particular, genetic algorithms (GAs), have been developed to search for the global optimum in parameter space. Recently, two alternative parameter calibration techniques were proposed, that is, Monte-Carlo force field optimizer (MCFF) and covariance matrix adaptation evolutionary strategy (CMA-ES). In this work, CMA-ES, MCFF, and a GA method (OGOLEM) are systematically compared using three training sets from the literature. By repeating optimizations with different random seeds and initial parameter guesses, it is shown that a single optimization run with any of these methods should not be trusted blindly: nonreproducible, poor or premature convergence is a common deficiency. GA shows the smallest risk of getting trapped into a local minimum, whereas CMA-ES is capable of reaching the lowest errors for two-third of the cases, although not systematically. For each method, we provide reasonable default settings, and our analysis offers useful guidelines for their usage in future work. An important side effect impairing parameter optimization is numerical noise. A detailed analysis reveals that it can be reduced, for example, by using exclusively unambiguous geometry optimization in the training set. Even without this noise, many distinct near-optimal parameter vectors can be found, which opens new avenues for improving the training set and detecting overfitting artifacts.