Multitask machine learning of collective variables for enhanced sampling of rare events

Multitask machine learning of collective variables for enhanced sampling of rare events
复制标题

DOI:
10.1021/acs.jctc.1c00143
复制
发表时间:
2020-12
影响因子:
5.5
通讯作者:
Lixin Sun;Jonathan Vandermause;Simon L. Batzner;Yu Xie;David J. Clark;Wei Chen;B. Kozinsky
Lixin Sun;Jonathan Vandermause;Simon L. Batzner;Yu Xie;David J. Clark;Wei Chen;B. Kozinsky
中科院分区:
化学1区
文献类型:
--
作者:
Lixin Sun;Jonathan Vandermause;Simon L. Batzner;Yu Xie;David J. Clark;Wei Chen;B. Kozinsky

文献摘要

相似文献

计算准确的反应速率是计算化学和生物学中的一个中心挑战,因为用无偏分子动力学估计自由能的成本很高。在这项工作中,设计了一种数据驱动的机器学习算法来学习多任务神经网络的集合变量,其中共同的上游部分将原子构型的高维降维到低维潜在空间,而单独的下游部分将潜在空间映射到盆地类别标签和势能的预测。由此产生的潜在空间被证明是一个有效的低维表示,捕捉到反应过程并指导有效的伞状采样以获得准确的自由能景观。该方法被成功地应用于模型体系,包括5D Müler Brown模型、5D三势垒模型、真空中的丙氨酸二肽和Au(110)表面重构单元反应。它实现了复杂系统中能量控制反应的自动降维,提供了一个统一的、数据高效的框架,可以用有限的数据进行训练,并且性能优于包括自动编码器在内的单任务学习方法。
Computing accurate reaction rates is a central challenge in computational chemistry and biology because of the high cost of free energy estimation with unbiased molecular dynamics. In this work, a data-driven machine learning algorithm is devised to learn collective variables with a multitask neural network, where a common upstream part reduces the high dimensionality of atomic configurations to a low dimensional latent space and separate downstream parts map the latent space to predictions of basin class labels and potential energies. The resulting latent space is shown to be an effective low-dimensional representation, capturing the reaction progress and guiding effective umbrella sampling to obtain accurate free energy landscapes. This approach is successfully applied to model systems including a 5D Müller Brown model, a 5D three-well model, the alanine dipeptide in vacuum, and an Au(110) surface reconstruction unit reaction. It enables automated dimensionality reduction for energy controlled reactions in complex systems, offers a unified and data-efficient framework that can be trained with limited data, and outperforms single-task learning approaches, including autoencoders.