Discovering Collective Variables of Molecular Transitions via Genetic Algorithms and Neural Networks.

Discovering Collective Variables of Molecular Transitions via Genetic Algorithms and Neural Networks.
复制标题

DOI:
10.1021/acs.jctc.0c00981
复制
发表时间:
2021-04-13
影响因子:
5.5
通讯作者:
Ensing B
Ensing B
中科院分区:
化学1区
文献类型:
--
作者:
Hooft F;Pérez de Alba Ortíz A;Ensing B

文献摘要

参考文献

被引文献

相似文献

随着计算机硬件和算法的不断改进,模拟已经成为理解各种(生物)分子过程的有力工具。为了处理大型模拟数据集并加速缓慢的活化跃迁,需要一组浓缩的描述符或集体变量(CV)来辨别描述感兴趣的分子过程的相关动力学。然而,提出一组足够的CV,可以捕获的内在反应坐标的分子过渡往往是非常困难的。在这里,我们提出了一个框架,找到一个最佳的一组CV的候选人使用人工神经网络和遗传算法的组合。该方法有效地用基因代替了自动编码器网络的编码器来表示潜在空间,即,CV。给定一组CV作为输入,网络将被训练以恢复过渡沿着点处CV值的原子坐标。网络性能被用作输入CV的适应度的估计器。两种遗传算法优化CV选择和神经网络结构。成功检索的最佳CV由这个框架说明在两个案例研究的手:众所周知的丙氨酸二肽分子的构象变化和更复杂的过渡B-DNA中的一个碱基对从经典的沃森-克里克配对的替代Hoogsteen配对。我们的框架的主要优点包括:最佳可解释CV,避免提交者或时间相关函数的昂贵计算,以及自动超参数优化。此外,我们表明,应用网络输入和输出之间的时间延迟允许增强慢变量的选择。此外,该网络还可以用于生成未探索的微观状态的分子构型,例如,用于增强模拟数据。
With the continual improvement of computing hardware and algorithms, simulations have become a powerful tool for understanding all sorts of (bio)molecular processes. To handle the large simulation data sets and to accelerate slow, activated transitions, a condensed set of descriptors, or collective variables (CVs), is needed to discern the relevant dynamics that describes the molecular process of interest. However, proposing an adequate set of CVs that can capture the intrinsic reaction coordinate of the molecular transition is often extremely difficult. Here, we present a framework to find an optimal set of CVs from a pool of candidates using a combination of artificial neural networks and genetic algorithms. The approach effectively replaces the encoder of an autoencoder network with genes to represent the latent space, i.e., the CVs. Given a selection of CVs as input, the network is trained to recover the atom coordinates underlying the CV values at points along the transition. The network performance is used as an estimator of the fitness of the input CVs. Two genetic algorithms optimize the CV selection and the neural network architecture. The successful retrieval of optimal CVs by this framework is illustrated at the hand of two case studies: the well-known conformational change in the alanine dipeptide molecule and the more intricate transition of a base pair in B-DNA from the classic Watson–Crick pairing to the alternative Hoogsteen pairing. Key advantages of our framework include the following: optimal interpretable CVs, avoiding costly calculation of committor or time-correlation functions, and automatic hyperparameter optimization. In addition, we show that applying a time-delay between the network input and output allows for enhanced selection of slow variables. Moreover, the network can also be used to generate molecular configurations of unexplored microstates, for example, for augmentation of the simulation data.
DOI: 10.1103/physreve.97.062412
发表时间: 2018-06
期刊: Physical review. E
影响因子: --
作者:
Hernández CX;Wayment-Steele HK;Sultan MM;Husic BE;Pande VS
通讯作者: Pande VS
DOI: 10.1073/pnas.100127697
发表时间: 2000-05-23
影响因子: 11.1
作者:
Bolhuis, PG;Dellago, C;Chandler, D
通讯作者: Chandler, D
DOI: 10.1038/nmeth.3658
发表时间: 2016-01
期刊: Nature methods
影响因子: 48
作者:
Ivani I;Dans PD;Noy A;Pérez A;Faustino I;Hospital A;Walther J;Andrio P;Goñi R;Balaceanu A;Portella G;Battistini F;Gelpí JL;González C;Vendruscolo M;Laughton CA;Harris SA;Case DA;Orozco M
通讯作者: Orozco M
DOI: 10.1063/1.470117
发表时间: 1995-11-15
影响因子: 4.4
作者:
ESSMANN, U;PERERA, L;PEDERSEN, LG
通讯作者: PEDERSEN, LG
DOI: 10.1073/pnas.202427399
发表时间: 2002-10-01
影响因子: 11.1
作者:
Laio, A;Parrinello, M
通讯作者: Parrinello, M