A Novel Monte Carlo Procedure for Protein Design

A Novel Monte Carlo Procedure for Protein Design
复制标题

DOI:
--
复制
发表时间:
1997-11
期刊:
--
影响因子:
--
通讯作者:
Anders Irbäck;Carsten Peterson;F. Potthast;Erik Sandelin
Anders Irbäck;Carsten Peterson;F. Potthast;Erik Sandelin
中科院分区:
其他
文献类型:
--
作者:
Anders Irbäck;Carsten Peterson;F. Potthast;Erik Sandelin

文献摘要

被引文献

相似文献

提出了一种新的蛋白质模型序列优化方法。该方法继承了多伊奇和库罗斯基[Phys. Rev. Lett. 76,323(1996)]通过最大化条件概率而不是最小化能量函数,基于一种新颖且非常有效的多序列蒙特卡罗方案。通过构造,该方法确保设计的序列在语义上代表良好的文件夹。该算法在链长为N = 16、18和32的二维HP模型上得到了成功的探索,并可以很容易地推广到非格模型。此外,序列空间搜索的引导程序的设计,这使得非常大的链可行。PACS编号:87.15.By,87.10.+ e irback@thep.lu.se carsten@thep.lu.se frank@thep.lu.se erik@thep.lu.se蛋白质折叠的“逆”,序列优化,在药物设计的背景下具有最大的相关性。这个问题,这相当于找到最佳的氨基酸序列给定的目标结构,也已经在理解蛋白质折叠的粗粒度模型的折叠特性的背景下进行了研究。这样的模型由能量函数E(r,σ)描述,其中r = {r1,r2,.,rN}表示氨基酸坐标,并且σ = {σ1,σ2,..,〇 N}氨基酸序列。良好的折叠序列快速并以稳定的方式折叠成所需的靶结构。即使在蛋白质折叠的极简模型中,对满足这些标准的序列的蛮力搜索也是非常耗时的。虽然可以将这种类型的标准应用于简单的螺旋线圈模型[1],但必须找到更有效的策略。在[2]中提出了一个相当激烈的简化,其中通过最小化关于σ的E来解决问题,其中r固定在目标结构r 0上。这种方法非常快,因为不涉及构象空间的探索,但不幸的是,它在许多例子中失败了(参见例如[1,3,4])。最近,提出了一种更通用的方案[3],其目的是优化条件概率P(r 0| σ),即玻尔兹曼权重,而不是E(r 0,σ)。使用统计学语言,这对应于贝叶斯回归,而不是将模型拟合到数据点(期望的结构),概率分布的宽度由温度T给出。该方法的优点是考虑了熵效应,但由于最大化P(r 0| σ)是一个不平凡的任务。事实上,[3]中的计算涉及简化关于P(r 0)的形式的假设|σ)和构象空间。在这封信中,我们提出了一个实用的蒙特卡罗(MC)程序执行最大化的P(r 0| σ)。好的折叠器的热力学特征是基态最小值与其他状态很好地分离-在有限的T下,系统在基态很长时间。在晶格模型中,对于相对较小的链,状态是可重性的,这通常被认为是基态的非简并性,并且后者与更高的能量状态很好地分离。人们期望在匹配过程中使用有限T分布来挑选出那些在非简并性方面具有良好折叠特性的最佳序列。事实上,在[3]中,当使用3× 3× 3晶格在27-mer链上探索该技术时,发现与[2]中获得的结果相比上级的结果,其中E(r 0,σ)被最小化,相对于所发现的最佳序列的非简并性。|σ)是极其乏味的。在这里,我们设计了一种有效的MC方法,它是基于多序列方法[5],其中序列和坐标自由度都受到随机移动。基本思想是执行联合概率分布P(r,σ)的单次模拟,而不是玻尔兹曼分布P(r)的重复模拟|σ),对于不同的固定σ。因此,我们的方法与[4]的方法根本不同。该方法丢弃具有与目标结构具有低结构相似性的低位能量最小值的所有序列。因此,通过构建确保了所设计序列的能隙的存在。我们开发的方法最大化P(r| σ)在二维HP晶格模型[6]上使用长度N = 16、18和32的链进行探索。对于N = 16,我们研究了[4]中使用的例子,以将他们的方法与[2,3]的方法进行比较。N = 16和18的结果都是根据精确枚举进行检查的,而对于N = 32,我们使用“手工”构建的目标结构。只要比较可行,我们的方法就能非常迅速地重现精确的结果。此外,该方法具有相当大的潜力,以处理非常大的链。
A new method for sequence optimization in protein models is presented. The approach, which has inherited its basic philosophy from recent work by Deutsch and Kurosky [Phys. Rev. Lett. 76, 323 (1996)] by maximizing conditional probabilities rather than minimizing energy functions, is based upon a novel and very efficient multisequence Monte Carlo scheme. By construction, the method ensures that the designed sequences represent good folders thermodynamically. The algorithm, which is successfully explored on the two-dimensional HP-model with chain lengths N = 16, 18 and 32, can easily be generalized to off-lattice models. Also, a bootstrap procedure for the sequence space search is devised, which makes very large chains feasible. PACS numbers: 87.15.By, 87.10.+e irback@thep.lu.se carsten@thep.lu.se frank@thep.lu.se erik@thep.lu.se The “inverse” of protein folding, sequence optimization, is of utmost relevance in the context of drug design. This problem, which amounts to finding optimal amino acid sequences given a target structure, has also been investigated in the context of understanding folding properties of coarsegrained models for protein folding. Such models are described by energy functions E(r, σ), where r = {r1, r2, .., rN} denotes the amino acid coordinates and σ = {σ1, σ2, .., σN} the amino acid sequence. Good folding sequences fold fast and in a stable way into the desired target structure. A brute force search for sequences meeting these criteria is prohibitively time-consuming even in minimalist models for protein folding. Although it has been possible to apply this type of criteria to a simple helix-coil model [1], it is essential to find more efficient strategies. A fairly drastic simplification was proposed in [2], where the problem is approached by minimizing E with respect to σ with r clamped to the target structure, r0. This method is very fast since no exploration of the conformational space is involved, but, unfortunately, it fails for a number of examples (see e.g. [1, 3, 4]). Recently, a more generic scheme was suggested [3], which aims at optimizing the conditional probability P (r0|σ), i.e. the Boltzmann weight, rather than E(r0, σ). Using statistics language this corresponds to Bayesian regression rather than fitting a model to data points (desired structure), with the width of the probability distribution given by the temperature T . This approach has the advantage that entropy effects are taken into account, but its usefulness is not obvious since maximizing P (r0|σ) is a non-trivial task. In fact, the calculations in [3] involved simplifying assumptions about both the form of P (r0|σ) and the conformational space. In this letter we present a practical Monte Carlo (MC) procedure for performing the maximization of P (r0|σ). Thermodynamical characteristics for good folders are that the ground state minima are well separated from other states — at finite T the system spends a long time in the ground state well. In lattice models, where for relatively small chains the states are enumerable, this is often taken as non-degeneracy of ground state and that the latter is well separated from higher energy states. One expects that working with finite T distributions in the matching process singles out those optimal sequences that have good folding properties in terms of non-degeneracy. Indeed, in [3] when exploring the technique on a 27-mer chain using a 3× 3× 3 lattice, one finds superior results as compared to what was obtained in [2], where E(r0, σ) was minimized, with respect to non-degeneracy of the optimal sequences found. Computationally, straightforward MC approaches for maximizing P (r0|σ) are extremely tedious. Here we devise an efficient MC methodology, which is based on the Multisequence Method [5], where both sequence and coordinate degrees of freedom are subject to stochastic moves. The basic idea is to perform a single simulation of a joint probability distribution P (r, σ) rather than repeated simulations of the Boltzmann distribution P (r|σ) for different fixed σ. Hence, our approach is fundamentally different from that of [4]. The method discards all sequences that have low-lying energy minima with low structural resemblance to the target structure. Thus, the existence of an energy gap for the designed sequences is ensured by construction. The method we develop for maximizing P (r|σ) is explored on the two-dimensional HP lattice model [6] using chains of length N = 16, 18 and 32. For N = 16 we study the example used in [4] to compare their method to those of [2, 3]. The results for both N = 16 and 18 are checked against exact enumerations, whereas for N = 32 we use a target structure constructed “by hand”. Our method reproduces the exact results extremely rapidly whenever comparisons are feasible. Furthermore, the method has quite some potential to deal with very large chains.