A Novel Monte Carlo Procedure for Protein Design
A Novel Monte Carlo Procedure for Protein Design
复制标题
DOI:
--
复制
发表时间:
1997-11
期刊:
影响因子:
--
通讯作者:
Anders Irbäck;Carsten Peterson;F. Potthast;Erik Sandelin
中科院分区:
文献类型:
--
作者:
Anders Irbäck;Carsten Peterson;F. Potthast;Erik Sandelin
A new method for sequence optimization in protein models is presented. The approach, which has inherited its basic philosophy from recent work by Deutsch and Kurosky [Phys. Rev. Lett. 76, 323 (1996)] by maximizing conditional probabilities rather than minimizing energy functions, is based upon a novel and very efficient multisequence Monte Carlo scheme. By construction, the method ensures that the designed sequences represent good folders thermodynamically. The algorithm, which is successfully explored on the two-dimensional HP-model with chain lengths N = 16, 18 and 32, can easily be generalized to off-lattice models. Also, a bootstrap procedure for the sequence space search is devised, which makes very large chains feasible. PACS numbers: 87.15.By, 87.10.+e irback@thep.lu.se carsten@thep.lu.se frank@thep.lu.se erik@thep.lu.se The “inverse” of protein folding, sequence optimization, is of utmost relevance in the context of drug design. This problem, which amounts to finding optimal amino acid sequences given a target structure, has also been investigated in the context of understanding folding properties of coarsegrained models for protein folding. Such models are described by energy functions E(r, σ), where r = {r1, r2, .., rN} denotes the amino acid coordinates and σ = {σ1, σ2, .., σN} the amino acid sequence. Good folding sequences fold fast and in a stable way into the desired target structure. A brute force search for sequences meeting these criteria is prohibitively time-consuming even in minimalist models for protein folding. Although it has been possible to apply this type of criteria to a simple helix-coil model [1], it is essential to find more efficient strategies. A fairly drastic simplification was proposed in [2], where the problem is approached by minimizing E with respect to σ with r clamped to the target structure, r0. This method is very fast since no exploration of the conformational space is involved, but, unfortunately, it fails for a number of examples (see e.g. [1, 3, 4]). Recently, a more generic scheme was suggested [3], which aims at optimizing the conditional probability P (r0|σ), i.e. the Boltzmann weight, rather than E(r0, σ). Using statistics language this corresponds to Bayesian regression rather than fitting a model to data points (desired structure), with the width of the probability distribution given by the temperature T . This approach has the advantage that entropy effects are taken into account, but its usefulness is not obvious since maximizing P (r0|σ) is a non-trivial task. In fact, the calculations in [3] involved simplifying assumptions about both the form of P (r0|σ) and the conformational space. In this letter we present a practical Monte Carlo (MC) procedure for performing the maximization of P (r0|σ). Thermodynamical characteristics for good folders are that the ground state minima are well separated from other states — at finite T the system spends a long time in the ground state well. In lattice models, where for relatively small chains the states are enumerable, this is often taken as non-degeneracy of ground state and that the latter is well separated from higher energy states. One expects that working with finite T distributions in the matching process singles out those optimal sequences that have good folding properties in terms of non-degeneracy. Indeed, in [3] when exploring the technique on a 27-mer chain using a 3× 3× 3 lattice, one finds superior results as compared to what was obtained in [2], where E(r0, σ) was minimized, with respect to non-degeneracy of the optimal sequences found. Computationally, straightforward MC approaches for maximizing P (r0|σ) are extremely tedious. Here we devise an efficient MC methodology, which is based on the Multisequence Method [5], where both sequence and coordinate degrees of freedom are subject to stochastic moves. The basic idea is to perform a single simulation of a joint probability distribution P (r, σ) rather than repeated simulations of the Boltzmann distribution P (r|σ) for different fixed σ. Hence, our approach is fundamentally different from that of [4]. The method discards all sequences that have low-lying energy minima with low structural resemblance to the target structure. Thus, the existence of an energy gap for the designed sequences is ensured by construction. The method we develop for maximizing P (r|σ) is explored on the two-dimensional HP lattice model [6] using chains of length N = 16, 18 and 32. For N = 16 we study the example used in [4] to compare their method to those of [2, 3]. The results for both N = 16 and 18 are checked against exact enumerations, whereas for N = 32 we use a target structure constructed “by hand”. Our method reproduces the exact results extremely rapidly whenever comparisons are feasible. Furthermore, the method has quite some potential to deal with very large chains.