Bayesian Top-Down Protein Sequence Alignment with Inferred Position-Specific Gap Penalties.

Bayesian Top-Down Protein Sequence Alignment with Inferred Position-Specific Gap Penalties.
复制标题

DOI:
10.1371/journal.pcbi.1004936
复制
发表时间:
2016-05
影响因子:
4.3
通讯作者:
Altschul SF
Altschul SF
中科院分区:
生物学2区
文献类型:
--
作者:
Neuwald AF;Altschul SF

文献摘要

被引文献

相似文献

我们描述了一个贝叶斯马尔可夫链蒙特卡罗(MCMC)采样器的蛋白质多序列比对(MSA),在程序GISMO和应用于大量的不同的序列,是更准确的比流行的MSA程序MUSCLE,MAFFT,Clustal-Ω和Kalign。GISMO的主要特点是:(1)它采用了一种“自上而下”的策略,具有良好的渐近时间复杂度,首先识别出所有输入序列通常共享的区域,然后重新排列密切相关的子组串联。(ii)它推断位置特异性空位罚分,其有利于每个序列内在比对位置处的插入或缺失(indel),其中在其他序列中调用indel。这有利于在保守区块之间插入,这可以被理解为构成蛋白质的结构核心。(iii)它使用基于最小描述长度原则和狄利克雷混合先验的贝叶斯统计比对质量度量。因此,GISMO仅在统计上合理时才比对序列区域。这与基于特别但广泛使用的配对和评分系统的方法不同,该系统将对齐随机序列。(iv)它定义了一个系统,用于探索对齐空间,通过开发新的采样策略,更有效地逃离次优陷阱,为进一步的实验提供了自然的途径。GISMO的上级性能说明使用408蛋白质组,平均包含235个序列。这些集合对应于NCBI保守结构域数据库比对,其已经根据可用的晶体结构手动地策划,并且因此提供了评估比对准确度的手段。GISMO填补了与其他MSA程序不同的利基,即识别和比对存在于大的、不同的全长序列组中的保守结构域。GISMO程序可在http://gismo.igs.umaryland.edu/上获得。现有的多重比对程序通常利用(i)自下而上的渐进策略,这需要对每对输入序列进行耗时的比对,(ii)比对质量的临时测量,以及(iii)预先指定的、统一定义的缺口罚分。在这里,我们描述了一种替代策略,首先临时对齐区域通常共享的所有输入序列,然后通过迭代重新排列串联相关序列来完善这种对齐。它直接从进化的比对中推断出位置特异性空位罚分。它避免了次优陷阱随机遍历复杂的,相关的空间对齐使用统计上严格的措施对齐质量。对于大的序列集,这种方法提供了明显的优势,在比对准确性超过目前最流行的程序。
We describe a Bayesian Markov chain Monte Carlo (MCMC) sampler for protein multiple sequence alignment (MSA) that, as implemented in the program GISMO and applied to large numbers of diverse sequences, is more accurate than the popular MSA programs MUSCLE, MAFFT, Clustal-Ω and Kalign. Features of GISMO central to its performance are: (i) It employs a “top-down” strategy with a favorable asymptotic time complexity that first identifies regions generally shared by all the input sequences, and then realigns closely related subgroups in tandem. (ii) It infers position-specific gap penalties that favor insertions or deletions (indels) within each sequence at alignment positions in which indels are invoked in other sequences. This favors the placement of insertions between conserved blocks, which can be understood as making up the proteins’ structural core. (iii) It uses a Bayesian statistical measure of alignment quality based on the minimum description length principle and on Dirichlet mixture priors. Consequently, GISMO aligns sequence regions only when statistically justified. This is unlike methods based on the ad hoc, but widely used, sum-of-the-pairs scoring system, which will align random sequences. (iv) It defines a system for exploring alignment space that provides natural avenues for further experimentation through the development of new sampling strategies for more efficiently escaping from suboptimal traps. GISMO’s superior performance is illustrated using 408 protein sets containing, on average, 235 sequences. These sets correspond to NCBI Conserved Domain Database alignments, which have been manually curated in the light of available crystal structures, and thus provide a means to assess alignment accuracy. GISMO fills a different niche than other MSA programs, namely identifying and aligning a conserved domain present within a large, diverse set of full length sequences. The GISMO program is available at http://gismo.igs.umaryland.edu/. Existing multiple alignment programs typically utilize (i) bottom-up progressive strategies, which require the time-consuming alignment of each pair of input sequences, (ii) ad hoc measures of alignment quality, and (iii) pre-specified, uniformly-defined gap penalties. Here we describe an alternative strategy that first provisionally aligns regions generally shared by all the input sequences, and then refines this alignment by iteratively realigning correlated sequences in tandem. It infers position-specific gap penalties directly from the evolving alignment. It avoids suboptimal traps by stochastically traversing the complex, correlated space of alignments using a statistically rigorous measure of alignment quality. For large sequence sets, this approach offers clear advantages in alignment accuracy over the most popular programs currently available.