Fold-specific sequence scoring improves protein sequence matching.

Fold-specific sequence scoring improves protein sequence matching.
复制标题

折叠特异性序列评分改善了蛋白质序列匹配。

DOI:
10.1186/s12859-016-1198-z
复制
发表时间:
2016-08-30
期刊:
影响因子:
3
通讯作者:
Jernigan RL
Jernigan RL
中科院分区:
生物学4区
文献类型:
--
作者:
Leelananda SP;Kloczkowski A;Jernigan RL

文献摘要

参考文献

被引文献

相似文献

序列匹配在整个生物学中的应用极其重要,特别是在发现功能和进化关系等信息方面,以及在区分无关紧要的突变和疾病突变方面。目前,很大一部分基因的功能尚不清楚;序列匹配的改进将改善基因注释。通用氨基酸替代矩阵,如Blosum 62,被用来测量序列相似性并识别远距离的同系物,而与结构类别无关。然而,这样的单一矩阵没有考虑到蛋白质不同拓扑结构中明显的重要结构信息,并以相同的方式对待所有蛋白质折叠中的取代。其他人提出,使用结构信息可以显著改善序列匹配,但这还不是很有效。在这里,我们开发了新的替换矩阵,它不仅包括一般的序列信息,而且还具有对于每个Cath拓扑唯一的拓扑特定分量。这种针对每个蛋白质拓扑使用序列和结构信息的组合的新特征显著提高了所测试的序列对的序列匹配分数。我们使用了一种新的多结构比对方法对CATH的每个同源水平进行比对,以提取拓扑信息。我们在73%的α螺旋测试用例中获得了统计上显著的改进的序列匹配分数。平均而言,当结构信息被纳入替换矩阵时,61%的测试用例在同源性检测方面表现出改进。对于所有病例,同源性检测的z得分平均提高了54%以上,一些个别病例的z得分是使用通用矩阵获得的两倍以上。我们的拓扑特定相似矩阵也优于其他传统的相似矩阵和基于单个矩阵的结构方法。将Psi-BLAST算法中默认的氨基酸替换矩阵替换为基于结构的矩阵,结构匹配比传统的Psi-BLAST算法有了显著的提高。它也比为每个拓扑生成的相应HMM简档获得的结果更好。我们发现,通过在特定的氨基酸替换矩阵中加入特定于拓扑的结构信息和序列信息,序列匹配分数和同源性检测都得到了显著提高。该方法在序列匹配方面优于传统的基于单矩阵的相似矩阵和基于Psi-BLAST和HMM Profile的方法。这些结果支持了新的氨基酸相似矩阵区分远距离同系物和结构不相似对的能力。本文的在线版本(doi:10.1186/s12859-0161198-z)包含补充材料,授权用户可以使用。
Sequence matching is extremely important for applications throughout biology, particularly for discovering information such as functional and evolutionary relationships, and also for discriminating between unimportant and disease mutants. At present the functions of a large fraction of genes are unknown; improvements in sequence matching will improve gene annotations. Universal amino acid substitution matrices such as Blosum62 are used to measure sequence similarities and to identify distant homologues, regardless of the structure class. However, such single matrices do not take into account important structural information evident within the different topologies of proteins and treats substitutions within all protein folds identically. Others have suggested that the use of structural information can lead to significant improvements in sequence matching but this has not yet been very effective. Here we develop novel substitution matrices that include not only general sequence information but also have a topology specific component that is unique for each CATH topology. This novel feature of using a combination of sequence and structure information for each protein topology significantly improves the sequence matching scores for the sequence pairs tested. We have used a novel multi-structure alignment method for each homology level of CATH in order to extract topological information. We obtain statistically significant improved sequence matching scores for 73 % of the alpha helical test cases. On average, 61 % of the test cases showed improvements in homology detection when structure information was incorporated into the substitution matrices. On average z-scores for homology detection are improved by more than 54 % for all cases, and some individual cases have z-scores more than twice those obtained using generic matrices. Our topology specific similarity matrices also outperform other traditional similarity matrices and single matrix based structure methods. When default amino acid substitution matrix in the Psi-blast algorithm is replaced by our structure-based matrices, the structure matching is significantly improved over conventional Psi-blast. It also outperforms results obtained for the corresponding HMM profiles generated for each topology. We show that by incorporating topology-specific structure information in addition to sequence information into specific amino acid substitution matrices, the sequence matching scores and homology detection are significantly improved. Our topology specific similarity matrices outperform other traditional similarity matrices, single matrix based structure methods, also show improvement over conventional Psi-blast and HMM profile based methods in sequence matching. The results support the discriminatory ability of the new amino acid similarity matrices to distinguish between distant homologs and structurally dissimilar pairs. The online version of this article (doi:10.1186/s12859-016-1198-z) contains supplementary material, which is available to authorized users.
DOI: 10.1002/prot.22458
发表时间: 2009-11-15
影响因子: 2.9
作者:
Illergard, Kristoffer;Ardell, David H.;Elofison, Arne
通讯作者: Elofison, Arne
DOI: 10.1073/pnas.89.22.10915
发表时间: 1992-11-15
影响因子: 11.1
作者:
HENIKOFF, S;HENIKOFF, JG
通讯作者: HENIKOFF, JG
DOI: 10.1186/1471-2105-8-435
发表时间: 2007-11-09
期刊: BMC bioinformatics
影响因子: 3
作者:
Bernardes JS;Dávila AM;Costa VS;Zaverucha G
通讯作者: Zaverucha G
DOI: 10.1101/gr.3866105
发表时间: 2005-12-01
期刊: GENOME RESEARCH
影响因子: 7
作者:
Brent, MR
通讯作者: Brent, MR
DOI: 10.1089/cmb.2011.0307
发表时间: 2012-07-01
影响因子: 1.7
作者:
Gniewek, Pawel;Kolinski, Andrzej;Gront, Dominik
通讯作者: Gront, Dominik