Improving model construction of profile HMMs for remote homology detection through structural alignment.

Improving model construction of profile HMMs for remote homology detection through structural alignment.
复制标题

DOI:
10.1186/1471-2105-8-435
复制
发表时间:
2007-11-09
期刊:
影响因子:
3
通讯作者:
Zaverucha G
Zaverucha G
中科院分区:
生物学4区
文献类型:
--
作者:
Bernardes JS;Dávila AM;Costa VS;Zaverucha G

文献摘要

参考文献

被引文献

相似文献

远程同源性检测是生物信息学中的一个具有挑战性的问题。可以论证的是,概要隐马尔可夫模型(phmm)是解决这一重要问题的最成功的方法之一。pHMM包的计算成本相对较小,并且在识别远程同源性方面表现得特别好。这就提出了一个问题,即结构比对是否会影响从模糊区蛋白质中培养的phmm的性能,因为在识别基序和功能残基方面,结构比对通常比序列比对更准确。接下来,我们评估使用结构对齐对pHMM性能的影响。我们使用SCOP数据库进行实验。使用3DCOFFEE和MAMMOTH-mult工具进行结构比对;利用CLUSTALW、TCOFFEE、MAFFT和PROBCONS进行序列比对。我们对超级家庭进行了留一个家庭的交叉验证。采用ROC曲线和配对双尾t检验评价疗效。我们观察到,在低同一性区域(主要低于20%),结构比对得到的phmm比序列比对得到的phmm表现明显更好。我们认为这是因为结构对齐工具更擅长于关注通过进化而更经常保存的重要模式,从而产生更高质量的phmm。另一方面,这些工具对这些低身份区域的敏感性仍然很低。我们的研究结果为这一领域的改进提出了一些可能的方向。
Remote homology detection is a challenging problem in Bioinformatics. Arguably, profile Hidden Markov Models (pHMMs) are one of the most successful approaches in addressing this important problem. pHMM packages present a relatively small computational cost, and perform particularly well at recognizing remote homologies. This raises the question of whether structural alignments could impact the performance of pHMMs trained from proteins in the Twilight Zone, as structural alignments are often more accurate than sequence alignments at identifying motifs and functional residues. Next, we assess the impact of using structural alignments in pHMM performance. We used the SCOP database to perform our experiments. Structural alignments were obtained using the 3DCOFFEE and MAMMOTH-mult tools; sequence alignments were obtained using CLUSTALW, TCOFFEE, MAFFT and PROBCONS. We performed leave-one-family-out cross-validation over super-families. Performance was evaluated through ROC curves and paired two tailed t-test. We observed that pHMMs derived from structural alignments performed significantly better than pHMMs derived from sequence alignment in low-identity regions, mainly below 20%. We believe this is because structural alignment tools are better at focusing on the important patterns that are more often conserved through evolution, resulting in higher quality pHMMs. On the other hand, sensitivity of these tools is still quite low for these low-identity regions. Our results suggest a number of possible directions for improvements in this area.
DOI: 10.1093/bioinformatics/bth091
发表时间: 2004-05-22
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Edgar, RC;Sjölander, K
通讯作者: Sjölander, K
DOI: 10.1093/nar/gkh039
发表时间: 2004-01-01
影响因子: 14.9
作者:
Andreeva, A;Howorth, D;Murzin, AG
通讯作者: Murzin, AG
DOI: 10.1006/jmbi.2001.5080
发表时间: 2001-11-02
影响因子: 5.6
作者:
Gough, J;Karplus, K;Chothia, C
通讯作者: Chothia, C
DOI: 10.1002/prot.10369
发表时间: 2003-06-01
影响因子: 2.9
作者:
Karchin, R;Cline, M;Karplus, K
通讯作者: Karplus, K
DOI: 10.1073/pnas.84.13.4355
发表时间: 1987-07-01
影响因子: 11.1
作者:
GRIBSKOV, M;MCLACHLAN, AD;EISENBERG, D
通讯作者: EISENBERG, D