Hidden Markov models that use predicted local structure for fold recognition: Alphabets of backbone geometry

Hidden Markov models that use predicted local structure for fold recognition: Alphabets of backbone geometry
复制标题

DOI:
10.1002/prot.10369
复制
发表时间:
2003-06-01
影响因子:
2.9
通讯作者:
Karplus, K
Karplus, K
中科院分区:
生物学4区
文献类型:
--
作者:
Karchin, R;Cline, M;Karplus, K

文献摘要

被引文献

相似文献

计算生物学中的一个重要问题是预测基因组测序项目发现的大量假定蛋白质的结构。折叠识别方法试图通过将目标蛋白与已知结构联系起来,寻找与目标蛋白同源的模板蛋白来解决这个问题。可能具有显著结构相似性的远端同系物通常不能仅通过序列相似性来检测。为了解决这一问题,我们将预测局部结构(二级结构的推广)引入到双道轮廓隐马尔可夫模型(HMM)中。我们没有依赖于简单的二级结构的螺旋-链-线圈定义,而是试验了各种局部结构描述,遵循一个原则性的协议来确定哪些描述对于提高折叠识别和比对质量最有用。在一组1298个非同源蛋白质的测试集上,MMS结合了3个字母的跨度字母表,通过ROC-65数字衡量,MMS比仅含氨基酸的HMM提高了15%的折叠识别准确率,比PSI-BLAST提高了23%。我们在200个蛋白质对的困难比对测试集上比较了双道MMS和仅氨基酸MMS(结构相似,序列同源性为3%-24%)。与DALI结构对齐相比,具有6个字母跨距二级轨道的MMS将对齐质量提高了62%,而具有STR轨道(将链细分为六种状态的扩展的DSSP字母表)的MMS相对于CE提高了40%。
An important problem in computational biology is predicting the structure of the large number of putative proteins discovered by genome sequencing projects. Fold-recognition methods attempt to solve the problem by relating the target proteins to known structures, searching for template proteins homologous to the target. Remote homologs that may have significant structural similarity are often not detectable by sequence similarities alone. To address this, we incorporated predicted local structure, a generalization of secondary structure, into two-track profile hidden Markov models (Hmms). We did not rely on a simple helix-strand-coil definition of secondary structure, but experimented with a variety of local structure descriptions, following a principled protocol to establish which descriptions are most useful for improving fold recognition and alignment quality. On a test set of 1298 nonhomologous proteins, mms incorporating a 3-letter STRIDE alphabet improved fold recognition accuracy by 15% over amino-acid-only Hmms and 23% over PSI-BLAST, measured by ROC-65 numbers. We compared two-track mms to amino-acid-only mms on a difficult alignment test set of 200 protein pairs (structurally similar with 3-24% sequence identity). mms with a 6-letter STRIDE secondary track improved alignment quality by 62%, relative to DALI structural alignments, while mms with an STR track (an expanded DSSP alphabet that subdivides strands into six states) improved by 40% relative to CE.