Hidden Markov models for detecting remote protein homologies

Hidden Markov models for detecting remote protein homologies
复制标题

DOI:
10.1093/bioinformatics/14.10.846
复制
发表时间:
1998-01-01
期刊:
影响因子:
5.8
通讯作者:
Hughey, R
Hughey, R
中科院分区:
生物学3区
文献类型:
--
作者:
Karplus, K;Barrett, C;Hughey, R

文献摘要

被引文献

相似文献

动机:描述和评估了一种用于寻找蛋白质序列远程同源物的新隐马尔可夫模型方法(SAM-T98)。该方法从简单的目标序列开始,并根据序列和使用 HMM 进行数据库搜索找到的同源物迭代构建隐马尔可夫模型 (HMM)。 SAM-T98 还用于根据结构数据库中的序列自动构建模型库。方法:我们使用错误的数据集评估 SAM-T98 方法。其中三个测试集是折叠识别测试,其中正确答案由结构相似性确定。第四个使用精选数据库。该方法与 WU-BLASTP 和 DOUBLE-BLAST 进行比较,DOUBLE-BLAST 是一种类似于 ISS 的两步方法,但使用 BLAST 而不是 FASTA。结果:SAM-T98 在所有测试中的错误最少,对于折叠识别测试而言也是如此。在 SCOP(蛋白质结构分类)域测试的最小误差点,SAM-T98 获得 880 个流感阳性和 68 个假阳性,DOUBLE-BLAST 获得 533 个真阳性和 71 个假阳性,而 WU-BLASTP 获得 353 个真阳性和 24 个假阳性。该方法经过优化以识别超家族,并且需要使用参数调整来查找家族或折叠关系。HMM 方法性能的一个关键是一种新的分数归一化技术,该技术将分数与反向模型的分数进行比较,而不是与统一的零模型进行比较。
Motivation: A new hidden Markov model method (SAM-T98) for finding remote homologs of protein sequences is described and evaluated. The method begins with a simple target sequence and iteratively builds a hidden Markov model (HMM) from the sequence and homologs found using die HMM for database search. SAM-T98 is also used to construct model libraries automatically, from sequences in structural databases.Methods: We evaluate the SAM-T98 method with foul datasets. Three of the test sets are fold-recognition tests, where the correct answers are determined by structural similarity. The fourth uses a curated database. The method is compared against WU-BLASTP and against DOUBLE-BLAST, a two-step method similar to ISS, but using BLAST instead of FASTA.Results: SAM-T98 had the fewest errors in all tests- dramatically so for the fold-recognition tests. At the minimum-error point on the SCOP (Structural Classification of Proteins)-domains test, SAM-T98 got 880 flue positives and 68 false positives, DOUBLE-BLAST got 533 true positives with 71 false positives, ann WU-BLASTP got 353 true positives with 24 false positives. The method is optimized to recognize superfamilies, and would require parameter adjustment to be used to find family or fold relationships, One key to the performance of the HMM method is a new score-normalization technique that compares the score to the score with a reversed model rather than to a uniform null model.