Assignment of homology to genome sequences using a library of hidden Markov models that represent all proteins of known structure

Assignment of homology to genome sequences using a library of hidden Markov models that represent all proteins of known structure
复制标题

DOI:
10.1006/jmbi.2001.5080
复制
发表时间:
2001-11-02
影响因子:
5.6
通讯作者:
Chothia, C
Chothia, C
中科院分区:
生物学2区
文献类型:
--
作者:
Gough, J;Karplus, K;Chothia, C

文献摘要

被引文献

相似文献

在序列比较方法中,基于轮廓的方法比使用成对比较的方法具有更高的选择性。在分析方法中,隐马尔可夫模型(HMM)显然是最好的。本文的第一部分描述了(I)改进隐马尔可夫模型的性能和(Ii)确定为已知结构的蛋白质序列创建隐马尔可夫模型的良好过程的计算。对于相关蛋白质家族,有更多的同源基因。使用从不同的单一种子序列建立的多个模型来检测,而不是从那些序列的良好比对建立的一个模型来检测。描述了一种新的程序,用于检测和纠正在该程序的建模阶段出现的那些错误。这两个改进大大提高了选择性和覆盖率。论文的第二部分描述了一个被称为超家族的HMM文库的构建,该文库基本上代表了所有已知结构的蛋白质。已知结构的蛋白质中结构域序列的识别率低于95%,被用作构建模型的种子。利用现有的数据,得到一个包含4894个模型的文库。论文的第三部分描述了使用超家族模型库来注释50多个基因组的序列。该模型匹配的靶序列是通过成对序列比较方法匹配的靶序列的两倍。对于每个基因组,近一半的序列全部或部分匹配,总体上,匹配覆盖了35%的真核基因组和45%的细菌基因组。平均而言,大约15%的基因组序列被标记为假想的,但与已知结构的蛋白质同源。从这些匹配中派生的注解可从公共Web服务器获得:http://stash.mrc-lmb.cam.ac.uk/SUPERFAMILY.该服务器还使用户能够将他们自己的序列与超家族模型库进行匹配。(C)2001年学术出版社。
Of the sequence comparison methods, profile-based methods perform with greater selectively than those that use pairwise comparisons. Of the profile methods, hidden Markov models (HMMs) are apparently the best. The first part of this paper describes calculations that (i) improve the performance of HMMs and (ii) determine a good procedure for creating HMMs for sequences of proteins of known structure. For a family of related proteins, more homologues. are detected using multiple models built from diverse single seed sequences than from one model built from a good alignment of those sequences. A new procedure is described for detecting and correcting those errors that arise at the model-building stage of the procedure. These two improvements greatly increase selectivity and coverage.The second part of the paper describes the construction of a library of HMMs, called SUPERFAMILY, that represent essentially all proteins of known structure. The sequences of the domains in proteins of known structure, that have identifies less than 95%, are used as seeds to build the models. Using the current data, this gives a library with 4894 models.The third part of the paper describes the use of the SUPERFAMILY model library to annotate the sequences of over 50 genomes. The models match twice as many target sequences as are matched by pairwise sequence comparison methods. For each genome, close to half of the sequences are matched in all or in part and, overall, the matches cover 35% of eukaryotic genomes and 45% of bacterial genomes. On average roughly 15% of genome sequences are labelled as being hypothetical yet homologous to proteins of known structure. The annotations derived from these matches are available from a public web server at: http://stash.mrc-lmb.cam.ac.uk/SUPERFAMILY. This server also enables users to match their own sequences against the SUPERFAMILY model library. (C) 2001 Academic Press.