HIDDEN MARKOV-MODELS IN COMPUTATIONAL BIOLOGY - APPLICATIONS TO PROTEIN MODELING

HIDDEN MARKOV-MODELS IN COMPUTATIONAL BIOLOGY - APPLICATIONS TO PROTEIN MODELING
复制标题

DOI:
10.1006/jmbi.1994.1104
复制
发表时间:
1994-02-04
影响因子:
5.6
通讯作者:
HAUSSLER, D
HAUSSLER, D
中科院分区:
生物学2区
文献类型:
--
作者:
KROGH, A;BROWN, M;HAUSSLER, D

文献摘要

被引文献

相似文献

隐马尔可夫模型(HiddenMarkovModels,简称HRM)被应用于蛋白质家族和蛋白质结构域的统计建模、数据库搜索和多序列比对等问题。这些方法被证明对珠蛋白家族,蛋白激酶催化结构域,EF-手钙结合基序。在每种情况下,HMM的参数都是从未对齐序列的训练集估计的。在HMM被建立之后,它被用于获得所有训练序列的多重比对。它还用于搜索SWISS-PROT 22数据库中作为给定蛋白质家族成员或包含给定结构域的其他序列。隐马尔可夫模型产生多个高质量的比对,这些比对与结合三维结构信息的程序产生的比对密切一致。当用于区分测试时(通过检查数据库中的序列与珠蛋白、激酶和EF-手障碍物的匹配程度),HMM能够以高度的准确度区分这些家族的成员和非成员。在这些测试中,HMM和PROFIELLESTRA(一种用于搜索蛋白质序列和多重比对序列之间关系的技术)都比PROSITE(蛋白质中位点和模式的字典)表现得更好。HMM似乎在较低的假阴性和假阳性率方面比PROFIELESTIC具有轻微的优势,即使HMM仅使用未对齐的序列进行训练,而PROFIELESTIC需要对齐的训练序列。我们的研究结果表明,在L-型钙通道α-1亚基的一个高度保守和进化保留的155个残基的推定细胞内区域中存在EF-手形钙结合基序,其在兴奋-收缩偶联中起重要作用。该区域已被认为含有所有L型钙通道的典型或必需的功能结构域,无论它们是否与兰尼碱受体、传导离子或两者偶联。
Hidden Markov Models (HMMs) are applied to the problems of statistical modeling, database searching and multiple sequence alignment of protein families and protein domains. These methods are demonstrated on the globin family, the protein kinase catalytic domain, and the EF-hand calcium binding motif. In each case the parameters of an HMM are estimated from a training set of unaligned sequences. After the HMM is built, it is used to obtain a multiple alignment of all the training sequences. It is also used to search the SWISS-PROT 22 database for other sequences that are members of the given protein family, or contain the given domain. The HMM produces multiple alignments of good quality that agree closely with the alignments produced by programs that incorporate three-dimensional structural information. When employed in discrimination tests (by examining how closely the sequences in a database fit the globin, kinase and EF-hand HMMs), the HMM is able to distinguish members of these families from non-members with a high degree of accuracy. Both the HMM and PROFILESEARCH (a technique used to search for relationships between a protein sequence and multiply aligned sequences) perform better in these tests than PROSITE (a dictionary of sites and patterns in proteins). The HMM appears to have a slight advantage over PROFILESEARCH in terms of lower rates of false negatives and false positives, even though the HMM is trained using only unaligned sequences, whereas PROFILESEARCH requires aligned training sequences. Our results suggest the presence of an EF-hand calcium binding motif in a highly conserved and evolutionary preserved putative intracellular region of 155 residues in the α-1 subunit of L-type calcium channels which play an important role in excitation-contraction coupling. This region has been suggested to contain the functional domains that are typical or essential for all L-type calcium channels regardless of whether they couple to ryanodine receptors, conduct ions or both.