Predicting protein secondary structure using stochastic tree grammars

Predicting protein secondary structure using stochastic tree grammars
复制标题

DOI:
10.1023/a:1007477814995
复制
发表时间:
1997-11-01
期刊:
影响因子:
7.5
通讯作者:
Mamitsuka, H
Mamitsuka, H
中科院分区:
计算机科学3区
文献类型:
--
作者:
Abe, N;Mamitsuka, H

文献摘要

被引文献

相似文献

我们提出了一种新的方法来预测蛋白质二级结构的一个给定的氨基酸序列,基于随机树文法的概率参数的训练算法。特别是,我们专注于预测β-折叠区域的问题,这在以前被认为是困难的,因为对应于β-折叠的序列所表现出的无界依赖性。为了科普这一困难,我们使用一个新的家庭的随机树文法,我们称之为随机排名节点重写文法,这是足够强大的捕获类型的依赖表现出的序列的P-片区域,如“平行”和“反平行”的依赖关系和它们的组合。我们使用的训练算法是随机上下文蜜蜂语法的“内外”算法的扩展,但有一些显着的修改。我们将我们的方法应用于从HSSP数据库获得的真实的数据!结果令人鼓舞:我们的方法能够在系统评估实验中正确预测大约75%的β链,其中测试序列不仅与训练序列的同一性不到25%,而且与它们完全无关。这个数字与该领域最先进的预测方法的预测准确性相比是有利的,即使我们的实验是在有限类型的β折叠结构上进行的,并且测试是在相对较小的数据大小上进行的。我们还强调,我们的方法可以预测的结构,以及β-折叠区域的位置,这是不可能通过传统的方法进行二级结构预测。本文所介绍的部分工作的扩展摘要已出现在(安倍和Mamitsuka,1994年)和(Mamitsuka和安倍,1993年)。
We propose a new method for predicting protein secondary structure of a given amino acid sequence, based on a training algorithm for the probability parameters of a stochastic tree grammar. In particular, we concentrate on the problem of predicting beta-sheet regions, which has previously been considered difficult because of the unbounded dependencies exhibited by sequences corresponding to beta-sheets. To cope with this difficulty, we use a new family of stochastic tree grammars, which we call Stochastic Ranked Node Rewriting Grammars, which are powerful enough to capture the type of dependencies exhibited by the sequences of P-sheet regions, such as the 'parallel' and 'anti-parallel' dependencies and their combinations. The training algorithm we use is an extension of the 'inside-outside' algorithm for stochastic context-bee grammars, but with a number of significant modifications. We applied our method on real data obtained from the HSSP database !Homology-derived Secondary Structure of Proteins Ver 1.0) and the results were encouraging: Our method was able to predict roughly 75 percent of the beta-strands correctly in a systematic evaluation experiment, in which the test sequences not only have less than 25 percent identity to the training sequences, but are totally unrelated to them. This figure compares favorably to the predictive accuracy of the state-of-the-art prediction methods in the field, even though our experiment was on a restricted type of beta-sheet structures and the test was done on a relatively small data size. We also stress that our method can predict the structure as well as the location of beta-sheet regions, which was not possible by conventional methods for secondary structure prediction. Extended abstracts of parts of the work presented in this paper have appeared in (Abe & Mamitsuka, 1994) and (Mamitsuka & Abe, 1993).