Using amphiphilic pseudo amino acid composition to predict enzyme subfamily classes

Using amphiphilic pseudo amino acid composition to predict enzyme subfamily classes
复制标题

DOI:
10.1093/bioinformatics/bth466
复制
发表时间:
2005-01-01
期刊:
影响因子:
5.8
通讯作者:
Chou, KC
Chou, KC
中科院分区:
生物学3区
文献类型:
--
作者:
Chou, KC

文献摘要

被引文献

相似文献

动机:随着蛋白质序列以爆炸性的速度进入数据库,早期确定新发现的酶分子的家族或亚家族类别变得重要,因为这直接关系到它作用于哪个特定靶点的详细信息,以及它的催化过程和生物功能。不幸的是,仅通过实验来做到这一点既耗时又昂贵。在以前的研究中,协变-判别算法被引入来识别氧化还原酶的16个亚家族类别。尽管结果相当令人鼓舞,但整个预测过程仅基于氨基酸组成,不包括任何序列顺序信息。结果:为了将序列顺序效应引入到预测因子中,引入了两亲性假氨基酸组成来表示蛋白质的统计样本。新的表示法包含20+2lambda离散数字:前20个数字是传统氨基酸组成的组成部分;接下来的2lambda数字是一组相关因子,它们反映了不同的疏水性和亲水性沿蛋白质链的分布模式。基于这一概念和公式,提出了一种新的预报器。自相合性检验、刀切检验和独立数据集检验表明,新预报器的预测成功率均显著高于以往的预报器。成功率的显著提高也意味着氨基酸残基的疏水性和亲水性在蛋白质链上的分布对其结构和功能起着非常重要的作用。
Motivation: With protein sequences entering into databanks at an explosive pace, the early determination of the family or subfamily class for a newly found enzyme molecule becomes important because this is directly related to the detailed information about which specific target it acts on, as well as to its catalytic process and biological function. Unfortunately, it is both time-consuming and costly to do so by experiments alone. In a previous study, the covariant-discriminant algorithm was introduced to identify the 16 subfamily classes of oxidoreductases. Although the results were quite encouraging, the entire prediction process was based on the amino acid composition alone without including any sequence-order information. Therefore, it is worthy of further investigation.Results: To incorporate the sequence-order effects into the predictor, the 'amphiphilic pseudo amino acid composition' is introduced to represent the statistical sample of a protein. The novel representation contains 20 + 2lambda discrete numbers: the first 20 numbers are the components of the conventional amino acid composition; the next 2lambda numbers are a set of correlation factors that reflect different hydrophobicity and hydrophilicity distribution patterns along a protein chain. Based on such a concept and formulation scheme, a new predictor is developed. It is shown by the self-consistency test, jackknife test and independent dataset tests that the success rates obtained by the new predictor are all significantly higher than those by the previous predictors. The significant enhancement in success rates also implies that the distribution of hydrophobicity and hydrophilicity of the amino acid residues along a protein chain plays a very important role to its structure and function.