DISCRIMINATION OF INTRACELLULAR AND EXTRACELLULAR PROTEINS USING AMINO-ACID-COMPOSITION AND RESIDUE-PAIR FREQUENCIES

DISCRIMINATION OF INTRACELLULAR AND EXTRACELLULAR PROTEINS USING AMINO-ACID-COMPOSITION AND RESIDUE-PAIR FREQUENCIES
复制标题

DOI:
10.1006/jmbi.1994.1267
复制
发表时间:
1994-04-22
影响因子:
5.6
通讯作者:
NISHIKAWA, K
NISHIKAWA, K
中科院分区:
生物学2区
文献类型:
--
作者:
NAKASHIMA, H;NISHIKAWA, K

文献摘要

被引文献

相似文献

细胞内和细胞外可溶性蛋白质序列的氨基酸组成和残基对频率进行了统计分析。计算从(n,n+ 1)到(n,n+ 5)连续分离的残基对频率,并转换为评分参数。然后,对于每种测试蛋白质,应用单残基和残基对参数来计算总分。根据我们的定义,产生阳性评分的蛋白质指示细胞内蛋白质,而阴性评分意味着细胞外蛋白质。参数集来自PIR数据库中构成不同蛋白质家族的894个序列,并通过应用于379种蛋白质的测试进行评估。结果表明,88%的胞内蛋白和84%的胞外蛋白被正确分配。与之前仅使用成分数据的研究相比,辨别力提高了约8%。还通过其他标准观察到细胞内/细胞外蛋白质的分离,例如结构类别(细胞内蛋白质偏好α和α/β型,细胞外蛋白质偏好β和α + β型)。分离序列被认为是一个更可靠的程序区分内/胞外蛋白质比使用结构类的方法。这种分离序列的可能原因进行了讨论。
Sequences of intracellular and extracellular soluble proteins were analyzed statistically in terms of amino acid composition and residue-pair frequencies. Residue-pair frequencies were calculated for sequential separations from (n, n+ 1) to (n, n+ 5), and converted into scoring parameters. Then, for each test protein, the single-residue and residue-pair parameters were applied to calculate a total score. According to our definition, a protein which yields a positive score is indicative of an intracellular protein, whereas a negative score implies an extracellular one. The parameter set was derived from 894 sequences constituting different protein families in the PIR database, and assessed by application to a test of 379 proteins. The results showed that 88% of intracellular and 84% of extracellular proteins were correctly assigned. The discrimination power was improved by about 8% in comparison with the previous study, which used composition data alone. Segregation of intra/extracellular proteins is also observed by other criteria, such as structural class (intracellular proteins prefer α and α/β types and extracellular proteins prefer β and α + β types). Segregation by sequence was found to be a more reliable procedure for distinguishing intra/extracellular protein than methods using structural class. Possible causes for this segregation by sequence are discussed.