Mining α-helix-forming molecular recognition features with cross species sequence alignments

Mining α-helix-forming molecular recognition features with cross species sequence alignments
复制标题

DOI:
10.1021/bi7012273
复制
发表时间:
2007-11-27
期刊:
影响因子:
2.9
通讯作者:
Dunker, A. Keith
Dunker, A. Keith
中科院分区:
生物学3区
文献类型:
--
作者:
Cheng, Yugong;Oldfield, Christopher J.;Dunker, A. Keith

文献摘要

被引文献

相似文献

先前描述的挖掘算法(x-螺旋形成分子识别元件(MORE),由Oldfield等人描述。(Oldfield,C.J.,程,Y.,Cortese,M.S.,Brown,C.J.,Uversky,V.N.和Dunker,A.K.(2005)比较和组合主要无序蛋白质的预测因子,生物化学44,1989-2000),也称为分子识别特征(Morf)(Mohan,A.,Oldfield,C.J.,Radivojac,P.,Vacic,V.,Cortese,M.S.,Dunker,A.K.,and Uversky,V.N.),分子识别特征(Morf)分析,J.Mol.比奥尔。362,1043-1059),揭示了经历无序到有序转变的区域参与了许多分子识别事件,并且对蛋白质-蛋白质相互作用至关重要。然而,这些算法是使用有限大小的训练数据集开发的。在这里,我们建议通过以下方法来改进预测算法:(1)在正训练集中包括额外的(α-Morf)示例及其跨物种同源物,(2)从蛋白质数据库(PDB)中仔细提取单体结构链作为负训练集,(3)包括来自最近开发的无序预测器、二级结构预测器和氨基酸指数的属性,以及(4)构建基于神经网络的预测器并进行验证。在PDB中确定的50多个经历从无序到有序转变的区域以及每个基于结构的例子的一组对应的跨物种同源物被包括在新的正训练集中。通过条件概率方法评估了1500多个属性,包括无序预测、二级结构预测和氨基酸指数。使用VSL2和VL3无序预测和氨基酸残基的几种物理化学倾向等顶级属性来建立前馈神经网络。在10次交叉验证中,预测因子α-Morf-PredII的敏感性、特异性和准确性分别为0.87+/-0.10、0.87+/-0.11和0.87+/-0.08。我们给出了这些分析的结果和验证实例,以讨论(α-Morf-PredII预测精度的潜在改进。
Previously described algorithms for mining (x-helix-forming molecular recognition elements (MoREs), described by Oldfield et al. (Oldfield, C. J., Cheng, Y., Cortese, M. S., Brown, C. J., Uversky, V. N., and Dunker, A. K. (2005) Comparing and combining predictors of mostly disordered proteins, Biochemistry 44, 1989-2000), also known as molecular recognition features (MoRFs) (Mohan, A., Oldfield, C. J., Radivojac, P., Vacic, V., Cortese, M. S., Dunker, A. K., and Uversky, V. N. (2006) Analysis of Molecular Recognition Features (MoRFs), J. Mol. Biol. 362, 1043-1059), revealed that regions undergoing disorder-to-order transition are involved in many molecular recognition events and are crucial for protein-protein interactions. However, these algorithms were developed using a training data set of a limited size. Here we propose to improve the prediction algorithms by (1) including additional (alpha-MoRF examples and their cross species homologues in the positive training set, (2) carefully extracting monomer structure chains from the Protein Data Bank (PDB) as the negative training set, (3) including attributes from recently developed disorder predictors, secondary structure predictions, and amino acid indices, and (4) constructing neural network based predictors and performing validation. Over 50 regions which undergo disorder-to-order transition that were identified in the PDB together with a set of corresponding cross species homologues of each structure-based example were included in a new positive training set. Over 1500 attributes, including disorder predictions, secondary structure predictions, and amino acid indices, were evaluated by the conditional probability method. The top attributes, including VSL2 and VL3 disorder predictions and several physicochemical propensities of amino acid residues, were used to develop the feed forward neural networks. The sensitivity, specificity and accuracy of the resulting predictor, alpha-MoRF-PredII, were 0.87 +/- 0.10, 0.87 +/- 0.11, and 0.87 +/- 0.08 over 10 cross validations, respectively. We present the results of these analyses and validation examples to discuss the potential improvement of the (alpha-MoRF-PredII prediction accuracy.