Improved recognition of native-like protein structures using a combination of sequence-dependent and sequence-independent features of proteins

Improved recognition of native-like protein structures using a combination of sequence-dependent and sequence-independent features of proteins
复制标题

DOI:
10.1002/(sici)1097-0134(19990101)34:1
复制
发表时间:
1999-01-01
影响因子:
2.9
通讯作者:
Baker, D
Baker, D
中科院分区:
生物学4区
文献类型:
--
作者:
Simons, KT;Ruczinski, I;Baker, D

文献摘要

被引文献

相似文献

我描述了基于与 P(序列\结构) *P(结构) 成正比的分解 P(结构\序列) 的评分函数的开发,该评分函数在正确识别大型紧凑诱饵集合中的类天然蛋白质结构方面优于以前的评分函数。第一个术语捕获蛋白质结构的序列依赖性特征,例如核心中疏水残基的埋藏,第二个术语捕获通用的序列独立特征,例如将β链组装成β折叠。使用 17 个小蛋白质结构域中每一个具有固定二级结构的 30,000 个类似紧凑构象的集合,系统地评估了蛋白质结构的各种序列依赖性和序列独立性特征在识别天然样结构方面的功效。最好的结果是使用核心评分函数获得的,其中 P(序列\结构)参数化与我们之前的工作类似(Simons 等人,J Mol Biol 1997;268:209-225),P(结构)专注于二级结构包装偏好;虽然几个附加特征本身具有一定的区分能力,但与核心评分函数结合时,它们没有提供任何额外的区分能力。我们的结果,在训练集和独立诱饵上Park 和 Levitt 的一组 (J Mol Biol 1996;258:367-392) 表明,该评分函数应有助于根据序列和二级结构的知识预测三级结构 (C) 1999 Wiley-Liss, Inc.。
me describe the development of a scoring function based on the decomposition P(structure\sequence) proportional to P(sequence\structure) *P(structure), which outperforms previous scoring functions in correctly identifying native-like protein structures in large ensembles of compact decoys. The first term captures sequence-dependent features of protein structures, such as the burial of hydrophobic residues in the core, the second term, universal sequence-independent features, such as the assembly of beta-strands into beta-sheets. The efficacies of a wide variety of sequence-dependent and sequence-independent features of protein structures for recognizing native-like structures were systematically evaluated using ensembles of similar to 30,000 compact conformations with fixed secondary structure for each of 17 small protein domains. The best results were obtained using a core scoring function with P(sequence\structure) parameterized similarly to our previous work (Simons et al., J Mol Biol 1997;268:209-225] and P(structure) focused on secondary structure packing preferences; while several additional features had some discriminately power on their own, they did not provide any additional discriminatory power when combined with the core scoring function. Our results, on both the training set and the independent decoy set of Park and Levitt (J Mol Biol 1996;258:367-392), suggest that this scoring function should contribute to the prediction of tertiary structure from knowledge of sequence and secondary structure. (C) 1999 Wiley-Liss, Inc.