Predicting disulfide connectivity from protein sequence using multiple sequence feature vectors and secondary structure

Predicting disulfide connectivity from protein sequence using multiple sequence feature vectors and secondary structure
复制标题

DOI:
10.1093/bioinformatics/btm505
复制
发表时间:
2007-12-01
期刊:
影响因子:
5.8
通讯作者:
Burrage, Kevin
Burrage, Kevin
中科院分区:
生物学3区
文献类型:
--
作者:
Song, Jiangning;Yuan, Zheng;Burrage, Kevin

文献摘要

被引文献

相似文献

动机:二硫键是蛋白质中两个半胱氨酸残基之间的主要共价交联,它们在稳定蛋白质结构中起着关键作用,并且通常在外胞外质或分泌的蛋白质中发现。在蛋白质折叠预测中,二硫键的定位可以大大减少构象空间中的搜索。因此,非常需要开发能够准确预测可能具有潜在重要应用的蛋白质中的二硫键连接模式的计算方法。回报:我们已经开发了一种新的方法来预测蛋白质主要序列的二硫键连通性模式,并使用支持矢量回归进行了支持。 (SVR)基于多个序列特征向量和通过PSIPRED程序预测二级结构的方法。结果表明,当使用4倍的跨验验作蛋白质和半胱氨酸对测量的蛋白质上,我们的方法可以分别达到74.4和77.9的预测准确性,分别达到74.4和77.9。同源数据集。我们评估了不同序列编码方案对二硫键连通性预测性能的影响。已经表明,基于多个序列特征向量与预测的二级结构相结合的序列编码方案可以显着提高预测准确性,从而使我们的方法能够胜过大多数其他当前可用的预测指标。我们的工作为当前算法提供了一种补充方法,该方法应在计算分配二硫键模式中有用,并有助于大规模全基因组项目产生的蛋白质序列的注释。
Motivation: Disulfide bonds are primary covalent crosslinks between two cysteine residues in proteins that play critical roles in stabilizing the protein structures and are commonly found in extracy-toplasmatic or secreted proteins. In protein folding prediction, the localization of disulfide bonds can greatly reduce the search in conformational space. Therefore, there is a great need to develop computational methods capable of accurately predicting disulfide connectivity patterns in proteins that could have potentially important applications.Results: We have developed a novel method to predict disulfide connectivity patterns from protein primary sequence, using a support vector regression (SVR) approach based on multiple sequence feature vectors and predicted secondary structure by the PSIPRED program. The results indicate that our method could achieve a prediction accuracy of 74.4 and 77.9, respectively, when averaged on proteins with two to five disulfide bridges using 4-fold cross-validation, measured on the protein and cysteine pair on a well-defined non-homologous dataset. We assessed the effects of different sequence encoding schemes on the prediction performance of disulfide connectivity. It has been shown that the sequence encoding scheme based on multiple sequence feature vectors coupled with predicted secondary structure can significantly improve the prediction accuracy, thus enabling our method to outperform most of other currently available predictors. Our work provides a complementary approach to the current algorithms that should be useful in computationally assigning disulfide connectivity patterns and helps in the annotation of protein sequences generated by large-scale whole-genome projects.