SOLpro: accurate sequence-based prediction of protein solubility

SOLpro: accurate sequence-based prediction of protein solubility
复制标题

DOI:
10.1093/bioinformatics/btp386
复制
发表时间:
2009-09-01
期刊:
影响因子:
5.8
通讯作者:
Baldi, Pierre
Baldi, Pierre
中科院分区:
生物学3区
文献类型:
--
作者:
Magnan, Christophe N.;Randall, Arlo;Baldi, Pierre

文献摘要

被引文献

相似文献

动机:蛋白质不溶性是许多实验研究的主要障碍。一种基于序列的预测方法能够准确地预测蛋白质在过表达时可溶的倾向,例如,可以用于大规模蛋白质组学项目中的优先目标,并识别可能增加不溶性蛋白质溶解度的突变。结果:在这里,我们首先策划了一个大型的,非冗余和平衡的训练集,超过17000个蛋白质。接下来,我们提取并研究了23组直接计算或预测的特征(例如二级结构)。数据和特征用于训练两阶段支持向量机(SVM)架构。由此产生的预测,SOLpro,直接与现有的方法进行比较,并显示出显着的改善,根据标准的评估指标,超过74%的整体准确性估计使用多个运行的10倍交叉验证。
Motivation: Protein insolubility is a major obstacle for many experimental studies. A sequence-based prediction method able to accurately predict the propensity of a protein to be soluble on overexpression could be used, for instance, to prioritize targets in large-scale proteomics projects and to identify mutations likely to increase the solubility of insoluble proteins.Results: Here, we first curate a large, non-redundant and balanced training set of more than 17 000 proteins. Next, we extract and study 23 groups of features computed directly or predicted (e.g. secondary structure) from the primary sequence. The data and the features are used to train a two-stage support vector machine (SVM) architecture. The resulting predictor, SOLpro, is compared directly with existing methods and shows significant improvement according to standard evaluation metrics, with an overall accuracy of over 74% estimated using multiple runs of 10-fold cross-validation.