Accurate estimation of isoelectric point of protein and peptide based on amino acid sequences.

Accurate estimation of isoelectric point of protein and peptide based on amino acid sequences.
复制标题

DOI:
10.1093/bioinformatics/btv674
复制
发表时间:
2016-03-15
期刊:
Bioinformatics (Oxford, England)
影响因子:
--
通讯作者:
Perez-Riverol Y
Perez-Riverol Y
中科院分区:
其他
文献类型:
--
作者:
Audain E;Ramos Y;Hermjakob H;Flower DR;Perez-Riverol Y

文献摘要

参考文献

被引文献

相似文献

动机:在任何大分子多质子系统中-例如蛋白质、DNA或RNA-等电点-通常称为pI--可以定义为滴定曲线中的奇点,对应于溶液pH值,在该pH值下两性电解质的净总表面电荷以及电泳迁移率总和为零。不同的现代分析生物化学和蛋白质组学方法依赖于等电点作为蛋白质和肽表征的主要特征。通过等电点分离蛋白质是2-D凝胶电泳的关键部分,2-D凝胶电泳是蛋白质组学的关键前体,其中离散点可以在凝胶中消化,随后通过分析质谱法鉴定蛋白质。在LC-MS/MS分析之前,根据其pI的肽分级也广泛用于当前蛋白质组学样品制备程序中。因此,准确的理论预测pI将加快这种分析。虽然这种pI计算被广泛使用,但它在很大程度上仍未经过测试,这促使我们努力对pI预测方法进行基准测试。 结果如下:使用来自数据库PIP-DB的数据和一个实验室可用的数据集作为我们的参考金标准,我们对pI计算方法进行了基准测试。我们发现,方法的准确性各不相同,并且对基组的选择非常敏感。机器学习算法,特别是基于SVM的算法,在研究肽混合物时表现出上级性能。一般来说,基于学习的pI预测方法(如Cofactor,SVM和Branca)需要大量的训练数据集,其结果性能将在很大程度上取决于该数据的质量。与迭代方法相比,机器学习算法具有能够添加新特征以提高预测准确性的优势。 联系方式:yperez@ebi.ac.uk可用性和实施:软件和数据可在https://github.com/ypriverol/pIR免费获得。 补充信息:补充数据可从在线生物信息学获得。
Motivation: In any macromolecular polyprotic system—for example protein, DNA or RNA—the isoelectric point—commonly referred to as the pI—can be defined as the point of singularity in a titration curve, corresponding to the solution pH value at which the net overall surface charge—and thus the electrophoretic mobility—of the ampholyte sums to zero. Different modern analytical biochemistry and proteomics methods depend on the isoelectric point as a principal feature for protein and peptide characterization. Protein separation by isoelectric point is a critical part of 2-D gel electrophoresis, a key precursor of proteomics, where discrete spots can be digested in-gel, and proteins subsequently identified by analytical mass spectrometry. Peptide fractionation according to their pI is also widely used in current proteomics sample preparation procedures previous to the LC-MS/MS analysis. Therefore accurate theoretical prediction of pI would expedite such analysis. While such pI calculation is widely used, it remains largely untested, motivating our efforts to benchmark pI prediction methods. Results: Using data from the database PIP-DB and one publically available dataset as our reference gold standard, we have undertaken the benchmarking of pI calculation methods. We find that methods vary in their accuracy and are highly sensitive to the choice of basis set. The machine-learning algorithms, especially the SVM-based algorithm, showed a superior performance when studying peptide mixtures. In general, learning-based pI prediction methods (such as Cofactor, SVM and Branca) require a large training dataset and their resulting performance will strongly depend of the quality of that data. In contrast with Iterative methods, machine-learning algorithms have the advantage of being able to add new features to improve the accuracy of prediction. Contact: yperez@ebi.ac.uk Availability and Implementation: The software and data are freely available at https://github.com/ypriverol/pIR. Supplementary information: Supplementary data are available at Bioinformatics online.
DOI: 10.1002/elps.200700701
发表时间: 2008-07-01
期刊: ELECTROPHORESIS
影响因子: 2.9
作者:
Cargile, Benjamin J.;Sevinsky, Joel R.;Stephenson, James L., Jr.
通讯作者: Stephenson, James L., Jr.
DOI: 10.1002/elps.11501401163
发表时间: 1993-10-01
期刊: ELECTROPHORESIS
影响因子: 2.9
作者:
BJELLQVIST, B;HUGHES, GJ;HOCHSTRASSER, D
通讯作者: HOCHSTRASSER, D
DOI: 10.1093/bioinformatics/btu637
发表时间: 2015-01-15
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Bunkute, Egle;Cummins, Christopher;Flower, Darren R.
通讯作者: Flower, Darren R.
DOI: 10.6026/97320630002101
发表时间: 2007-01-01
期刊: BIOINFORMATION
影响因子: 1.9
作者:
Carugo, Oliviero
通讯作者: Carugo, Oliviero
DOI: 10.1016/j.bbapap.2013.02.032
发表时间: 2014-01
期刊: Biochimica et biophysica acta
影响因子: --
作者:
Perez-Riverol Y;Wang R;Hermjakob H;Müller M;Vesada V;Vizcaíno JA
通讯作者: Vizcaíno JA