pTop 1.0: A High-Accuracy and High-Efficiency Search Engine for Intact Protein Identification

pTop 1.0: A High-Accuracy and High-Efficiency Search Engine for Intact Protein Identification
复制标题

DOI:
10.1021/acs.analchem.5b03963
复制
发表时间:
2016-03-15
影响因子:
7.4
通讯作者:
He, Si-Min
He, Si-Min
中科院分区:
化学1区
文献类型:
--
作者:
Sun, Rui-Xiang;Luo, Lan;He, Si-Min

文献摘要

被引文献

相似文献

近5年来,自上而下的蛋白质组学研究取得了巨大的进展,特别是在完整蛋白分离和高分辨率质谱分析方面。然而,处理大规模质谱的生物信息学在算法研究和软件开发方面都比较落后。在本研究中,我们开发了一种新的软件工具pTop 1.0,以显着提高TDP质谱数据分析的准确性和效率。前体质量为推断蛋白质上潜在的翻译后修饰提供了重要线索,其可靠性在很大程度上依赖于其质量准确性。为了更准确地检测前体,在pTop中通过支持向量机(SVM)在线训练了一个包含各种光谱特征的机器学习模型。pTop利用从MS/MS谱中提取的序列标签和动态规划算法来加快搜索速度,特别是对于那些有多个翻译后修饰的谱。我们在三个公开可用的数据集上测试了pTop,并将其与ProSight和MS-Align+在召回率、精确度、运行时间等方面进行了比较。结果表明,pTop总体上优于ProSight和MS-Align+。尽管pTop从人类组蛋白数据集中输出的前体比Xtract(在ProSight中)少30%,但其正确的前体召回率提高了22%。pTop的运行速度比MS-Align+快1 ~ 2个数量级。pTop的这种算法的进步,包括准确性和速度,将激发其他类似软件的开发,以分析整个蛋白质的质谱。
There has been tremendous progress in top-down proteomics (TDP) in the past 5 years, particularly in intact protein separation and high resolution mass spectrometry. However, bioinformatics to deal with large-scale mass spectra has lagged behind, in both algorithmic research and software development. In this study, we developed pTop 1.0, a novel software tool to significantly improve the accuracy and efficiency of mass spectral data analysis in TDP. The precursor mass offers crucial clues to infer the potential post translational modifications co-occurring on the protein, the reliability of which relies heavily on its mass accuracy. Concentrating on detecting the precursors more accurately, a machine-learning model incorporating a variety of spectral features was trained online in pTop via a support vector machine (SVM). pTop employs the sequence tags extracted from the MS/MS spectra and a dynamic programming algorithm to accelerate the search speed, especially for those spectra with multiple post-translational modifications. We tested pTop on three publicly available data sets and compared it with ProSight and MS-Align+ in terms of its recall, precision, running time, and so on. The results showed that pTop can, in general, outperform ProSight and MS-Align+. pTop recalled 22% more correct precursors, although it exported 30% fewer precursors than Xtract (in ProSight) from a human histone data set. The running speed of pTop was about 1 to 2 orders of magnitude faster than that of MS-Align+. This algorithmic advancement in pTop, including both accuracy and speed, will inspire the development of other similar software to analyze the mass spectra from the entire proteins.