Evaluating de novo sequencing in proteomics: already an accurate alternative to database-driven peptide identification?

Evaluating de novo sequencing in proteomics: already an accurate alternative to database-driven peptide identification?
复制标题

DOI:
10.1093/bib/bbx033
复制
发表时间:
2018-09-01
影响因子:
9.5
通讯作者:
Renard, Bernhard Y.
Renard, Bernhard Y.
中科院分区:
生物学2区
文献类型:
--
作者:
Muth, Thilo;Renard, Bernhard Y.

文献摘要

被引文献

相似文献

虽然基于质谱(MS)的鸟枪蛋白质组学中的肽鉴定主要使用数据库搜索方法获得,但是来自现代MS仪器的高分辨率光谱数据如今提供了改善计算从头肽测序的性能的前景。从头测序的主要好处是它不需要参考数据库来直接从实验串联质谱光谱推断全长或部分基于标签的肽序列。虽然已经开发了各种算法用于自动从头测序,但在独立的基准研究中很少评估所提出的解决方案的预测准确性。这项工作的主要目的是提供一个详细的从头测序算法的性能评价高分辨率数据。为此目的,我们处理了四个实验数据集,从不同的仪器类型,碰撞诱导解离和高能碰撞解离(HCD)碎裂模式,使用软件包Novor,PEAKS和PepNovo。此外,这些算法的准确性也测试地面实况数据的基础上产生的峰值强度预测软件的模拟光谱。我们发现Novor在正确的全肽、基于标签和单残基预测的准确性方面显示出与PEAKS和PepNovo相比的整体最佳性能。此外,同一个工具在运行时间加速方面超过了商业竞争对手PEAKS大约12-17倍。尽管HCD数据集上完整肽序列的预测准确率约为35%,但作为一个整体,评估的算法在实验数据上表现适中,但在模拟数据上表现出明显更好的性能(高达84%的准确率)。此外,我们描述了最常见的从头测序错误,并评估了丢失的碎片离子峰和光谱噪声对准确度的影响。最后,我们讨论了从头测序的潜力,现在成为更广泛地应用于该领域。
While peptide identifications in mass spectrometry (MS)-based shotgun proteomics are mostly obtained using database search methods, high-resolution spectrum data from modern MS instruments nowadays offer the prospect of improving the performance of computational de novo peptide sequencing. The major benefit of de novo sequencing is that it does not require a reference database to deduce full-length or partial tag-based peptide sequences directly from experimental tandem mass spectrometry spectra. Although various algorithms have been developed for automated de novo sequencing, the prediction accuracy of proposed solutions has been rarely evaluated in independent benchmarking studies. The main objective of this work is to provide a detailed evaluation on the performance of de novo sequencing algorithms on high-resolution data. For this purpose, we processed four experimental data sets acquired from different instrument types from collision-induced dissociation and higher energy collisional dissociation (HCD) fragmentation mode using the software packages Novor, PEAKS and PepNovo. Moreover, the accuracy of these algorithms is also tested on ground truth data based on simulated spectra generated from peak intensity prediction software. We found that Novor shows the overall best performance compared with PEAKS and PepNovo with respect to the accuracy of correct full peptide, tag-based and single-residue predictions. In addition, the same tool outpaced the commercial competitor PEAKS in terms of running time speedup by factors of around 12-17. Despite around 35% prediction accuracy for complete peptide sequences on HCD data sets, taken as a whole, the evaluated algorithms perform moderately on experimental data but show a significantly better performance on simulated data (up to 84% accuracy). Further, we describe the most frequently occurring de novo sequencing errors and evaluate the influence of missing fragment ion peaks and spectral noise on the accuracy. Finally, we discuss the potential of de novo sequencing for now becoming more widely used in the field.