In-depth analysis of protein inference algorithms using multiple search engines and well-defined metrics

In-depth analysis of protein inference algorithms using multiple search engines and well-defined metrics
复制标题

DOI:
10.1016/j.jprot.2016.08.002
复制
发表时间:
2017-01-06
影响因子:
3.3
通讯作者:
Perez-Riverol, Yasset
Perez-Riverol, Yasset
中科院分区:
生物学2区
文献类型:
--
作者:
Audain, Enrique;Uszkoreit, Julian;Perez-Riverol, Yasset

文献摘要

被引文献

相似文献

在基于质谱的鸟枪法蛋白质组学中,蛋白质鉴定通常是期望的结果。然而,大多数分析方法都是基于可靠肽的鉴定,而不是完整蛋白质的直接鉴定。因此,将从串联质谱鉴定的肽组装成蛋白质列表,称为蛋白质推断,是蛋白质组学研究的关键步骤。目前,不同的蛋白质推理算法和工具可用于蛋白质组学社区。在这里,我们评估了五个软件工具的蛋白质推断(PIA,ProteinProphet,Fido,ProteinLP,MSBayesPro)使用三个流行的数据库搜索引擎:吉祥物,X! Tandem和MS-GF+。所有算法都使用高度可定制的KNIME工作流程进行评估,使用四种不同的公共数据集,具有不同的复杂性(不同的样品制备,物种和分析仪器)。我们定义了一组质量控制指标来评估搜索引擎、蛋白质推理算法和参数在每个数据集上的每个组合的性能。我们发现,复杂样品的结果不仅关于报告的蛋白质组的实际数量,而且关于组的实际组成不同。此外,当使用不同复杂性的数据库时,报告的蛋白质的鲁棒性强烈依赖于所应用的推理算法。最后,合并多个搜索引擎的标识并不一定会增加报告的蛋白质的数量,但确实会增加每个蛋白质的肽的数量,因此通常可以recommended.Significance:蛋白质推断是当今基于MS的蛋白质组学的主要挑战之一。目前,有大量的蛋白质推理算法和实现可用于蛋白质组学社区。蛋白质组装影响研究的最终结果,定量值和研究手稿中的最终声明。尽管蛋白质推断是蛋白质组学数据分析中的关键步骤,但从未对许多不同的推断方法进行过全面评估。此前《蛋白质组学杂志》已经发表了多篇关于生物信息学算法其他基准的研究(PMID:26585461; PMID:22728601)在蛋白质组学研究中的重要性,明确了这些研究对蛋白质组学社区和期刊读者的重要性。这篇手稿提出了一种基于KNIME/OpenMS平台的新的生物信息学解决方案,旨在提供蛋白质推理算法的公平比较(https://github.com/KNIME-OMICS).六种不同的算法-ProteinProphet,MSBayesPro,ProteinLP,Fido和PIA-使用高度可定制的工作流程在四个具有不同复杂性的公共数据集上进行了评估。五个流行的数据库搜索引擎吉祥物,X!针对每种蛋白质推断工具评估串联、MS-GF+及其组合。总共分析了> 186个蛋白质列表,并使用三个度量仔细比较蛋白质推断结果的质量评估:1)报告的蛋白质的数量,2)每种蛋白质的肽,和3)每种推断方法独特报告的蛋白质的数量,以解决每种推断方法的质量。我们还通过选择搜索引擎、蛋白质推理算法和每个数据集上的参数的每种组合来检查报告了多少蛋白质。结果表明,在研究分析工作流程的结果时,使用1)PIA或Fido似乎是一个很好的选择,不仅考虑到报告的蛋白质和高质量的鉴定,而且考虑到所需的运行时间。2)合并多个搜索引擎的识别几乎总是提供更可靠的结果,并增加每个蛋白质组的肽的数量。3)使用的数据库,不仅包含典型的,但也已知的蛋白质亚型的报告蛋白质的数量有一个小的影响。关于研究背后的问题,特定亚型的检测可以弥补使用简约报告的略短报告。4)目前的工作流程可以很容易地扩展,以支持新的算法和搜索引擎组合。(C)2016由Elsevier B.V.出版
In mass spectrometry-based shotgun proteomics, protein identifications are usually the desired result. However, most of the analytical methods are based on the identification of reliable peptides and not the direct identification of intact proteins. Thus, assembling peptides identified from tandem mass spectra into a list of proteins, referred to as protein inference, is a critical step in proteomics research. Currently, different protein inference algorithms and tools are available for the proteomics community. Here, we evaluated five software tools for protein inference (PIA, ProteinProphet, Fido, ProteinLP, MSBayesPro) using three popular database search engines: Mascot, X!Tandem, and MS-GF+. All the algorithms were evaluated using a highly customizable KNIME workflow using four different public datasets with varying complexities (different sample preparation, species and analytical instruments). We defined a set of quality control metrics to evaluate the performance of each combination of search engines, protein inference algorithm, and parameters on each dataset. We show that the results for complex samples vary not only regarding the actual numbers of reported protein groups but also concerning the actual composition of groups. Furthermore, the robustness of reported proteins when using databases of differing complexities is strongly dependant on the applied inference algorithm. Finally, merging the identifications of multiple search engines does not necessarily increase the number of reported proteins, but does increase the number of peptides per protein and thus can generally be recommended.Significance: Protein inference is one of the major challenges in MS-based proteomics nowadays. Currently, there are a vast number of protein inference algorithms and implementations available for the proteomics community. Protein assembly impacts in the final results of the research, the quantitation values and the final claims in the research manuscript. Even though protein inference is a crucial step in proteomics data analysis, a comprehensive evaluation of the many different inference methods has never been performed. Previously journal of proteomics has published multiple studies about other benchmark of bioinformatics algorithms (PMID: 26585461; PMID: 22728601) in proteomics studies making clear the importance of those studies for the proteomics community and the journal audience.This manuscript presents a new bioinformatics solution based on the KNIME/OpenMS platform that aims at providing a fair comparison of protein inference algorithms (https://github.com/KNIME-OMICS). Six different algorithms - ProteinProphet, MSBayesPro, ProteinLP, Fido and PIA- were evaluated using the highly customizable workflow on four public datasets with varying complexities. Five popular database search engines Mascot, X!Tandem, MS-GF+ and combinations thereof were evaluated for every protein inference tool. In total >186 proteins lists were analyzed and carefully compare using three metrics for quality assessments of the protein inference results: 1) the numbers of reported proteins, 2) peptides per protein, and the 3) number of uniquely reported proteins per inference method, to address the quality of each inference method. We also examined how many proteins were reported by choosing each combination of search engines, protein inference algorithms and parameters on each dataset.The results show that using 1) PIA or Fido seems to be a good choice when studying the results of the analyzed workflow, regarding not only the reported proteins and the high-quality identifications, but also the required runtime. 2) Merging the identifications of multiple search engines gives almost always more confident results and increases the number of peptides per protein group. 3) The usage of databases containing not only the canonical, but also known isoforms of proteins has a small impact on the number of reported proteins. The detection of specific isoforms could, concerning the question behind the study, compensate for slightly shorter reports using the parsimonious reports. 4) The current workflow can be easily extended to support new algorithms and search engine combinations. (C) 2016 Published by Elsevier B.V.