Comparison and Evaluation of Clustering Algorithms for Tandem Mass Spectra

Comparison and Evaluation of Clustering Algorithms for Tandem Mass Spectra
复制标题

DOI:
10.1021/acs.jproteome.7b00427
复制
发表时间:
2017-11-01
影响因子:
4.4
通讯作者:
Rahnenfuehrer, Joerg
Rahnenfuehrer, Joerg
中科院分区:
生物学2区
文献类型:
--
作者:
Rieder, Vera;Schork, Karin U.;Rahnenfuehrer, Joerg

文献摘要

被引文献

相似文献

在蛋白质组学中,建立了液相色谱串联质谱(LC MS/MS)来鉴定肽和蛋白质。重复的光谱,即同一肽的多个光谱,在单次 MS/MS 运行和大型光谱库中都会出现。聚类串联质谱用于寻找共有光谱,具有多种应用。首先,它加快了数据库搜索速度,例如由 Mascot 执行的。其次,它有助于识别跨物种的新型肽。第三,它用于质量控制以检测错误注释的光谱。我们根据光谱之间的余弦距离比较不同的聚类算法。 CAST、MS-Cluster 和 PRIDE Cluster 是常用的串联质谱聚类算法。我们添加了用于大型数据集、层次聚类、DBSCAN 和图的连接组件的著名算法,以及新方法 N-Cluster。所有算法均根据具有不同参数设置的真实数据进行评估。将聚类结果相互比较,并根据纯度等验证措施与肽注释进行比较。针对示例性结果簇讨论了关于检测错误(未)注释的光谱的质量控制。事实证明,N-Cluster 具有很强的竞争力。所有聚类结果都受益于所谓的 DISMS2 过滤器,该过滤器集成了附加信息,例如前体质量信息。
In proteomics, liquid chromatography tandem mass spectrometry (LC MS/MS) is established for identifying peptides and proteins. Duplicated spectra, that is, multiple spectra of the same peptide, occur both in single MS/MS runs and in large spectral libraries. Clustering tandem mass spectra is used to find consensus spectra, with manifold applications. First, it speeds up database searches, as performed for instance by Mascot. Second, it helps to identify novel peptides across species. Third, it is used for quality control to detect wrongly annotated spectra. We compare different clustering algorithms based on the cosine distance between spectra. CAST, MS-Cluster, and PRIDE Cluster are popular algorithms to cluster tandem mass spectra. We add well-known algorithms for large data sets, hierarchical clustering, DBSCAN, and connected components of a graph, as well as the new method N-Cluster. All algorithms are evaluated on real data with varied parameter settings. Cluster results are compared with each other and with peptide annotations based on validation measures such as purity. Quality control, regarding the detection of wrongly (un)annotated spectra, is discussed for exemplary resulting clusters. N-Cluster proves to be highly competitive. All clustering results benefit from the so-called DISMS2 filter that integrates additional information, for example, on precursor mass.