Recognizing millions of consistently unidentified spectra across hundreds of shotgun proteomics datasets.

Recognizing millions of consistently unidentified spectra across hundreds of shotgun proteomics datasets.
复制标题

DOI:
10.1038/nmeth.3902
复制
发表时间:
2016-08
期刊:
影响因子:
48
通讯作者:
Vizcaíno JA
Vizcaíno JA
中科院分区:
生物学1区
文献类型:
--
作者:
Griss J;Perez-Riverol Y;Lewis S;Tabb DL;Dianes JA;Del-Toro N;Rurik M;Walzer MW;Kohlbacher O;Hermjakob H;Wang R;Vizcaíno JA

文献摘要

被引文献

相似文献

质谱(MS)是蛋白质组学方法中使用的主要技术。然而,在MS实验中分析的光谱中平均75%仍然未被识别。我们建议使用光谱聚类在一个大的规模,以阐明这些未识别的光谱。PRoteomics IDEntifications database(PRIDE)Archive是全球最大的MS蛋白质组学公共数据库之一。通过对PRIDE Archive中公开提供的来自数百个数据集的所有串联MS光谱进行聚类,我们能够一致地表征三组不同的光谱:1)错误识别的光谱,2)正确识别但低于设定的评分阈值的光谱,以及3)真正未识别的光谱。使用多种互补的分析方法,我们能够识别不到20%的一致未识别的光谱。完整的频谱聚类结果可通过新版PRIDE集群资源(http://www.ebi.ac.uk/pride/cluster)获得。除其他目的外,该资源旨在鼓励和简化对这些未识别光谱的进一步调查。
Mass spectrometry (MS) is the main technology used in proteomics approaches. However, on average 75% of spectra analysed in an MS experiment remain unidentified. We propose to use spectrum clustering at a large-scale to shed a light on these unidentified spectra. PRoteomics IDEntifications database (PRIDE) Archive is one of the largest MS proteomics public data repositories worldwide. By clustering all tandem MS spectra publicly available in PRIDE Archive, coming from hundreds of datasets, we were able to consistently characterize three distinct groups of spectra: 1) incorrectly identified spectra, 2) spectra correctly identified but below the set scoring threshold, and 3) truly unidentified spectra. Using a multitude of complementary analysis approaches, we were able to identify less than 20% of the consistently unidentified spectra. The complete spectrum clustering results are available through the new version of the PRIDE Cluster resource (http://www.ebi.ac.uk/pride/cluster). This resource is intended, among other aims, to encourage and simplify further investigation into these unidentified spectra.