Centralized data analysis of a large interlaboratory proteomics project: A feasibility study

Centralized data analysis of a large interlaboratory proteomics project: A feasibility study
复制标题

大型实验室间蛋白质组学项目的集中数据分析:可行性研究

DOI:
10.1002/pmic.200401336
复制
发表时间:
2005
期刊:
影响因子:
3.4
通讯作者:
A. Admon
A. Admon
中科院分区:
生物学3区
文献类型:
--
作者:
Ilan Beer;Eilon Barnea;A. Admon

文献摘要

参考文献

被引文献

相似文献

人血浆蛋白质组计划(PPP)是许多实验室之间的大规模合作。PPP中要求最高的任务之一是分析参与者产生的大量原始MS/MS数据。管理这项任务的主要方法是让参与者分析他们自己的数据,并将结果作为已识别蛋白质和肽的列表提交给中央PPP存储库。为了补充这种分布式方法,我们还对参与者提供的原始MS/MS数据进行了集中分析。由于这种项目中固有的数据冗余,集中式分析有可能通过在分析之前减少冗余来减少计算工作量。集中分析还可以统一流程,并利用实验室之间的数据共享来改进蛋白质鉴定和验证。我们采用的过程包括去除低质量光谱,通过相互相似性聚类光谱,并应用统一的肽和蛋白质鉴定程序。为了证明这一过程,我们分析了来自8个实验室的血清和血浆蛋白的胰蛋白酶肽的528万个MS/MS光谱。
The human Plasma Proteome Project (PPP) is a large‐scale collaboration between many laboratories. One of the most demanding tasks in the PPP involved the analysis of very large amounts of raw MS/MS data produced by the participants. The main approach for managing this task was letting the participants analyze their own data and submit the results to the central PPP repository as lists of identified proteins and peptides. To complement this distributed approach, we also performed centralized analysis of the raw MS/MS data provided by the participants. Due to the data redundancy inherent in such a project, centralized analysis has the potential to reduce the computational effort by reducing redundancy before the analysis. Centralized analysis can also unify the process and take advantage of data sharing among laboratories to improve protein identification and validation. The process we employed included removing low‐quality spectra, clustering spectra by mutual similarity, and applying uniform peptide and protein identification procedures. To demonstrate the process, we analyzed 5.28 million MS/MS spectra derived by eight laboratories from tryptic peptides of serum and plasma proteins.
DOI: 10.1021/ac034869m
发表时间: 2004-02-15
影响因子: 7.4
作者:
Shen, YF;Jacobs, JM;Tompkins, RG
通讯作者: Tompkins, RG