Overview of the HUPO Plasma Proteome Project: Results from the pilot phase with 35 collaborating laboratories and multiple analytical groups, generating a core dataset of 3020 proteins and a publicly-available database

Overview of the HUPO Plasma Proteome Project: Results from the pilot phase with 35 collaborating laboratories and multiple analytical groups, generating a core dataset of 3020 proteins and a publicly-available database
复制标题

DOI:
10.1002/pmic.200500358
复制
发表时间:
2005-08-01
期刊:
影响因子:
3.4
通讯作者:
Hanash, SM
Hanash, SM
中科院分区:
生物学3区
文献类型:
--
作者:
Omenn, GS;States, DJ;Hanash, SM

文献摘要

被引文献

相似文献

人类蛋白质组组织(HUPO)于2002年发起了血浆蛋白质组项目(PPP)。其试点阶段已经(1)评估了许多去除、分级分离和质谱技术平台的优势与局限性;(2)比较了人类血清以及乙二胺四乙酸(EDTA)、肝素和枸橼酸盐抗凝血浆的PPP参考样本;(3)创建了一个公共知识库(www.bioinformatics.med.umich.edu/hupo/ppp;www.ebi.ac.uk/pride)。13个国家的35个参与实验室提交了数据集。各个工作组研究了(a)样本稳定性和蛋白质浓度;(b)来自18个串联质谱数据集的蛋白质鉴定;(c)对原始串联质谱光谱的独立分析;(d)搜索引擎性能、亚蛋白质组分析和生物学见解;(e)抗体阵列;以及(f)直接质谱/表面增强激光解吸电离分析。串联质谱数据集有15710种不同的国际蛋白质索引(IPI)蛋白质编号;我们应用于肽序列多重匹配的整合算法得出了9504种由一种或多种肽鉴定出的IPI蛋白质以及3020种由两种或多种肽鉴定出的蛋白质(核心数据集)。这些蛋白质已通过基因本体论、蛋白质家族数据库(InterPro)、诺华图谱、人类孟德尔遗传在线(OMIM)以及基于免疫测定的浓度测定进行了特征描述。该数据库允许对许多其他子集进行检查,例如1274种由三种或更多种肽鉴定出的蛋白质。蛋白质与DNA的反向匹配为118个先前未鉴定的开放阅读框鉴定出了蛋白质。我们建议使用血浆而非血清,并使用EDTA(或枸橼酸盐)进行抗凝。为了提高肽鉴定和蛋白质匹配的分辨率、灵敏度和重现性,我们建议结合使用去除、分级分离和串联质谱/质谱技术,并明确光谱评估标准、搜索算法的使用以及同源蛋白质匹配的整合。本期《蛋白质组学》特刊展示了合作分析不可或缺的论文以及关于PPP工作计划各个方面的许多补充工作报告。这些关于血浆和血清蛋白质的复杂性、动态范围、不完全采样、假阳性匹配以及不同数据集整合的PPP结果为健康和疾病中循环蛋白质生物标志物的开发和验证奠定了基础。
HUPO initiated the Plasma Proteome Project (PPP) in 2002. Its pilot phase has (1) evaluated advantages and limitations of many depletion, fractionation, and MS technology platforms; (2) compared PPP reference specimens of human serum and EDTA, heparin, and citrate-anticoagulated plasma; and (3) created a publicly-available knowledge base (www.bioinformatics.med.umich.edu/hupo/ppp; www.ebi.ac.uk/pride). Thirty-five participating laboratories in 13 countries submitted datasets. Working groups addressed (a) specimen stability and protein concentrations; (b) protein identifications from 18 MS/MS datasets; (c) independent analyses from raw MS-MS spectra; (d) search engine performance, subproteome analyses, and biological insights; (e) antibody arrays; and (f) direct MS/SELDI analyses. MS-MS datasets had 15 710 different International Protein Index (IPI) protein IDs; our integration algorithm applied to multiple matches of peptide sequences yielded 9504 IPI proteins identified with one or more peptides and 3020 proteins identified with two or more peptides (the Core Dataset). These proteins have been characterized with Gene Ontology, InterPro, Novartis Atlas, OMIM, and immunoassay-based concentration determinations. The database permits examination of many other subsets, such as 1274 proteins identified with three or more peptides. Reverse protein to DNA matching identified proteins for 118 previously unidentified ORFs.We recommend use of plasma instead of serum, with EDTA (or citrate) for anticoagulation. To improve resolution, sensitivity and reproducibility of peptide identifications and protein matches, we recommend combinations of depletion, fractionation, and MS/MS technologies, with explicit criteria for evaluation of spectra, use of search algorithms, and integration of homologous protein matches.This Special Issue of PROTEOMICS presents papers integral to the collaborative analysis plus many reports of supplementary work on various aspects of the PPP workplan. These PPP results on complexity, dynamic range, incomplete sampling, false-positive matches, and integration of diverse datasets for plasma and serum proteins lay a foundation for development and validation of circulating protein biomarkers in health and disease.