Do we want our data raw? Including binary mass spectrometry data in public proteomics data repositories

Do we want our data raw? Including binary mass spectrometry data in public proteomics data repositories
复制标题

我们想要原始数据吗?

DOI:
--
复制
发表时间:
2005
期刊:
影响因子:
3.4
通讯作者:
K. Gevaert
K. Gevaert
中科院分区:
生物学3区
文献类型:
--
作者:
L. Martens;A. Nesvizhskii;H. Hermjakob;M. Adamski;G. Omenn;J. Vandekerckhove;K. Gevaert

文献摘要

被引文献

相似文献

随着人类血浆蛋白质组计划(PPP)试点阶段的完成,迄今为止最大和最雄心勃勃的蛋白质组学实验已经达到了第一个里程碑。相应地,来自该试点项目的令人印象深刻的数据量强调了对集中传播机制的需求,并导致密歇根大学安阿伯开发了详细的PPP特定数据收集基础设施,以及欧洲生物信息学研究所的蛋白质鉴定数据库项目作为一般蛋白质组学数据库。在讨论为PPP存储哪些数据时出现的一个问题是,是否应该存储来自质谱仪的原始二进制数据,或者更确切地说,是更紧凑且已经显著处理的峰列表。由于这场辩论并不局限于PPP,而是涉及到蛋白质组学社区一般,我们将尝试详细的相对优点和警告与集中存储和传播的原始数据和/或峰值列表,建立在PPP试点阶段获得的广泛经验。最后,提出了一些建议,为当前和未来的存储MS数据在公共知识库。
With the human Plasma Proteome Project (PPP) pilot phase completed, the largest and most ambitious proteomics experiment to date has reached its first milestone. The correspondingly impressive amount of data that came from this pilot project emphasized the need for a centralized dissemination mechanism and led to the development of a detailed, PPP specific data gathering infrastructure at the University of Michigan, Ann Arbor as well as the protein identifications database project at the European Bioinformatics Institute as a general proteomics data repository. One issue that crept up while discussing which data to store for the PPP concerns whether the raw, binary data coming from the mass spectrometers should be stored, or rather the more compact and already significantly processed peak lists. As this debate is not restricted to the PPP but relates to the proteomics community in general, we will attempt to detail the relative merits and caveats associated with centralized storage and dissemination of raw data and/or peak lists, building on the extensive experience gained during the PPP pilot phase. Finally, some suggestions are made for both immediate and future storage of MS data in public repositories.