MS1, MS2, and SQT - three unified, compact, and easily parsed file formats for the storage of shotgun proteomic spectra and identifications

MS1, MS2, and SQT - three unified, compact, and easily parsed file formats for the storage of shotgun proteomic spectra and identifications
复制标题

DOI:
10.1002/rcm.1603
复制
发表时间:
2004-01-01
影响因子:
2
通讯作者:
Yates, JR
Yates, JR
中科院分区:
化学3区
文献类型:
--
作者:
McDonald, WH;Tabb, DL;Yates, JR

文献摘要

被引文献

相似文献

随着蛋白质组实验室生成数据的速度随着其正在进行的项目规模的增加而增加,由此产生的数据存储和数据处理问题将继续对计算资源提出挑战。对于鸟枪法蛋白质组学技术来说尤其如此,每天每台仪器可以生成数万个光谱。导致许多这些问题的一个设计因素是由于将光谱和给定光谱的数据库标识存储为单独的文件而引起的。虽然可以通过将所有光谱和搜索结果存储在大型关系数据库中来解决这些问题,但实施这种策略的基础设施可能超出了学术实验室的能力。我们在此报告了一系列用于存储光谱数据(MS1 和 MS2)和搜索结果(SQT)的统一文本文件格式,这些格式紧凑,易于机器和人类解析,但足够灵活,可以与新算法和数据挖掘策略相结合。版权所有 (C) 2004 John Wiley Sons, Ltd.
As the speed with which proteomic labs generate data increases along with the scale of projects they are undertaking, the resulting data storage and data processing problems will continue to challenge computational resources. This is especially true for shotgun proteomic techniques that can generate tens of thousands of spectra per instrument each day. One design factor leading to many of these problems is caused by storing spectra and the database identifications for a given spectrum as individual files. While these problems can be addressed by storing all of the spectra and search results in large relational databases, the infrastructure to implement such a strategy can be beyond the means of academic labs. We report here a series of unified text file formats for storing spectral data (MS1 and MS2) and search results (SQT) that are compact, easily parsed by both machine and humans, and yet flexible enough to be coupled with new algorithms and data-mining Strategies. Copyright (C) 2004 John Wiley Sons, Ltd.