mzDB: A File Format Using Multiple Indexing Strategies for the Efficient Analysis of Large LC-MS/MS and SWATH-MS Data Sets

mzDB: A File Format Using Multiple Indexing Strategies for the Efficient Analysis of Large LC-MS/MS and SWATH-MS Data Sets
复制标题

DOI:
10.1074/mcp.o114.039115
复制
发表时间:
2015-03-01
影响因子:
7
通讯作者:
Monsarrat, Bernard
Monsarrat, Bernard
中科院分区:
生物学1区
文献类型:
--
作者:
Bouyssie, David;Dubois, Marc;Monsarrat, Bernard

文献摘要

被引文献

相似文献

MS数据的分析和管理,特别是那些由数据独立的MS采集产生的数据,例如SWATH-MS,对蛋白质组学生物信息学提出了重大挑战。这些数据集固有的大尺寸和大量信息需要适当地结构化,以实现用于识别特定靶肽的信号的有效和直接的提取。标准的基于XML的格式不太适合于大型MS数据文件,例如SWATH-MS生成的数据文件,并损害了高吞吐量数据处理和存储。我们开发了mzDB,一种用于大型MS数据集的高效文件格式。它依赖于SQLite软件库,由一个标准化和可移植的无服务器单文件数据库组成。采用优化的3D索引方法,其中LC-MS坐标(保留时间和m/z)以及SWATH-MS数据的前体m/z沿着用于查询数据库以进行数据提取。与XML格式相比,mzDB节省了大约25%的存储空间,并将访问时间缩短了两倍甚至2000倍,这取决于特定的数据访问。类似地,mzDB与mz 5等其他格式相比,访问时间也略有降低。C++和Java实现,将原始或XML格式转换为mzDB并提供访问方法,将在许可证下发布。mzDB可以通过SQLite C库及其所有主要语言的驱动程序轻松访问,并使用现有的专用GUI浏览。这里描述的mzDB可以提升现有的质谱数据分析管道,在效率,便携性,紧凑性和灵活性方面提供前所未有的性能。
The analysis and management of MS data, especially those generated by data independent MS acquisition, exemplified by SWATH-MS, pose significant challenges for proteomics bioinformatics. The large size and vast amount of information inherent to these data sets need to be properly structured to enable an efficient and straightforward extraction of the signals used to identify specific target peptides. Standard XML based formats are not well suited to large MS data files, for example, those generated by SWATH-MS, and compromise high-throughput data processing and storing. We developed mzDB, an efficient file format for large MS data sets. It relies on the SQLite software library and consists of a standardized and portable server-less single-file database. An optimized 3D indexing approach is adopted, where the LC-MS coordinates (retention time and m/z), along with the precursor m/z for SWATH-MS data, are used to query the database for data extraction. In comparison with XML formats, mzDB saves similar to 25% of storage space and improves access times by a factor of twofold up to even 2000-fold, depending on the particular data access. Similarly, mzDB shows also slightly to significantly lower access times in comparison with other formats like mz5. Both C++ and Java implementations, converting raw or XML formats to mzDB and providing access methods, will be released under permissive license. mzDB can be easily accessed by the SQLite C library and its drivers for all major languages, and browsed with existing dedicated GUIs. The mzDB described here can boost existing mass spectrometry data analysis pipelines, offering unprecedented performance in terms of efficiency, portability, compactness, and flexibility.