pymzML v2.0: introducing a highly compressed and seekable gzip format

pymzML v2.0: introducing a highly compressed and seekable gzip format
复制标题

DOI:
10.1093/bioinformatics/bty046
复制
发表时间:
2018-07-15
期刊:
影响因子:
5.8
通讯作者:
Fufezan, C.
Fufezan, C.
中科院分区:
生物学3区
文献类型:
--
作者:
Koesters, M.;Leufken, J.;Fufezan, C.

文献摘要

被引文献

相似文献

动机:在新版本的pymzML(v2.0)中,我们优化了这一已建立的质谱数据分析工具的速度,以适应质谱数据量的增加。因此,我们集成了更快的数值计算库,改进了数据检索算法,并优化了源代码。重要的是,为了适应快速增长的文件大小,我们开发了一个可推广的压缩方案非常快速的随机访问,并将这一概念应用到mzML文件检索光谱data.Results:pymzML执行与建立的C程序在处理时间。然而,它提供了脚本语言的多功能性,同时增加了对压缩文件的前所未有的快速随机访问。此外,我们设计了我们的压缩方案,在这样一个一般的方式,它可以应用于任何领域的快速随机访问压缩文件中的大数据块是必要的。
Motivation: In the new release of pymzML (v2.0), we have optimized the speed of this established tool for mass spectrometry data analysis to adapt to increasing amounts of data in mass spectrometry. Thus, we integrated faster libraries for numerical calculations, improved data retrieving algorithms and have optimized the source code. Importantly, to adapt to rapidly growing file sizes, we developed a generalizable compression scheme for very fast random access and applied this concept to mzML files to retrieve spectral data.Results: pymzML performs at par with established C programs when it comes to processing times. However, it offers the versatility of a scripting language, while adding unprecedented fast random access to compressed files. Additionally, we designed our compression scheme in such a general way that it can be applied to any field where fast random access to large data blocks in compressed files is desired.