Aligned genomic data compression via improved modeling.

Aligned genomic data compression via improved modeling.
复制标题

通过改进的建模来对齐基因组数据压缩。

DOI:
10.1142/s0219720014420025
复制
发表时间:
2014
影响因子:
1
通讯作者:
Weissman,Tsachy
Weissman,Tsachy
中科院分区:
生物学4区
文献类型:
--
作者:
Ochoa,Idoia;Hernaez,Mikel;Weissman,Tsachy

文献摘要

相似文献

随着Illumina公司最新的下一代测序(NGS)机器HiSeq X的发布,人类全基因组测序的成本预计将降至仅1000美元。测序史上的这一里程碑标志着个人负担得起的测序时代,并打开了个性化医疗的大门。与此同时,前所未有的基因组数据将需要存储处理。不仅迫切需要压缩对齐的数据,而且还需要生成可以直接提供给下游应用程序的压缩文件,以促进对数据的分析和推断。文献中提出了几种应对这一挑战的方法;然而,到目前为止,重点一直放在低覆盖率制度上,大多数建议的压缩器都不是基于有效的数据建模。我们将演示数据建模对压缩对齐读取的好处。具体来说,我们表明,通过使用为对齐数据设计的数据模型,我们可以大大提高以前提出的算法所达到的最佳压缩比。我们的研究结果表明,Bonfield和Mahoney(2013)提出的压缩率和速度的帕累托最优障碍[Bonfield JK和Mahoney MV], FASTQ和SAM格式测序数据的压缩,PLOS ONE,8(3):e59190, 2013。并不适用于高覆盖率对齐数据。此外,我们改进的压缩比是通过以有利于下游应用程序在压缩域中操作的方式拆分数据来实现的。
With the release of the latest Next-Generation Sequencing (NGS) machine, the HiSeq X by Illumina, the cost of sequencing the whole genome of a human is expected to drop to a mere $1000. This milestone in sequencing history marks the era of affordable sequencing of individuals and opens the doors to personalized medicine. In accord, unprecedented volumes of genomic data will require storage for processing. There will be dire need not only of compressing aligned data, but also of generating compressed files that can be fed directly to downstream applications to facilitate the analysis of and inference on the data. Several approaches to this challenge have been proposed in the literature; however, focus thus far has been on the low coverage regime and most of the suggested compressors are not based on effective modeling of the data.We demonstrate the benefit of data modeling for compressing aligned reads. Specifically, we show that, by working with data models designed for the aligned data, we can improve considerably over the best compression ratio achieved by previously proposed algorithms. Our results indicate that the pareto-optimal barrier for compression rate and speed claimed by Bonfield and Mahoney (2013) [Bonfield JK and Mahoneys MV, Compression of FASTQ and SAM format sequencing data, PLOS ONE,8(3):e59190, 2013.] does not apply for high coverage aligned data. Furthermore, our improved compression ratio is achieved by splitting the data in a manner conducive to operations in the compressed domain by downstream applications.