Efficient storage of high throughput DNA sequencing data using reference-based compression

Efficient storage of high throughput DNA sequencing data using reference-based compression
复制标题

DOI:
10.1101/gr.114819.110
复制
发表时间:
2011-05-01
期刊:
影响因子:
7
通讯作者:
Birney, Ewan
Birney, Ewan
中科院分区:
生物学1区
文献类型:
--
作者:
Fritz, Markus Hsi-Yang;Leinonen, Rasko;Birney, Ewan

文献摘要

被引文献

相似文献

在DNA序列数据的创建和分析中,数据存储成本已成为总成本中一个可观的比例。特别值得关注的是,DNA测序的增长速度大大超过了磁盘存储容量的增长速度。本文提出了一种新的基于参考的压缩方法,可以有效地压缩DNA序列进行存储。我们的方法适用于以研究充分的基因组为目标的重测序实验。我们将新序列与参考基因组对齐,然后将新序列与参考基因组之间的差异编码存储。当我们允许在保存质量信息和未对齐序列时控制数据丢失时,我们的压缩方法是最有效的。使用这种新的压缩方法,我们观察到随着读取长度的增加,效率呈指数增长,并且这种效率增长的幅度可以通过改变存储的质量信息的数量来控制。我们的压缩方法是可调的:质量分数和未对齐序列的存储可以根据不同的实验进行调整,以保存信息或最小化存储成本,并提供一个机会来解决增加DNA序列体积将克服我们存储序列的能力的威胁。
Data storage costs have become an appreciable proportion of total cost in the creation and analysis of DNA sequence data. Of particular concern is that the rate of increase in DNA sequencing is significantly outstripping the rate of increase in disk storage capacity. In this paper we present a new reference-based compression method that efficiently compresses DNA sequences for storage. Our approach works for resequencing experiments that target well-studied genomes. We align new sequences to a reference genome and then encode the differences between the new sequence and the reference genome for storage. Our compression method is most efficient when we allow controlled loss of data in the saving of quality information and unaligned sequences. With this new compression method we observe exponential efficiency gains as read lengths increase, and the magnitude of this efficiency gain can be controlled by changing the amount of quality information stored. Our compression method is tunable: The storage of quality scores and unaligned sequences may be adjusted for different experiments to conserve information or to minimize storage costs, and provides one opportunity to address the threat that increasing DNA sequence volumes will overcome our ability to store the sequences.